Together AI has announced significant enhancements to its fine-tuning service, designed to streamline the process of adapting open-weight models for production use. The core expansion focuses on providing engineers with greater visibility and control over their training workflows. This includes support for a growing selection of over 40 models, including GLM-5.3, Kimi K2.7, Qwen 3.8-27B, and Gemma 4, allowing teams to quickly integrate the latest advancements in model architecture.
Key improvements center around real-time monitoring. The service now offers live experiment tracking, enabling users to observe training progress directly through the Together API, CLI, and UI dashboard. Metrics, including loss, gradient norm, and learning rate, are recorded at every step and exposed for analysis. This allows for immediate adjustments to the training recipe, such as modifying hyperparameters or selecting the optimal checkpoint based on validation loss.
Furthermore, the service introduces features for enhanced data management. Users can now inspect and validate datasets before training, and the system automatically stops runs when validation loss plateaus, preventing unnecessary training costs. A notable technical advancement is the introduction of expert adapters for Mixture-of-Experts models, which allows training of the model’s core knowledge layers, resulting in up to 89% knowledge retention compared to 15% with attention-only adapters, as demonstrated through MMLU-Pro benchmarks.
The service also offers greater flexibility in batch size handling, accommodating large models and long sequences by utilizing gradient accumulation. Finally, Together AI is reducing prices on selected models to further incentivize adoption and streamline cost management. These changes are intended to accelerate the deployment of high-performing models across a variety of applications.
.png)


