Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Fine-tuning search agents with multi-turn reinforcement learning on Amazon SageMaker AI - OpenSmartRoute
Fine-tuning search agents with multi-turn reinforcement learning on Amazon SageMaker AI
Amazon SageMaker AI now supports fine-tuning search agents using multi-turn reinforcement learning, improving retrieval quality and reliability with minimal setup.
Key points
Supports models like Qwen3.6-27B in the US West (Oregon) region
Improved retrieval scores: +23.7% on BrowseComp-Plus, +18.4% on WixQA
Failure rate on some benchmarks dropped from 22.89% to 0.68%
Uses reward based on nDCG@10 for training and evaluation
Why it matters: Improves search agent quality and reliability while reducing cost and latency through fine-tuning smaller models.
By OpenSmartRoute editorial · written through the router by llm-small
From AWS machine learning blog - “Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI”
Line chart of nDCG@10 reward rising over training steps for both the training and validation sets. Image: AWS machine learning blog (original)
Introduction - Overview of multi-turn reinforcement learning for search agents
Recently, Amazon SageMaker AI introduced a new way to improve search agents using multi-turn reinforcement learning (MTRL). This method helps search systems learn better strategies across multiple rounds of interaction. Instead of just one response, the agent makes a series of decisions, refining its search each time. This approach aims to make search agents more reliable and effective in real-world tasks.
Traditional search models often struggle with multi-step behavior. They may not understand how to decide what to search for next or when to stop. Fine-tuning models with reinforcement learning allows them to learn these behaviors from experience. This method is especially useful for enterprise search, where accuracy and reliability matter a lot. It also helps smaller models perform like larger, more expensive ones, but at a lower cost and faster speed.
What Amazon SageMaker AI MTRL is and how it works
Amazon SageMaker AI MTRL is a tool that helps developers fine-tune large language models (LLMs) using reinforcement learning in multi-turn settings. It treats the search task as a sequence of decisions, where each step depends on the previous ones. The system generates training data through multi-turn rollouts, which simulate real interactions. It then uses policy gradient algorithms to optimize the model’s decision-making process.
This system makes it easy to integrate custom rewards and tools. Developers can define what success looks like, such as retrieving relevant documents. They can also connect different search tools, like keyword or semantic search, into the training process. The interface is designed to require little coding, making it accessible for teams with limited reinforcement learning experience.
The platform supports serverless execution, meaning there is no need to manage hardware or clusters. It runs at per-token pricing, which can reduce costs significantly. During training, multiple rollouts happen in parallel, speeding up the process. The system also allows for long training runs by splitting them into smaller jobs. This flexibility helps teams improve their models without extensive infrastructure management.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.
Google launched EmbeddingGemma 2, a compact open-weight model that maps text, images, and audio into one vector space. It runs locally on phones with minimal RAM and enables instant on-device semantic search.
Key features of SageMaker AI MTRL: modular interface, serverless, parallel rollout
One of the main advantages of SageMaker AI MTRL is its modular interface. Users can easily define custom rewards, tool loops, and conversation shapes. This flexibility allows the system to adapt to different search environments and goals. For example, a team can set rewards based on retrieval quality or task completion speed.
The platform runs in a serverless mode, removing the need for provisioning or managing GPU clusters. This setup simplifies deployment and reduces operational overhead. It also makes it easier to scale training up or down depending on the project needs.
Another key feature is asynchronous rollout and trajectory collection. Multiple generation processes and gradient updates happen simultaneously, with controlled off-policy staleness. This means training remains fast and efficient without drifting away from the current policy. The system also provides observability tools, allowing users to inspect what the agent did during each turn and training step.
SageMaker AI MTRL includes a native library of algorithms. Developers can choose from options like Proximal Policy Optimization (PPO), Clipped Importance Sampling Policy Optimization (CISPO), and importance-sampling (IS) losses. These algorithms are paired with advantage estimators, such as group-based methods, to improve training stability. The system supports resumable training, so long runs can be split and continued later.
The platform also offers evaluation jobs that report metrics like reward, pass@k, and trajectory quality. These metrics help assess how well the model is performing before deployment. The observability and evaluation features make it easier to monitor progress and adjust training as needed.
Supported algorithms and training setup details
SageMaker AI MTRL supports several reinforcement learning algorithms suitable for training search agents. The most common are Proximal Policy Optimization (PPO), Clipped Importance Sampling Policy Optimization (CISPO), and importance-sampling (IS) losses. These algorithms help the model learn decision policies that maximize the reward signal over multiple turns.
Training setup involves a few key steps. First, datasets are prepared and uploaded to Amazon S3 in a format compatible with MTRL. These datasets include examples for training and validation. The training process uses a small set of hyperparameters, such as the number of epochs, batch size, and rollout concurrency. These are the main settings that influence training speed and quality.
The training job runs in the cloud, with multiple rollouts happening in parallel. Developers can monitor progress through metrics and logs. The system also allows for checkpointing, so training can be paused and resumed without losing progress. This setup makes it easier to experiment and refine models over time.
Datasets used for training and evaluation
A variety of datasets are used to train and evaluate search agents with MTRL. These include multi-hop factoid question-answering datasets that require combining information from multiple sources. Such datasets test the model’s ability to synthesize knowledge across different documents.
Other datasets focus on reasoning-intensive retrieval tasks across multiple domains. These datasets evaluate how well the model can find relevant information that requires understanding context and making inferences. They are designed to challenge the model’s ability to handle complex queries.
Internal company datasets are also used, such as synthetic enterprise documents and questions. These datasets help train retrieval systems for internal knowledge bases. They include labeled query-product pairs and multilingual product-search data, which test the model’s ability to handle different languages and relevance types.
Additional datasets include web-based question-answering, support QA over knowledge bases, and deep-research queries. These cover a broad range of real-world scenarios, from technical Q&A to product searches. Using diverse datasets helps ensure the model performs reliably across different tasks.
Reward function and metrics: nDCG@10 and failure penalties
The main metric used for training and evaluation is nDCG@10, which stands for Normalized Discounted Cumulative Gain at rank 10. It measures how well the top 10 retrieved documents match the ideal ranking. A score of 1.0 indicates perfect ranking, while 0.0 means no relevant documents are retrieved.
In training, nDCG@10 is used as the reward signal. The model gets a higher reward when it places relevant documents near the top. This encourages the system to improve its ranking quality over time. The reward is calculated after the full multi-turn search, reflecting the final outcome.
To discourage undesirable behaviors, a penalty of -1 is assigned if the agent reaches the maximum number of turns or tokens without success. This penalty teaches the model to avoid long or unproductive searches. It also helps the agent learn to finish tasks within its turn and token limits, improving overall reliability.
Training process: configuration, hyperparameters, and execution
Training with SageMaker AI MTRL involves setting a few key hyperparameters. The main ones are the number of epochs, batch size, and rollout concurrency. For example, training might run for a single epoch with a batch size of 128 prompts per step. Rollout concurrency controls how many search simulations happen simultaneously.
The training job is launched in the cloud, where multiple rollouts generate data in parallel. The system automatically handles gradient updates and trajectory collection. Users can monitor training progress through metrics like reward scores and failure rates. The training process can span multiple days, with support for checkpointing and resuming.
Adjusting hyperparameters allows developers to tune the training process. For instance, increasing the number of epochs or batch size can improve results but may require more compute resources. The default settings are designed to work well out of the box, making it easier for teams to get started without deep reinforcement learning expertise.
Results: improvements in retrieval quality and reliability
Fine-tuning with MTRL has shown significant improvements in search performance. The trained models perform better on several benchmark datasets, with higher nDCG@10 scores. The largest gains are seen in datasets that require complex reasoning or multi-hop retrieval.
For example, the nDCG@10 score on some datasets increased by over 20 percent after fine-tuning. This indicates that the model ranks relevant documents more accurately at the top of the list. The improvements mean users get better results faster and with fewer irrelevant documents.
Reliability also improved markedly. The failure rate, which measures how often the agent hits errors or exceeds limits, dropped sharply. In one case, the failure rate fell from nearly 23 percent to less than 1 percent. This shows the model learned to finish searches within the turn and token budgets, making it more dependable in real use.
Training curves display steady growth in reward scores during the process. Both training and validation rewards increase over time, then plateau. This pattern suggests that further training might not yield much more improvement, indicating the model has reached a good level of performance.
Practical steps: resource cleanup and deployment considerations
After training, it is important to clean up resources to avoid ongoing charges. Stopping or deleting the training job in the SageMaker console is the first step. Any model artifacts stored in Amazon S3 should be deleted if they are no longer needed. This prevents unnecessary storage costs.
If the trained model is deployed for inference, the endpoint should be deleted when it is no longer in use. This helps control costs and maintain security. It is also good practice to review logs and metrics to understand the model’s behavior and performance.
Deploying the fine-tuned model involves creating an endpoint that exposes search tools, such as BM25 and vector search. These tools are called during search rollouts to gather information. Proper deployment ensures the model can be used reliably in production environments.
Teams should also monitor the model’s performance over time. Adjustments to hyperparameters or retraining may be necessary as data or requirements change. Regular evaluation helps maintain high retrieval quality and system reliability.
Reflection released Beam, a 501 billion parameter text-only model for coding and science. Apache 2.0 weights are available this month after training on 23.8 trillion tokens.