Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
AWS releases AgentCore Evaluations for multi-agent systems - OpenSmartRoute
AWS releases AgentCore Evaluations for multi-agent systems
Amazon Bedrock AgentCore adds a platform to evaluate agent performance and explainability in production.
Key points
AWS launched AgentCore Evaluations on October 5, 2026.
The system supports both built-in and custom evaluators for agents.
Evaluations run in on-demand or online modes for testing and monitoring.
Explainability is now a first-class dimension alongside accuracy.
Why it matters: Teams can measure task success, instruction following, and decision transparency together.
By OpenSmartRoute editorial · written through the router by writer-small
From AWS machine learning blog - “Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore”
Architecture of the multi-agent supply chain decisioning solution on Amazon Bedrock AgentCore. Image: AWS machine learning blog (original)
Amazon Bedrock AgentCore Evaluations launch today. This new platform helps teams test multi-agent systems in production. Enterprises now use these agents to solve complex real-world problems. They coordinate multiple specialized agents to make decisions and execute workflows. Traditional evaluation methods often fail when checking agent performance. These older methods focus only on model response quality. They miss critical checks for tool selection and workflow execution. AgentCore Evaluations address this gap with a fully managed capability. It assesses agent performance across development and production environments. Teams can now measure accuracy, task success, and behavior across multiple dimensions. This shift ensures agents follow instructions reliably in real-world scenarios.
The challenge lies in ensuring consistent helpfulness and explainability. Multi-agent systems move from experimentation to production quickly. Enterprises need deeper guarantees than simple fluent responses. Agents must select the right tools and respect business constraints. They also need clear reasoning behind their outputs. Current evaluation approaches are insufficient for these agentic systems. Correctness depends on more than just language quality. It requires checking tool usage and adherence to constraints. Production deployment also needs responsible AI controls. Amazon Bedrock Guardrails provides safeguards like content filtering. These tools enforce safety constraints during execution. Evaluations assess agent quality after the work is done. Guardrails prevent harmful actions while the work happens. This combination covers both performance and safety requirements.
Amazon Bedrock AgentCore serves as a platform to build agents at scale. It connects and optimizes agents using any framework or model. The new evaluation capability is designed for production environments. It allows teams to measure quality across different dimensions. Built-in evaluators offer pre-defined assessments for common issues. They cover helpfulness, task success, and instruction following. Teams can quickly baseline performance without extra setup. However, enterprise use cases need deeper validation. Custom evaluators allow you to define business-aware checks. You can create rules specific to your industry needs. This ensures agents meet domain-specific business validity standards.
Amazon Bedrock in AWS GovCloud now supports Claude Opus 5.5 and Claude Sonnet 5.5. These models hold FedRAMP Class D certification and DoD Impact Level 4 or 5 authorization.
Explainability is now a first-class evaluation dimension in this framework. Built-in evaluators assess general response clarity effectively. Custom evaluators verify that agents articulate decision rationale clearly. They check if agents reference supporting data or tool outputs. The system also explains trade-offs like cost versus service level. Combining these evaluators provides structured insights into agent decisions. It goes beyond surface-level response quality checks. Teams gain measurable data on how and why agents decide. This transparency is crucial for building trust in automated systems.
To understand the architecture, consider a fictitious global retail company called AnyCompany Retail. This multinational retailer operates ecommerce channels and fulfillment centers. They face frequent inventory imbalances across different regions. Some areas face stockouts during promotions while others have excess inventory. Transportation teams must balance delivery speed with carrier capacity and cost. The company wants an agentic assistant to optimize inventory allocation. This assistant will recommend distribution adjustments and analyze inventory health. It can also simulate routing or fulfillment scenarios for planners.
The solution uses Strands Agents SDK along with Amazon Bedrock AgentCore components. You build a multi-agent supply chain decisioning system using these tools. The system includes an orchestrator agent and four specialized sub-agents. These include an optimization agent, distribution agent, routing agent, and analytics agent. Each agent runs on the Amazon Bedrock AgentCore runtime. Memory and Observability features are enabled for all agents. The orchestrator receives the planner's request and delegates work to specialized tools. It acts as the central coordinator for the entire workflow.
The optimization agent calls MCP tools backed by mock Amazon API Gateway REST interfaces. These interfaces return optimization decisions based on current data. The distribution agent calls recommendation APIs to suggest inventory rebalancing. It adjusts stock levels across fulfillment centers, stores, and digital channels. The routing agent calls logistics APIs to recommend carrier and route options. Finally, the analytics agent answers supply chain diagnostics questions directly. This solution uses foundation models from Amazon Bedrock for the agent loop. Model availability varies by AWS Region according to official documentation.
The solution supports both built-in and custom evaluators for comprehensive testing. Built-in evaluators assess general quality dimensions like helpfulness and task completion. Custom evaluators check supply-chain-specific behaviors such as constraint satisfaction. They verify route feasibility, SQL correctness, inventory grounding, and explanation quality. AnyCompany can evaluate both the language quality of the response and business validity. This dual approach ensures agents are accurate and logically sound.
The system supports on-demand and online modes for evaluation flexibility. On-demand mode is meant for development benchmarking and regression testing. It integrates with continuous integration and continuous delivery gates. Online mode handles continuous production monitoring and alerts. Both modes help close the loop and act on user feedback. Custom evaluators used for on-demand evaluations are repurposed for online mode. An OnlineEvaluationConfig object references the Amazon Resource Names (ARNs) of the evaluators. It specifies a sampling rate, such as 1–10% of production traces. Optional session filters allow further refinement of the data set. The service automatically reads traces from AgentCore Observability. It scores them and streams results to Amazon CloudWatch dashboards and alarms.
The architecture follows a three-layer evaluation approach for multi-agent systems. This approach progressively builds enterprise trust in automated decision-making. The first layer uses built-in evaluators requiring no setup. It applies Helpfulness as a universal baseline for all agents. A second agent-specific evaluator targets each agent's primary failure mode. For the orchestrator, this is Tool Selection Accuracy. Optimization and distribution agents face Response Relevance issues. Routing agents struggle with Instruction Following, while analytics agents suffer from Faithfulness problems.
The second layer adds custom evaluators encoding domain-specific business rules. These include constraint satisfaction for optimization tasks. Data grounding checks ensure distribution recommendations use real inventory data. Route feasibility validates that routing respects delivery windows and carrier capacity. SQL correctness ensures analytics queries match user intent accurately. Plan coherence evaluates if the orchestrator combines sub-agent outputs into valid recommendations. These evaluators validate business validity by checking budget limits and operational correctness.
The third layer applies explainability evaluators as distinct, cross-cutting checks. These assess whether agents articulate decision rationale clearly. They check if agents cite supporting evidence from tool outputs. The system verifies which constraints shaped the final response. It also explains trade-offs among cost, service level, and inventory risk. Agents must clarify why specific sub-agents were invoked during execution. They disclose assumptions when data is incomplete or missing. Separating explainability into its own layer allows independent measurement of transparency. A recommendation can be accurate but unexplainable, which signals a need for better reasoning articulation rather than decision logic.
Before deploying this solution, set up your development environment with specific tools. Install the AWS Command Line Interface (AWS CLI) on your local machine. You also need the AWS Serverless Application Model (AWS SAM) CLI version 1.100.0 or later. The Strands Agents implementation requires additional dependencies packaged in the DockerFile. These include strands-agents, strands-agents-tools, and bedrock-agentcore. The solution is available for download from the official GitHub repository. It provides a single step deployment to access the solution in your AWS environment. Edit terraform.tfvars to set the vpc_id and runtime_subnet_azs parameters. Copy the example configuration file to your local terraform.tfvars file. The outputs include runtime ARNs, AnyCompany Retail API URL, evaluator API URL, and memory ARN.
The test_client folder contains a Python script that invokes the deployed Supply Chain agent. It uses 20 sample queries with five per sub-agent to validate end-to-end functionality. Each category runs as a multi-turn session, generating unique session IDs. You can use these session IDs with the evaluators API for detailed analysis. Run the test_agent.py script with the runtime-arn and region parameters. This ensures your local tests mirror the production environment structure accurately.
Why it matters
This improves safety by enforcing constraints during execution while measuring quality after execution. It increases speed by automating regression testing and continuous monitoring without manual intervention. Operational trust grows because teams get measurable insights into agent reasoning processes. Safety controls prevent harmful actions while evaluation ensures decisions are accurate and explainable.
What to do
Start by installing the AWS CLI and SAM CLI on your development machine. Download the Strands Agents dependencies needed for your specific framework version. Clone the solution repository from GitHub to access the deployment scripts. Edit the terraform.tfvars file with your specific VPC ID and subnet details. Run the test_agent.py script to validate end-to-end functionality locally. Monitor the results using Amazon CloudWatch dashboards to track performance metrics.
Announcement - AWS introduces AgentCore Evaluations for multi-agent systems
Amazon Bedrock released new evaluation tools for complex agent teams. This update targets organizations running multiple specialized agents together. The goal is to measure accuracy, task success, and behavior across quality dimensions. Teams can now assess performance during development and production phases. Amazon Bedrock AgentCore serves as the central platform for building these systems. It supports any framework or model type for maximum flexibility. These evaluations replace older methods that only checked text fluency. They focus on whether agents follow instructions and respect business rules.
The Challenge - Why current evaluation methods fail for complex agents
Old evaluation techniques often miss critical failures in multi-agent workflows. A single fluent response might hide a wrong tool choice or skipped step. Enterprises need guarantees that agents select the right tools every time. They must also respect constraints like budget limits or safety policies. Current models struggle to prove they understand these hidden business rules. Traditional tests rarely check if an agent explains its reasoning clearly. Without deep validation, teams cannot trust automated decisions in real-world scenarios. Production deployment requires strict controls beyond simple quality checks.
Architecture Overview - How the supply chain example uses Strands Agents
The example uses a fictitious global retail company named AnyCompany Retail. This company faces frequent inventory imbalances across different regions and stores. Transportation teams must balance delivery speed, carrier capacity, and total cost. The goal is an agentic assistant that optimizes inventory allocation automatically. It recommends distribution adjustments and analyzes overall inventory health status. The system simulates routing scenarios to find the best fulfillment paths. We use Strands Agents SDK as the core framework for this implementation. An orchestrator agent manages the workflow and delegates tasks to specialized sub-agents. There are four distinct agents handling optimization, distribution, routing, and analytics. Each agent runs on Amazon Bedrock AgentCore runtime with memory enabled. Observability features track every action taken by the system during execution.
Evaluation Layers - The three-step approach to building trust
The solution builds trust through three specific layers of evaluation checks. First, built-in evaluators assess general quality dimensions like helpfulness and task completion. Second, custom evaluators verify domain-specific behaviors such as constraint satisfaction or SQL correctness. Third, an explainability layer ensures agents articulate their decision rationale clearly. This separation allows independent measurement of transparency versus logic accuracy. A recommendation can be accurate but still lack necessary explanation details. Teams need to measure both the output quality and the reasoning process separately. This layered approach prevents false positives from simple text generation metrics alone.
Built-in Evaluators - Pre-defined checks for general quality dimensions
Built-in evaluators provide pre-defined assessments for common quality dimensions automatically. They cover helpfulness, task success rates, and instruction following capabilities. Teams can quickly baseline agent performance without additional setup requirements. These checks run on every interaction to establish a standard of operation. They ensure the language quality of responses meets enterprise communication standards. General quality metrics help identify basic issues before diving into complex logic.
Custom Evaluators - Domain-specific rules for business validity
Custom evaluators address specific business-aware checks required for unique use cases. They verify constraint satisfaction, route feasibility, and inventory grounding accuracy. SQL correctness checks ensure data queries do not break database integrity rules. These rules validate the business validity of the agent's proposed decisions. AnyCompany Retail can evaluate both language quality and business logic validity simultaneously. Domain-specific validation prevents generic agents from making contextually inappropriate suggestions. Custom rules allow organizations to encode their unique operational constraints directly into the evaluation framework.
Explainability Layer - Cross-cutting checks for transparency and reasoning
The explainability layer performs cross-cutting checks for transparency and reasoning clarity. It verifies that agents explicitly articulate the decision rationale behind every action. Agents must reference supporting data or tool outputs in their explanations. They need to explain tradeoffs such as cost versus service level implications. This ensures users understand why a specific recommendation was made. Cross-cutting checks apply regardless of the specific task or domain involved. Transparency builds operational trust when humans review automated agent decisions.
Why it matters - How this improves safety, speed, and operational trust
This approach improves safety by enforcing constraints during execution while measuring quality after execution. It increases speed by automating regression testing and continuous monitoring without manual intervention. Operational trust grows because teams get measurable insights into agent reasoning processes. Safety controls prevent harmful actions while evaluation ensures decisions are accurate and explainable. Teams no longer rely on guesswork when agents make critical business recommendations.
What to do - Steps to set up the development environment and run tests
Start by installing the AWS Command Line Interface (AWS CLI) on your development machine. You also need the AWS Serverless Application Model (AWS SAM) CLI version 1.100.0 or later. The Strands Agents implementation requires additional dependencies packaged in a DockerFile. These include strands-agents, strands-agents-tools, and bedrock-agentcore packages. Download the solution repository from the official GitHub page for the latest code. Clone the repository to access the deployment scripts and configuration files. Edit the terraform.tfvars file with your specific VPC ID and subnet details. Run the test_agent.py script to validate end-to-end functionality locally. Monitor the results using Amazon CloudWatch dashboards to track performance metrics. Use session IDs generated during testing for detailed analysis with the evaluators API.