Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
AI agent teams waste tokens for tiny quality gains - OpenSmartRoute
Researchers tested if adding more AI agents improves results. They found that teams often cost much more than single agents. The extra cost can be between 1.8 times and 5.1 times higher. This huge price jump comes from using many more tokens. Tokens are the basic units of data that models consume.
The study looked at two specific models: GPT-6 Sol and Claude Opus 5.5. These models ran as solo agents and as part of a team. Teams worked together to solve coding problems. The team setup required more reasoning effort than the solo setup.
Only one comparison showed a real improvement in results. That was GPT-6 Sol working at a medium reasoning level. The team scored 7.3 points higher than the solo agent. This was the only statistically significant win for a team.
At maximum reasoning effort, teams failed to beat solo agents. Neither GPT-6 Sol nor Claude Opus 5.5 improved their scores. The extra agents added noise instead of helpful input. The results suggest teams are not worth the extra money.
Most models already run at full compute capacity. Adding more agents does not help them think faster. The extra cost of running a team is not justified. Companies should reconsider building large agent teams.
The Vals AI benchmark results on GPT-6 Sol and Claude Opus 5.5
Vals AI created a benchmark to measure agent performance. They called it the Vibe Code Bench. This test measures how well agents write code. The benchmark compares solo agents against teams of agents.
The visual results show arrows pointing from single agents to teams. These arrows represent the same reasoning effort level. The teams cost far more but barely improved scores. The graph clearly shows the high price of scaling up.
GPT-6 Sol showed the best performance in the entire study. It improved its score when working as a team. However, the improvement was small compared to the cost. The team used five times more tokens than the solo agent.
Claude Opus 5.5 showed no improvement at all. Its team setup performed worse than its solo version. The extra agents confused the model instead of helping it. This happened even at medium reasoning levels.
Vision-language models link visual data with text to describe, compare, and reason about images. They differ from standard computer vision by using open-ended language instead of fixed labels.
OpenAI plans to dump hundreds of AI-solved math problems on GitHub without publishing papers. Mathematicians want formal verification and proper credit before accepting the results.
The benchmark measured cost per application versus score. It showed a steep line for teams. The line stayed flat for solo agents. This proves that teams waste money on tokens.
How other companies like Anthropic and Fable tested scaling agents
Anthropic ran its own tests with the Opus 5.5 model. They added more agents to see if quality would rise. They found that quality gains actually shrank as they added agents. This happened in two of their specific tests.
Larger teams reached a performance level faster initially. But going from ten agents to one hundred only nudged scores up slightly. This happened after twenty-four hours of running. The marginal gain was not worth the massive token usage.
Fable 5.1 tested a different task involving theorem proving. This task requires strict mathematical logic. The model showed stronger quality gains above ten agents. It still scored below Opus 5.5 across all tests.
On a knowledge base task, Fable's score dipped slightly. This happened when scaling from thirty to one hundred agents. More data did not mean better answers here. The model struggled with the extra context.
These tests confirm that scaling agents has limits. Companies must be careful about how many agents they use. More agents do not always mean better results.
Why multi-agent systems mainly buy speed instead of quality
OpenAI researcher Noam Brown discussed agent scaling recently. He confirmed that multi-agent systems mainly buy speed, not better quality. Four agents solved tasks twice as fast. But they also cost twice as much in tokens.
At sixteen agents, the pattern held true. Efficiency grew slightly less as numbers increased. The system became less efficient with more participants. Speed comes from parallel processing, not better thinking.
The effect depends heavily on the specific task. Web research and math parallelize well. Many agents can check different sources at once. Writing a novel does not parallelize well. One agent must write the whole story.
Throwing ten thousand agents at a novel is pointless. It is just as useless as throwing ten thousand people. The work cannot be split up effectively. Some tasks require deep focus that teams disrupt.
OpenAI developer Eric Provencher warned against using agent swarms. He argued they are most likely wasted money. Coordination between agents breaks down over time. He called this the coordination tax.
The tax represents the cost of managing many agents. Agents argue with each other or repeat work. This waste adds up quickly in token costs. Teams pay for mistakes they could avoid alone.
Why it matters for teams running models and agents
Teams running models face high costs every day. They must decide if agent teams make sense. The research shows that teams cost significantly more. The price difference ranges from 1.8 times to 5.1 times.
Engineers often run models at full compute capacity. Adding agents does not help them think faster. The extra cost is not worth the tiny quality gains. Managers should question the ROI of large teams.
Safety becomes harder with more agents. Each agent introduces a potential failure point. Prompt injection attacks become more likely in large teams. The coordination tax can hide security risks.
Budgets suffer when teams consume massive tokens. Companies lose money on unnecessary token usage. The savings ledger shows these losses clearly. Teams drain budgets without delivering value.
Quality metrics remain flat for most teams. The Vals AI benchmark shows this clearly. Scores barely move while costs skyrocket. Engineers need better tools to manage these teams.
How OpenSmartRoute helps manage agent costs and coordination
OpenSmartRoute is an open-source router for AI requests. It sends each request to the best-fit model from a catalogue. The team defines the catalogue with specific models and agents.
The router scores every candidate on quality, cost, speed, and safety. The team sets the weights for each score. Hard rules pin a request to a specific model. Text with personal data stays on an on-premises model.
A region or a cost cap is never crossed by the router. It learns from outcomes automatically. A model that answers well gets more traffic. One that fails gets less traffic over time.
A new model is just one catalogue entry. It competes on the next request immediately. Nothing else changes in the app for the user. The hosted platform keeps a models catalogue with prices.
Public rankings are built from real traffic data. The savings ledger shows what each routed request cost. It compares this to what the most expensive model would have cost.
An input guard spots prompt injection before a request leaves. This prevents personal data from leaking out. The 'osr eval' tool measures routing accuracy on the team's own prompts. It can fail a build if accuracy drops.
Teams that use OpenSmartRoute avoid the coordination tax. They route requests to the cheapest effective agent. The router prevents wasted tokens on failing teams.
This setup saves money by default. The system optimizes for the team's specific goals. It ensures safety rules are followed automatically. The team gains control over their agent spending.
How it compares
Old systems tried to make teams work by adding more agents. They hoped more brains meant better answers. This research shows that is mostly false. Teams cost much more than single agents. They rarely give better results.
Single agents handle most coding and writing tasks well. They finish work faster when you need speed. Teams add complexity without adding much value. The extra cost is real and hard to ignore.
Some tasks still need many agents. Web research and math problems benefit from parallel work. Multiple agents can split a big problem into smaller pieces. They solve each piece at the same time. This speeds up the overall process.
Creative tasks do not work the same way. Writing a story needs one focused mind. Splitting that work breaks the flow. The result is often worse than a single agent.
The main change is the cost. Teams spend 1.8 times to 5.1 times more. They do not get 1.8 times better quality. The math does not add up.
What stays the same is the need for good models. A strong single agent beats a weak team. You must pick the right base model first. Then decide if a team is even useful.
The research confirms what engineers already suspected. Complexity adds risk, not just speed. Simple systems are often the best choice.
Questions this leaves open
The study tested specific models and tasks. It did not test every possible combination. We do not know how teams behave with 10,000 agents. The costs are too high to test easily.
Researchers say scaling to very large numbers is unexplored. They warn that costs are simply too high. No one has tried it on a massive scale yet.
We do not know if coordination breaks down in all cases. Some tasks might need many agents to succeed. The research only tested a few specific scenarios.
The "Vibe Code Bench" is a new test. It measures how well agents write code. We do not know how other benchmarks compare. Different tests might show different results.
The "coordination tax" is a new term. It describes the cost of broken communication. We need more data to understand its full impact.
Readers can check their own projects. Run tests with single agents and teams. Measure the token cost and quality score. Compare the results carefully.
Look for the 1.8 times to 5.1 times cost difference. If your team costs that much more, ask why. Check if the quality actually improved.
Ask your team about coordination issues. Do agents talk to each other well? Do they get stuck? These questions matter for future planning.
The research suggests simplicity wins. But we need more proof for every task type. Engineers should test their own setups. They should look for patterns in their data.
Readers should also check the speed gains. Teams often run faster for math tasks. But speed does not always mean better quality.
The study used specific reasoning levels. Medium and maximum effort levels. We do not know how teams perform at low effort.
Fable 5.1 showed some gains above ten agents. But it still scored below Opus 5.5. This suggests a ceiling on team performance.
Anthropic saw quality gains shrink as it added agents. Their tests with Opus 5.5 support this view. Larger teams reached performance levels faster. But going from ten to 100 agents only nudged scores up slightly.
OpenAI researcher Noam Brown confirmed these trends in a podcast. He explained that multi-agent systems mainly buy speed. They do not buy better quality.
Eric Provencher warned against using agent swarms. He called it wasted money. He argued that coordination breaks down.
These open questions mean we must be careful. Do not assume teams are better. Test your specific use case first. Measure the cost and quality before buying.
The research leaves room for future discovery. New models might change the rules. New tasks might need new approaches. Keep an eye on the field.
Readers should stay curious. Ask questions about the coordination tax. Check if their teams are actually helping.
The bottom line is that more is not always better. Sometimes less is more. Simple systems often outperform complex ones.
What to do when planning your next agent project
Start by testing single agents before building teams. Measure the token cost and quality score carefully. Compare the results against a team setup. Look for the 1.8 times to 5.1 times cost difference.
Check if your task parallelizes well. Web research and math might benefit from teams. Writing a novel or creative tasks do not. Adjust your agent count based on the task type.
Use a tool like OpenSmartRoute to manage costs. Set hard rules for data safety and budget caps. Monitor the savings ledger to track token waste.
Run evaluations on your own prompts regularly. The 'osr eval' tool can catch routing issues early. Fail a build if the routing accuracy drops too low.
Consider the coordination tax when planning large teams. More agents increase the risk of breakdowns. Keep teams small unless the task truly requires it.
Focus on quality metrics that matter to your users. Do not chase speed at the expense of accuracy. The Vals AI benchmark shows quality gains are small.
Review your current agent architecture quarterly. Remove agents that do not improve scores. Reallocate budget to better models or tools.
The research suggests that simplicity often wins. Single agents can outperform complex teams. Keep your system lean and efficient. This approach saves money and reduces risk.