Agents11 min read
AWS releases AgentCore Evaluations for multi-agent systems
Amazon Bedrock AgentCore adds a platform to evaluate agent performance and explainability in production.
From AWS machine learning blog
Blog
Daily notes on new models, LLM releases, agent frameworks and AI research, written from the sources we follow and delivered as a newsletter every day.
Get the morning and evening issues
Every new post of the day, in one email. Confirmation required.
Agents11 min read
Amazon Bedrock AgentCore adds a platform to evaluate agent performance and explainability in production.
From AWS machine learning blog
Agents1 min read
While end-to-end benchmarks like SWE-bench provide broad performance scores for AI agents, they are often expensive, slow, and lack the root-cause diagnostics needed to explain exactly where an agent's logic broke down. To solve this, de...
From Google AI developers blog
How this blog is made
POST /api/v1/route with execute: true.Photo: Yan Krukau
Agents1 min read
Skills let you encode domain-specific procedures as reusable, portable instructions for agents, but a fluent answer doesn't prove the agent picked the right skill or followed it. Learn how to measure skill selection and instruction follo...
From AWS machine learning blog
LLMs1 min read
arXiv:2609.13520v1 Announce Type: new Abstract: While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relat...
From arXiv cs.CL
Research1 min read
arXiv:2609.13637v1 Announce Type: new Abstract: Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a ...
From arXiv cs.AI
LLMs1 min read
Researchers introduced ASCIL, a post-ASR correction framework that learns from misclassifications to reduce false wake-up activations in conversational AI. Evaluations on a proprietary dataset showed a significant relative error reduction, alongside improvements in intentional acceptance rates.
From arXiv cs.CL
LLMs1 min read
The RIPPLE system improves workflow synthesis agents by isolating policy edits and evaluating their downstream effects through replay. Evaluations on Flow-HO show up to 23.1% improvement in validation success and maintain edit efficiency.
From arXiv cs.CL
Research1 min read
Researchers released Occamy-1.0, a 35 billion parameter model trained on a Qwen3.6-35B checkpoint. Evaluations across co-work benchmarks show it consistently performs among the strongest models of its size, achieving a competitive position on a cost-performance Pareto frontier.
From arXiv cs.AI
Research1 min read
Researchers developed WinSyn, an automated pipeline for generating realistic synthetic datasets of emails simulating enterprise scenarios. Evaluations using standard agentic baselines on these datasets showed aggregate scores below 80%, highlighting the need for more complex evaluation data.
From arXiv cs.AI
LLMs1 min read
Research reveals a significant disconnect between LLM-as-a-judge rankings and actual task success in user-simulated evaluations of task-oriented agents. The study, GAUGE, identifies a satisfaction-success gap and resolution loss among closely-ranked agents, highlighting the need for a calibrated approach to agent selection.
From arXiv cs.CL
Research1 min read
A new scheduling method reduces tail latency in agentic LLM workflows by strategically releasing turns based on evolving tail risk. Evaluations using real execution traces show significant improvements in workflow flow time under contention.
From arXiv cs.AI
Research1 min read
Benchmark Radar is a new database and search engine designed to streamline the discovery of AI benchmarks and evaluations for LLMs and other AI systems. It provides a centralized resource for locating benchmark datasets, code, and score histories, aiding in informed model selection and evaluation.
From arXiv cs.AI
Research1 min read
A new method, Specified-Foil Counterfactuals, allows engineers to identify past conditions that would have led to an alternative prediction in temporal graphs. The approach reduces predictor evaluations by 75-80% while maintaining 85.7-93.6% of black-box greedy successes.
From arXiv cs.AI
Agents1 min read
This post details a two-layered approach to monitoring production agent systems. It combines Amazon Bedrock AgentCore Evaluations for continuous quality scoring with AWS DevOps Agent for autonomous infrastructure investigation, using a four-agent airline reservation system as an example.
From AWS machine learning blog
Research1 min read
LexAgentHallu is a new benchmark designed to evaluate legal agent hallucinations across multi-step trajectories. The benchmark contains 3414 instances with detailed annotation and metrics, revealing patterns in agentic failures that are invisible with standard outcome-level evaluations.
From arXiv cs.AI
LLMs1 min read
A new benchmark, HybridDeepResearch, assesses autonomous agent performance requiring both web search and SQL queries to solve complex analytical tasks. Evaluations show current state-of-the-art models struggle with directional reasoning when integrating structured and unstructured data.
From arXiv cs.CL
LLMs1 min read
SEA-SpeechBench is a new large-scale benchmark evaluating speech understanding across 11 Southeast Asian languages. Initial evaluations of leading models reveal significant performance gaps, particularly in temporal understanding and low-resource language prompting.
From arXiv cs.CL
Research1 min read
RAPID achieves higher mean accuracy in text classification tasks by separating reliability-gated relational distillation from adaptive pair proposal. Evaluations using BERT-DistilBERT and DistilBERT-DistilBERT show significant improvements compared to a cross-entropy baseline.
From arXiv cs.AI
Research1 min read
arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We in...
From arXiv cs.AI
Posts are drafted from public feeds by models OpenSmartRoute routes to - the same router, skill and metering customers use - and always link to the original source. Corrections: support.
Archive (93)