Agents1 min read
Google Adds Native Live Voice Evaluation to ADK
Google AI developers now support evaluating live voice agents directly inside the ADK framework.
From Google AI developers blog
Blog
Daily notes on new models, LLM releases, agent frameworks and AI research, written from the sources we follow and delivered as a newsletter every day.
Get the morning and evening issues
Every new post of the day, in one email. Confirmation required.
Agents1 min read
Google AI developers now support evaluating live voice agents directly inside the ADK framework.
From Google AI developers blog
AI8 min read
Jas Khaira from Blackstone discussed how to build AI companies that last, focusing on funding, growth, and evaluating long-term potential.
From TechCrunch AI
How this blog is made
POST /api/v1/route with execute: true.Photo: Yan Krukau
AI transforms how developers work, shifting focus from implementation to defining problems, evaluating output, and making technical decisions.
From GitHub blog: AI & ML
Research1 min read
Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and...
From Apple machine learning research
AI1 min read
Thousands of applications. Multiple rounds of review. Hundreds of hours spent evaluating startups from around the world. Now comes the part everyone has been waiting for. Startup Battlefield 200 is almost here, and with today’s announcem...
From TechCrunch AI
LLMs1 min read
An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software...
From NVIDIA technical blog
LLMs1 min read
arXiv:2609.13389v1 Announce Type: new Abstract: Mapping textual specifications into formal representations is essential for ensuring the correctness of protocol designs and implementations. LLM-generated mappings, used for networking sec...
From arXiv cs.CL
Research1 min read
arXiv:2609.13422v1 Announce Type: new Abstract: LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patent...
From arXiv cs.AI
LLMs1 min read
Research reveals a significant disconnect between LLM-as-a-judge rankings and actual task success in user-simulated evaluations of task-oriented agents. The study, GAUGE, identifies a satisfaction-success gap and resolution loss among closely-ranked agents, highlighting the need for a calibrated approach to agent selection.
From arXiv cs.CL
LLMs1 min read
The RIPPLE system improves workflow synthesis agents by isolating policy edits and evaluating their downstream effects through replay. Evaluations on Flow-HO show up to 23.1% improvement in validation success and maintain edit efficiency.
From arXiv cs.CL
LLMs1 min read
The ORQA framework provides a method for testing large language model knowledge across 116 occupations using data sourced from trusted occupation-specific websites. Testing of 15 models revealed performance variations, with Claude Opus and GPT-5.4 achieving approximately 58-62% accuracy, while open-weight models showed 33-41% accuracy.
From arXiv cs.CL
Research1 min read
A new survey defines AI agent capabilities across five dimensions – environmental interaction, learning, autonomy, goals, and temporal coherence. The resulting Agent Compendium provides a structured resource for evaluating and comparing AI agents, promoting reproducibility and clearer research.
From arXiv cs.AI
Research1 min read
Research assessed generative AI agent performance across repeated interactions and accumulating disruptions in simulated healthcare workflows. The study examined operational resilience and considerate participation, revealing distinct behavioral patterns as challenges increased and informing deployment dilemmas.
From arXiv cs.AI
Agents1 min read
This post introduces a benchmarking harness to assess OpenAI models on Amazon Bedrock, focusing on cost per correct answer and deliverable quality rather than just price per token. The harness provides a more accurate measure of production workload costs.
From AWS machine learning blog
LLMs1 min read
The CyberGym benchmark, designed for evaluating AI security, is now publicly available on GitHub. This allows for high-score competition and potential model weight sharing on Hugging Face.
From Simon Willison
Research1 min read
A new benchmark, RESCUE, has been created to assess LLMs' ability to understand and utilize evolving interpersonal relationships in multi-party emotional support conversations. Experiments with ten models reveal limitations in capturing relation dynamics, particularly in tasks requiring relation pattern prediction.
From arXiv cs.AI
LLMs1 min read
The SWORD benchmark identifies inconsistencies in LLM factual error rejection by evaluating syntactically correct, but factually incorrect, statements across multiple languages. Results show models prioritize distributional familiarity over genuine verification and exhibit significant performance drops in East Asian languages.
From arXiv cs.CL
Research1 min read
XAI-Arena, a new framework, uses LLMs to evaluate the quality of XAI explanations, offering a scalable and reproducible method for comparison. Human ratings correlate strongly with LLM assessments, providing a systematic approach to evaluating XAI methods.
From arXiv cs.AI
Research1 min read
OpenDiscoveryTrace provides a public dataset of 558 AI scientist trajectories, capturing the reasoning process of 124 scientific tasks across drug discovery and other domains. Analysis reveals significant behavioral differences between models, particularly highlighting error profiles and confidence levels, enabling more robust evaluation.
From arXiv cs.AI
Posts are drafted from public feeds by models OpenSmartRoute routes to - the same router, skill and metering customers use - and always link to the original source. Corrections: support.
Archive (93)