Agents1 min read
[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%
overshadowing more efficient GPT6 models from OpenAI
From Latent Space
Blog
Daily notes on new models, LLM releases, agent frameworks and AI research, written from the sources we follow and delivered as a newsletter every day.
Get the daily issue
Every new post of the day, in one email. Confirmation required.
Agents1 min read
overshadowing more efficient GPT6 models from OpenAI
From Latent Space
LLMs4 min read
Yesterday was Grok 4.7 (pelicans) and MiMo v2.6 Flash/Pro (more pelicans). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It's going to take a while to get a good read on all ...
From Simon Willison on LLMs
How this blog is made
Each feed entry becomes one request to OpenSmartRoute: the router picks a model with a cost-weighted objective, the editorial-writer skill is layered on the prompt, and the outcome trains the learners - the same pipeline available to every workspace.
Open any post to see which target answered, its confidence, the alternatives and what the request cost. Run the same pipeline yourself: register feeds in the operator console, map a small model under Providers, or call POST /api/v1/route with execute: true.
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the level of Claude Fable 5.1 on most work. It also costs 40% less to run than Opus 5 on typical workloads at default se...
From MarkTechPost
Agents1 min read
Claude Opus 5.5, Anthropic's most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS. This post covers what's new in Opus 5.5, practical guidance, ...
From AWS machine learning blog
AI1 min read
Anthropic called it "the strongest-performing model we've tested to date."
From TechCrunch AI
AI1 min read
Hacktron AI used a new version of Claude to breach OpenAI's security.
From TechCrunch AI
Research1 min read
A study comparing Claude-agent-sdk and deepagents across Claude-opus-4-8 and openai-codex with gpt-5.5, gemini-3.5-flash and deepseek-v3.2 found no significant advantage for either vendor-native harness. The study also examined harness cost and completion rates, revealing higher costs and a complex, unresolved billing structure.
From arXiv cs.AI
LLMs1 min read
The ORQA framework provides a method for testing large language model knowledge across 116 occupations using data sourced from trusted occupation-specific websites. Testing of 15 models revealed performance variations, with Claude Opus and GPT-5.4 achieving approximately 58-62% accuracy, while open-weight models showed 33-41% accuracy.
From arXiv cs.CL
Agents2 min read
GitHub introduces HydraFusion, a research preview that dynamically orchestrates multiple models for coding tasks, optimizing for quality, cost, and latency. This runtime orchestration system leverages a selection of execution patterns to deliver frontier-level intelligence, reducing costs by up to 67% compared to models like Claude Opus 5.
From GitHub blog: AI & ML
Posts are drafted from public feeds by models OpenSmartRoute routes to - the same router, skill and metering customers use - and always link to the original source. Corrections: support.
Archive (39)