AI1 min read
Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear
Salesforce Koa is built on Nvidia's open-weight Nemotron model and is trained to do sales, marketing, and customer-support tasks.
From TechCrunch AI
Blog
Daily notes on new models, LLM releases, agent frameworks and AI research, written from the sources we follow and delivered as a newsletter every day.
Get the daily issue
Every new post of the day, in one email. Confirmation required.
AI1 min read
Salesforce Koa is built on Nvidia's open-weight Nemotron model and is trained to do sales, marketing, and customer-support tasks.
From TechCrunch AI
AI1 min read
Jensen Huang engaged in a conversation with President Trump at the All-In conference, discussing AI safety while demonstrating a new, unreleased foldable phone from Apple. This event highlighted Nvidia’s reliance on AI labs and Apple’s renewed interest in its hardware.
From TechCrunch AI
How this blog is made
Each feed entry becomes one request to OpenSmartRoute: the router picks a model with a cost-weighted objective, the editorial-writer skill is layered on the prompt, and the outcome trains the learners - the same pipeline available to every workspace.
Open any post to see which target answered, its confidence, the alternatives and what the request cost. Run the same pipeline yourself: register feeds in the operator console, map a small model under Providers, or call POST /api/v1/route with execute: true.
AI1 min read
OpenAI has purchased smartphone camera maker Glass Imaging for $300 million, leveraging expertise from former Apple engineers to improve smartphone camera image quality using AI. This acquisition aligns with OpenAI’s rumored development of its own hardware devices.
From TechCrunch AI
AI1 min read
iOS 27’s updated Siri leverages Google’s Gemini models for enhanced contextual understanding and complex task execution. This update offers new capabilities like multi-step directions and integration with the Photos app, improving Siri’s utility for daily tasks.
From TechCrunch AI
Agents1 min read
Richard Socher’s Recursive is building an AI system designed to accelerate AI research itself, achieving human-level performance on optimization tasks in under two days. The company’s $4.65B seed round focuses on recursive self-improvement and tackling complex scientific problems.
From Latent Space
AI1 min read
Founders must consider how their AI products maintain value amidst rapid advancements from platform models like OpenAI and Anthropic. The session explores strategies for building defensible AI companies beyond simply competing with model capabilities.
From TechCrunch AI
AI1 min read
A Real World AI Stage fireside chat featuring Ben Lamm, CEO of Colossal Biosciences, explores the intersection of AI, synthetic biology, and the potential for engineering nature’s comeback. The discussion examines the technologies behind de-extinction and the broader implications for biodiversity and conservation.
From TechCrunch AI
LLMs1 min read
The ORQA framework provides a method for testing large language model knowledge across 116 occupations using data sourced from trusted occupation-specific websites. Testing of 15 models revealed performance variations, with Claude Opus and GPT-5.4 achieving approximately 58-62% accuracy, while open-weight models showed 33-41% accuracy.
From arXiv cs.CL
LLMs2 min read
SynthSentry is a model-agnostic tool for detecting synthetic data contamination in language model training corpora. It analyzes lexical diversity, n-gram tails, and perplexity variance across reference models, offering a pre-training screening method without requiring access to the generating model.
From arXiv cs.CL
LLMs1 min read
The R2VC architecture, combining retrieval, verification, and confidence calibration, achieves higher accuracy in open-domain fact-checking compared to baseline models. Ablation studies highlight the critical roles of candidate selection and confidence calibration in improving both predictive accuracy and reliability.
From arXiv cs.CL
Research1 min read
The BlueLM-GUI system, a 35B model, achieves strong performance on mobile GUI tasks through a real-device training approach. This flywheel system utilizes a multi-stage process of data salvage, real-device rollout, and evolving benchmarks to improve agent capabilities and transfer directly to production environments.
From arXiv cs.AI
Research1 min read
T-GADE, a system using thermodynamic genetic algorithms and LLMs, evolved structured artifacts like code paired with descriptions. Training experiments on the online bin-packing task showed a 29% reduction in median training excess at a temperature of 0.003, validating the approach’s effectiveness.
From arXiv cs.AI
LLMs1 min read
The research introduces Chopthin-Consensus Power Sampling (CCPS), a method for LLM decoding that preserves reasoning diversity and improves accuracy. Evaluation across open-weight models and benchmarks shows CCPS matches or exceeds the Power-SMC baseline in 14 of 15 settings, achieving gains of up to 10.6 percentage points.
From arXiv cs.CL
Research1 min read
Researchers introduced GT Bench, a benchmark with over 100,000 examples across four graph representations, to evaluate LLM performance on algorithmic graph problems. The Graph Theory Agent (GTA) system, combining a representation selector and a plan-and-decompose agent, achieved improved results on the GT Bench benchmark.
From arXiv cs.AI
Research1 min read
Researchers released Occamy-1.0, a 35 billion parameter model trained on a Qwen3.6-35B checkpoint. Evaluations across co-work benchmarks show it consistently performs among the strongest models of its size, achieving a competitive position on a cost-performance Pareto frontier.
From arXiv cs.AI
AI1 min read
Recent statements from AI researchers, including a >10% risk of human eradication from Anthropic’s alignment lead, have sparked debate about the potential dangers of advanced AI models. This follows incidents involving OpenAI’s internal model and increased model capabilities.
From TechCrunch AI
AI2 min read
This article details four mechanisms – context budgeting, memory strategy, todo-state, and compaction – used within harness systems like LangChain Deep Agents and Claude Code to mitigate context overflow and goal loss in long-horizon agent tasks. These techniques involve offloading, summarization, and structured data management to improve agent performance.
From MarkTechPost
LLMs1 min read
ChatGPT Work, using GPT-6 Astra, created 5K and 10K running routes based on a user’s address, leveraging OpenStreetMap data. The process highlights challenges with LLM transparency and the need for agent tool call preservation.
From Simon Willison on LLMs
AI1 min read
Cognition’s SWE-2, a post-trained model based on Kimi K3, achieves 50% accuracy on FrontierCode 1.1 Main, costing 64% less than Fable 5.1. It’s currently available only within the Devin platform.
From MarkTechPost
Posts are drafted from public feeds by models OpenSmartRoute routes to - the same router, skill and metering customers use - and always link to the original source. Corrections: support.
Archive (34)