AI9 min read
Reflection's Beam beats Chinese models with less compute
Reflection released Beam, an open-weight model that matches GLM 5.2 and Qwen 3.8 on benchmarks while using three to four times less compute.
From The Decoder
Blog
Daily notes on new models, LLM releases, agent frameworks and AI research, written from the sources we follow and delivered as a newsletter every day.
Get the morning and evening issues
Every new post of the day, in one email. Confirmation required.
AI9 min read
Reflection released Beam, an open-weight model that matches GLM 5.2 and Qwen 3.8 on benchmarks while using three to four times less compute.
From The Decoder
AI1 min read
Alibaba's Qwen went from an invite-only chatbot in April 2023 to a 2.4-trillion-parameter open-weight model in August 2026. This is the full story, release by release: every major model, its key feature, and how its license changed. Each...
From MarkTechPost
How this blog is made
POST /api/v1/route with execute: true.Photo: Yan Krukau
A test showed Qwen3.8-27B computes sums and returns the answer in English words. It succeeded when reasoning was enabled but failed without it.
From Simon Willison on LLMs
AI1 min read
Chinese AI models often follow the party line on politically sensitive questions, according to an Aleph Alpha study that rated only 17 to 41 percent of answers as balanced. Aleph Alpha sells "sovereign AI" to governments, giving it a com...
From The Decoder
AI1 min read
With Clef and Clef-flash, Cloudflare is challenging TypeSafe AI's Jev decision model. Clef-flash delivers classifications in about 39 milliseconds, making it more than ten times faster than Jev. Both models are built on Qwen, licensed un...
From The Decoder
AI1 min read
AWS's Strands Agents team released Strands Decider 2B, an Apache-2.0 decision model built on Qwen3.5-2B-Base. It returns choices, yes/no probabilities and scores with calibrated confidence in one forward pass, never text. It runs at a 11...
From MarkTechPost
AI1 min read
Alibaba's Qwen team launched Qwen-Audio-3.1-Realtime, a full-duplex voice model for agents. It supports tool calling, has a large context window, and offers significant price reductions.
From MarkTechPost
Agents1 min read
This update shows how to deploy a text-to-speech model on SageMaker AI using vLLM-Omni. It streams speech over a WebSocket connection.
From AWS machine learning blog
Agents1 min read
Amazon SageMaker now supports deploying Qwen3-TTS for real-time voice cloning. Users can generate speech in a target speaker’s voice from a short reference.
From AWS machine learning blog
AI1 min read
BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B that spends 37.2% fewer thinking tokens across 12 benchmarks. Macro accuracy moves from 86.65% to 85.79%, and long-context AA-LCR improves by 2.25pp. The model...
From MarkTechPost
AI1 min read
Contrastive-LM has released CLM-8B, an open System One model that scores candidate actions against a state instead of generating text. It adds 2 small projection heads to a frozen Qwen3-8B encoder and trains them with a contrastive InfoN...
From MarkTechPost
LLMs1 min read
We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!
From Together AI blog
AI1 min read
Alibaba's Qwen team has released Qwen-Image-2.1, a 7B diffusion transformer that handles text-to-image generation, multi-reference editing, and native RGBA transparency in one checkpoint. A prefix KV cache speeds up edits with up to 10 r...
From MarkTechPost
AI1 min read
Qwen3.8-LiveTranslate is a new model that translates speech in real time with less delay. It supports 60 languages and adds features like speaker identification.
From MarkTechPost
AI1 min read
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B. The language model occupies 5.93 GB, against 53.80 GB in FP16. PrismML reports that it keeps 98.2% of the parent model’s average across 20 benchmarks. Th...
From MarkTechPost
AI1 min read
Alibaba's Qwen3.8-Omni-Flash understands audio and video, plans tasks, calls tools, and reports about 45.7% fewer tokens on OmniVideoBench. The post Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Age...
From MarkTechPost
AI1 min read
PrismML launched Bonsai 2, a 5.9 GB model based on Qwen3.8. It compresses the original by nine to ten times while keeping most performance.
From TechCrunch AI
Agents1 min read
AWS used a pipeline with Qwen-Image-Edit-2509 and Amazon Rekognition to create realistic training images. Experiments showed up to 160 percent improvement in person detection accuracy without dangerous photography.
From AWS machine learning blog
LLMs1 min read
TensorRT Edge-LLM ran Qwen3.6-27B on a Jetson AGX Thor. It finished the MLPerf Edge Agentic benchmark 6.4 times faster than the reference run.
From NVIDIA technical blog
Posts are drafted from public feeds by models OpenSmartRoute routes to - the same router, skill and metering customers use - and always link to the original source. Corrections: support.
Archive (93)