Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Reka Unveils Rho-1: A Single Model Handling Text, Video, and Robot Actions - OpenSmartRoute
Reka Unveils Rho-1: A Single Model Handling Text, Video, and Robot Actions
Reka released a research preview of Rho-1, a 19B model that generates video and outputs robot actions. It replaces pipelines by handling all tasks within one context window.
Key points
Rho-1 is a 19B omni-reasoning model trained from scratch.
It generates video at 0.79x real-time speed in 5 turns.
A distilled variant produces clips in about one second.
The model outputs seven action channels for robotics simulations.
Why it matters: Reka's Rho-1 reduces latency and complexity by removing handoffs between specialized models for text, vision, and robotics.
By OpenSmartRoute editorial · written through the router by writer-small
From MarkTechPost - “Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One”
Reka released a research preview of Rho-1 on October 5, 2026. This is a 19-billion parameter model trained from scratch. It understands text, images, video, and robot actions in one system. No other models need to handle these different tasks together.
The team made this available only for researchers to test. There are no public weights or APIs for commercial use yet. Pricing remains unknown because the project is still in preview mode. Engineers must wait for official announcements before building production systems.
Announcement - Reka releases Rho-1 as a research preview only
Reka announced Rho-1 as a new omni-reasoning model on October 5, 2026. The publication appeared on MarkTechPost and included technical details from the team. Researchers can access the project through official channels like X or Telegram. Public weights are not available for download at this time.
This release marks a shift toward unified reasoning instead of modular pipelines. Most companies rely on separate tools for vision, language, and robotics. Rho-1 aims to replace those complex chains with a single neural network. The team emphasized that this is a research preview, not a product launch.
Architecture - Two expert streams share one context window
Rho-1 uses a unique architecture where two expert streams operate inside every transformer block. One stream focuses on understanding language and visual data through discrete tokens. The other stream handles generation by denoising continuous latents into images or video. Both streams share the same attention mechanism and key-value cache.
This design allows the model to switch between modes without losing context. When a task requires pixels, the understanding stream sends a handoff token. The generation stream then renders the final output using the full accumulated state. Training combines next-token prediction for discrete sequences with flow matching for continuous outputs.
Multimodal Capabilities - Understanding and generating text, images, video, actions
Mirror Particle raises capital to create an AI engine that simulates changing human motivations. The company rejects large language models in favor of a foundation model trained on longitudinal data.
The model processes inputs as either discrete tokens or continuous tokens. Discrete tokens carry text, symbolic reasoning, and high-level commands. Continuous tokens carry image latents, video frames, and robot action signals. Proprioception data also flows through this unified token system.
Rho-1 can draw a lighthouse, box it, animate it, and edit the scene into a snowstorm. It explains the differences between these steps all within one session. No tool calls or second models are needed to complete this full loop. The system handles text, vision, and robotic actions as native capabilities.
Performance Metrics - Video generation speed and token efficiency
Video generation runs at 0.79 times real-time speed based on median measurements. A watchable stream starts in roughly six seconds for a standard clip. Reka measured the time to generate a first clip at seven seconds. This compares against an illustrative thirteen point eight seconds for multi-agent pipelines.
Token efficiency remains high because the model avoids redundant encoding steps. Bounding boxes appear as coordinate tokens instead of separate detector outputs. The first frame of a video reuses the in-context image representation. It does not create a new encoded copy for every new scene.
Distillation Results - Cutting denoising steps for faster inference
A distilled variant reduces denoising steps from ninety-nine to just eight. This change happens with minimal quality loss according to internal reports. The model now returns a five point three second clip in about one second. It matched the fastest dedicated image models in vendor-run tests.
The distilled version was also the quickest model tested to produce the first text token. These results come from Reka's internal testing rather than independent benchmarks. Engineers should verify these numbers before deploying them in critical systems. The speed gain comes directly from cutting unnecessary computational steps.
Why it matters - Replacing agentic pipelines with a single model
Replacing agentic pipelines reduces latency by removing multiple handoffs between models. Each specialist in a pipeline sees only a narrow part of the request. Rho-1 keeps the full context inside one window for better reasoning. This approach could lower costs and improve reliability for complex tasks.
Safety becomes easier to manage when one model handles all outputs. Teams no longer need to coordinate between vision, language, and action systems. The unified design simplifies debugging and monitoring for production environments. Engineers can focus on optimizing the single model instead of many tools.
How it compares - What existed before, what this changes and what stays the same
Old systems used pipelines to handle different data types. A central model planned the task first. It then handed off work to specialists for images or video. Each handoff added delay to the process. Specialists saw only a narrow part of the request. They did not see the full context.
Rho-1 removes those handoffs entirely. Text, vision, and robotic actions become tokens inside one window. The model draws a lighthouse, boxes it, animates it, edits the video into a snowstorm, and explains the difference. That entire loop happens in five turns. It requires no tool calls. No second model handles the work.
The base model generates video at 0.79 times real-time speed. This means the median generation time is roughly four seconds per second of video. A watchable stream starts in about six seconds total. Reka team measured seven seconds to get a first clip. An illustrative multi-agent pipeline took thirteen point eight seconds for the same task.
Every input and output uses one of two native formats. Discrete tokens carry text, symbolic reasoning, and high-level commands. Continuous tokens carry image latents, video frames, robot actions, and proprioception. Each transformer block holds two expert weight streams. The understanding stream handles language and visual parsing. The generation stream denoises latents into images and video. Both streams share attention and operate over the same KV cache.
When a reply needs pixels, the understanding stream emits a discrete handoff token. The generation stream then renders from the full accumulated state. Training combines next-token prediction for discrete sequences with flow matching for continuous outputs. This design has practical effects on bounding boxes. They come out as coordinate tokens instead of separate detector outputs. A video's first frame reuses the in-context image representation. It does not create a new encoded copy for every new scene.
Questions this leaves open - What the source does not say and how a reader can check it
The source states data verified on October 5, 2026 from official sources. This date is far in the future relative to current real-world time. Readers must treat all figures as hypothetical or part of a simulation environment. The text explicitly mentions "Research preview only." No public weights are available yet. There is no public API for this model. Pricing information remains completely undisclosed.
Engineers cannot run Rho-1 locally without access to the training infrastructure. Reka trained on 320 H100s for three months. Video output is capped at a resolution of 672 by 384 pixels. This resolution limit might affect use cases requiring high-definition video. The text does not specify the exact hardware required for inference. It also does not mention power consumption or energy costs per token generated.
Independent verification of the speed claims is currently impossible. Results come from vendor-run tests rather than independent benchmarks. Engineers should verify these numbers before deploying them in critical systems. The source mentions a distilled variant cuts denoising from 99 steps to 8. It reports minimal quality loss but does not quantify that loss. Readers need to check this metric themselves for their specific applications.
Safety considerations remain unaddressed in the provided text. One model handles all outputs including robot actions. Teams no longer need to coordinate between vision, language, and action systems. This simplification might introduce new safety risks if one model fails. The source does not mention how the model handles adversarial inputs or malicious prompts. It also does not discuss regulatory compliance for autonomous agent systems.
Feature comparisons against competitors like BAGEL, Emu3.5, and Genie 3 are incomplete. Rho-1 takes navigation actions while others do not emit them. Access to Genie 3 requires a Project Genie subscription. Emu3.5 parameter count is only listed on its Hugging Face model card. The source lacks direct performance comparisons for video generation quality. Readers must wait for official benchmark releases to make fair comparisons.
The text credits the researcher of this project but does not name them publicly. Michal Sutter is mentioned as a data science professional with a Master of Science from the University of Padova. This credit line appears in the source footer rather than the main technical content. Readers should verify the authorship through official Reka channels or academic publications. The text encourages following the researcher on X but does not provide a direct link.
Data privacy and training data provenance are not discussed in the article. The model understands and generates text, images, and video from scratch. It is trained on proprietary datasets that are not publicly disclosed. Engineers cannot audit the training process for bias or copyright issues. This lack of transparency might be a barrier for some organizations adopting the technology.
The source mentions pairing Rho-1 with its Inverse Dynamics Model to scale past scarce teleoperation logs. This suggests a modular architecture where external models assist the core system. The text does not explain how these models communicate or share state. Readers need to understand the integration points before building a full robotics stack.
What to do - Monitoring research previews and technical details
Readers should monitor official channels for updates on weights and APIs. Check the Technical details page for more information on the architecture. Follow the researcher on X or join the ML SubReddit for community discussions. Subscribe to newsletters if you want regular updates on model releases.
Compare Rho-1 against other omni-modal models like BAGEL or Emu3.5 when available. Look for independent benchmarks that verify video generation speed and quality. Test the distilled variant if your team needs faster inference times. Always check licensing terms before considering any commercial use cases.