Reka AI released a research preview of its new model Rho-1. This system handles text, images, video, and robot control in one network. It does not route tasks to separate specialized models. Instead, it runs all modalities as tokens in a single context window. The company calls this architecture an omni-model. Users can generate continuous video in real time with it. The model responds to new instructions without restarting the process.
Most AI systems require different tools for different tasks. They often call external functions or use separate models for specific jobs. Rho-1 breaks this pattern by keeping everything inside one neural network. Engineers can now test text, vision, and robotics together in one run. This simplifies the workflow for building complex applications. The preview allows researchers to see how the system behaves under load.
How Rho-1 processes text, images, video, and robot actions together
The same weights predict camera images and drive robot movements simultaneously. This means the model understands visual data just like it understands written words. It generates control signals for robots based on what it sees in a video. There are no tool calls or external models involved in this process. Everything happens within the shared context window of the single network.
Real-time generation is possible because the system processes data as it arrives. New instructions appear and the model adapts instantly to them. It does not need to reload data or switch between different software components. This creates a seamless experience for users interacting with physical robots. The integration of vision and action happens at the same computational level.
The technical trick using inverse dynamics to train on internet videos
Training robot models usually requires expensive, scarce datasets of human demonstrations. Reka AI faced this problem when trying to teach its model to control robots. To solve it, they built an inverse dynamics model. This tool pulls control signals from ordinary internet videos. It learns what actions humans perform in public footage. The system maps visual inputs to motor outputs automatically.