Reka's Rho-1 Does Text, Video, and Robotics in One Unified Model
Research·October 6, 2026
Reka AI has released Rho-1, a 19-billion-parameter model that attempts something most AI labs have avoided. Rather than assembling separate components for different tasks, the entire system was trained from scratch to handle text, images, video, and robot control signals through a single shared architecture.
The unified approach represents a departure from the current industry standard, where companies typically combine specialized models to handle different modalities. A video generator, a text encoder, a robotics controller. Reka instead built one network with a shared key-value cache that maintains reasoning across all input and output types. This allows the model to process a video and generate corresponding robot actions or natural language descriptions without architectural fragmentation.
The practical advantage becomes clear in speed tests. A distilled version of Rho-1 generates a 5.3-second video in roughly one second, a significant performance gain that could matter for real-time robotics applications where latency compounds problems. The existence of a faster distilled variant suggests Reka found ways to compress the model without collapsing its capabilities entirely.
Rho-1 remains a research preview for now. Reka has not released model weights to the public, keeping the system under controlled access while the company presumably evaluates next steps. This pattern has become common in AI launches: announce the breakthrough, gather feedback, decide later whether to open-source or commercialize.
The move reflects a broader question the field is beginning to ask. As individual models become more powerful, do researchers gain more by building everything into one system or by maintaining modular specialized components? Rho-1 is one company's answer to that trade-off, though whether unified architectures ultimately win out over specialized alternatives remains unsettled.
Reporting based on an external source.