Reka Rho-1: One 19B Model for Text, Video and Robot Actions
Reka Rho-1 is a 19B research-preview model that reads and generates text, images and video and outputs robot actions in a single network. No weights or API yet.

- Reka says Rho-1 handles text, image, video and robotic actions as tokens in a single context window, with no tool calls and no second model.
- The base model generates video at 0.79x real time, and a distilled variant cuts denoising from 99 steps to 8, returning a 5.3-second clip in about a second in Reka's tests.
- It was trained on 320 H100 GPUs for about three months. Reka calls it a proof of concept, not a finished product, and lists real limitations.
in this block
Reka Rho-1 is a weird and interesting new AI model. Reka released it on October 5 as a research preview: a 19 billion parameter "omni-reasoning" model trained from scratch that reads and generates text, images and video, and even outputs robot actions, all inside one network.
What actually happened
Reka published the announcement on its official blog under the title "Collapsing the multimodal stack." The idea: most AI systems today are pipelines, with a central model handing jobs to specialist models for images, video or actions. Each handoff adds lag and loses context. Reka Rho-1 tries to do it all in one brain.
The demo shows a single five-turn session. The model draws a lighthouse, puts a box around it, animates a drone flying toward it, edits the clip into a snowstorm, then explains what changed. Reka says every step came from the same model and the same memory, with no external detector or video tool.
The Decoder confirmed the release the same day and summed it up the same way: one 19B network for text, images, video and robot control, generating continuous video in real time and taking new instructions on the fly without restarting.
How Reka Rho-1 works
Reka describes a "symmetric" architecture. Inputs and outputs use the same two formats: discrete tokens for text and commands, and continuous tokens for image latents, video frames, robot actions and proprioception. Because outputs look like inputs, anything the model makes can be fed straight back into its context.
Training mixes two objectives at once: next-token prediction for text and flow matching for images, video and actions. Reka's hypothesis, and it is stated as a hypothesis, is that gains in one skill bleed into the others because everything shares attention.
Speed is the flex. Reka says a watchable stream starts in roughly six seconds, and the distilled model made images about as fast as the quickest dedicated image models it tested. It also says it was the fastest model it timed to the first token of a text reply. These are internal tests, not third-party benchmarks.
One model that can imagine a scene, answer questions about it and move a robot arm in it. That is the pitch. The receipts are still early.
Robots and world models
The robotics angle is the long game. Reka says the same weights that predict future camera frames also output joint actions, so a policy can "imagine" the next few seconds before acting. Demos show a gripper lifting a banana, carrying a mug by the handle and pulling folded cloth out of a basket, and one episode runs on a LIBERO simulation task.
To get more training data, Reka pairs Rho-1 with an inverse dynamics model that infers control signals from ordinary internet video. That is how it hopes to scale past scarce robot logs. If you followed our MindOn Mind-1 robot speed piece, this is the same race from the model side.
Reka is honest about the gaps. Long rollouts drift structurally, grounding works on images but not yet across video, editing is brittle, and native video is capped at 672x384. It blames modest compute and data for most of it.
What to do as a reader
If you are a builder, this is a research preview, not an API you can ship on. There are no public weights or published benchmark scores in the announcement. Reka is asking robotics, simulation and vision-action teams to reach out at [email protected].
If you track AI agents, note the direction. Agent stacks that juggle five models, like the ones in our Holo4 agent release coverage, may face competition from single models that do the whole loop.
If you are a crypto degen, expect copycat tokens named after hot models. Reka has not announced any token. Reka Rho-1 is a lab result, not a coin. Anything on a launchpad using its name is a stranger's meme.
The real tell will be scale. Reka says Rho-1 was built on a tiny fraction of frontier compute. Watch whether a bigger version fixes the drift. Not investment advice, ser. Just a model worth bookmarking.
Not financial advice. DYOR, ser.