Skip to main content

Reka’s Rho-1 combines video, images, text and robot actions

6 OCTOBER 2026·2 MIN READ·4 SOURCES

Reka describes Rho-1 as a 19-billion-parameter model trained from scratch that can understand and generate text, images and video, reason over them, and take actions within one neural network. Its launch materials also describe a robotics simulation demonstration, but not control of a physical robot.

Reka’s Rho-1 combines video, images, text and robot actions

Key takeaways · 3

  • 01

    Reka’s demonstration joined scene generation, object locating, video editing and explanation in five turns, using one model and conversation state without a tool call or second model.

  • 02

    The robotics evidence is from a LIBERO simulation episode with seven action channels; the announcement did not demonstrate control of a physical robot.

  • 03

    Before assessing video use, account for the 672-by-384-pixel output limit and Reka’s reported issues with long rollouts and targeted editing.

One model, shared context

Reka describes Rho-1 as a 19-billion-parameter model trained from scratch.[1] The company says it can understand and generate text, images and video, reason over them, and take actions within one neural network.[1] Text, vision and robotic actions are represented as tokens in a shared context window, according to Reka.[1] MarkTechPost reports that the streams share attention and use the same key-value cache.[2] Reka says new instructions can arrive at any moment.[1]

A multimodal demo, plus simulated actions

In one demonstration, Rho-1 generated a scene, located an object, animated and edited the video, then explained the changes in five turns.[1] Reka says this used one model and one conversation state, without a tool call or second model.[1] The robotics demonstration was narrower: MarkTechPost reports that Rho-1 produced seven action channels in a LIBERO simulation episode.[2] The Decoder reports Reka’s claim that the same weights predict camera images and drive robot movements.[3] RuntimeWire noted that the announcement did not demonstrate Rho-1 controlling a physical robot.[4] The simulation result therefore should not be read as a physical-robot deployment demonstration.[2][4]

Speed and output trade-offs

Reka reports that the base model generates video at a median 0.79 times real time and begins a watchable stream in roughly six seconds.[1] The company also says it returned a 5.3-second video clip in about a second.[1] A distilled version reduces video denoising from 99 steps to eight, according to Reka.[1] Those figures describe different reported measures; they do not establish that every task or output runs at the same speed.[1] Rho-1’s native video output is currently limited to 672 by 384 pixels.[1] Teams assessing it for video workflows should weigh these reported speeds against resolution needs and task-specific results.[1]

Preview limits and availability

Reka says Rho-1 can drift structurally during long video rollouts, even while retaining visual detail.[1] The company says object detection and coordinate grounding work reliably on static images but not yet across video.[1] It also describes targeted visual editing as nascent and consistency across diverse prompts as brittle.[1] The research preview has no public weights, API or pricing yet, according to Reka’s launch materials.[2] Reka reports that the model was trained on 320 H100 GPUs over about three months.[3] For now, the announcement offers evidence of a unified research direction, not enough information to plan a public API integration.[3][2]

Rho-1 points toward systems that combine perception, generation and action in a shared model, which may be relevant to teams evaluating multimodal or robotics research. The reported simulation and video results come with clear limits, and the preview currently offers no public weights, API or pricing. Professionals should separate promising demonstrations from capabilities they can test or deploy today.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 6 October 2026

    Reka’s Rho-1 combines video, images, text and robot actions

Sources

AI fluency, one session a day, built for your work.