block #0008 in --:--:--Join the pool
PENDING…
ai ✓ confirmed 6/6 1h ago · 3 min read

Odyssey-3 world model Claims Physics Crown With 66.1 Score

Odyssey-3 world model launched with a claimed state-of-the-art 66.1 on Physics-IQ Verified, using best-of-8 sampling. The live leaderboard ranks it second by default.

Odyssey-3 world model Claims Physics Crown With 66.1 Score
tl;dr
  • Odyssey-3 is an autoregressive diffusion transformer that predicts how scenes evolve as an agent acts. A research preview is live, and developers have to contact the company for access.
  • Odyssey says the Pro model scored 66.1 on Physics-IQ Verified with best-of-8 sampling. The live leaderboard's default view showed FLUX 3 [large] first at 64.36 and Odyssey-3 Pro at 61.78 when we checked.
  • Partners include Poke & Wiggle for robot benchmarking and Flexion, which built humanoid control policies on top of the model.
in this block
  1. What actually happened
  2. The benchmark fight
  3. The cost footnote
  4. How it was trained
  5. Robots, cars and agents
  6. What to do as a reader

The Odyssey-3 world model dropped on October 8 with a bold claim: state-of-the-art physics. Odyssey, a startup founded by people who spent a decade on self-driving cars, says Odyssey-3 Pro hit 66.1 on the Physics-IQ Verified video-to-video benchmark, the highest score reported. The fine print, though, includes best-of-8 sampling and a live leaderboard that ranks things differently.

What actually happened

Odyssey's launch post calls Odyssey-3 its most powerful foundation world model yet. It's a learned dynamical system that picks up physics, dynamics and cause-and-effect from visual data. Developers can use it both to simulate environments and to train policies for robots and other machines.

The preview lets you generate an embodied environment from a prompt, walk through it in first or third person, move the camera on its own and drop in events mid-generation. The model then predicts in real time how the world reacts.

On resolution, footnotes say Odyssey-3 renders at 832×480 and Pro at 1280×720. Odyssey also says the model ranks first in 3 of 4 WorldMark categories in its own evaluations.

The benchmark fight

Physics-IQ, from Anates Labs and DeepMind, asks models to continue videos of real physical experiments, covering fluids, optics, solid mechanics, magnetism and thermodynamics. Predictions are then compared with what actually happened, so pretty pixels alone don't score points.

Odyssey's chart, sourced to the leaderboard on October 7, shows the base Odyssey-3 at 51.8 with plain prompts, 61.6 with enhanced prompts and 64.4 with best-of-8. Pro scored 63.4 with enhanced prompts and 66.1 with best-of-8. In image-to-video, Pro scored 54.7.

When we loaded the live Physics-IQ Verified leaderboard on October 9, its default ranking listed FLUX 3 [large] at 64.36 and Odyssey-3 Pro at 61.78. Different views, prompt setups and sampling configurations explain the gap. We're reporting both numbers rather than picking one.

Best-of-8 is a real result, but it isn't your first try.

The cost footnote

Odyssey says Odyssey-3 improves the tradeoff between physical accuracy and generation cost. Its footnote, though, says the cost "assumes $1 per MI355X GPU-hour, excluding prompt-rewriting fees," while rival costs include prompt fees where reported. ExplainX called that an apples-to-oranges comparison.

That doesn't make the model bad. It means cost-per-score charts from a launch post deserve the same skepticism as any vendor benchmark.

How it was trained

Odyssey says the training data mixes internet video with time-stamped event annotations, gameplay recordings with matching keyboard and mouse inputs, and simulated rigid-body interactions with captions. The goal is to connect what happens on screen with the actions that caused it.

The Odyssey-3 world model starts as a multi-step video diffusion transformer and is then extended autoregressively so it can keep predicting from previous frames and action inputs. A distillation step produces a few-step variant fast enough for real-time interaction.

Robots, cars and agents

The more interesting part is transfer. Odyssey says that with only tens of hours of robot demos, Odyssey-3-based policies completed manipulation tasks and showed recovery behaviors that weren't in the demos, like reorienting a gripper after a missed grasp.

Flexion built humanoid control policies on the model and says they beat the VLA baselines it tested under lighting and environment changes. Odyssey also says it trained a driving policy on just 20 hours of Indian road data while keeping the Odyssey-3 backbone frozen.

All of these are company-reported demos. Poke & Wiggle's job is to test how consistent the capabilities are across different robots, so its results will matter more than launch clips.

What to do as a reader

If you build robotics or simulation tools, request preview access from the Odyssey-3 announcement and run your own evals before trusting the leaderboard. Check which configuration a score uses: base prompts, enhanced prompts or best-of-N.

For everyone else, the Odyssey-3 world model is a sign that world models are becoming the next AI hype lane after chatbots. Compare it with other physical-AI stories like MindOn's robot speed claims and Liquid AI's D1. Nothing here is investment advice.

Not financial advice. DYOR, ser.

More in the pool

all ai
gm ser

Get confirmed before the crowd

Daily block at 07:00 UTC. No spam, just the block, ser.