← Back to Issue

Dyna-2 Proves Scaling Laws for Robotics: 1 Million Hours of Human Video Unlocks Zero-Shot Dexterity

From aiste.ulozaite@gmail.com · original ↗ · unsubscribe

Dyna-2 is a world-action model (WAM) pre-trained on over one million hours of human video data. Its existence proves that scaling human video data predictably improves zero-shot performance on unseen robot hardware. Dyna-2 achieved an 87% pass rate in real-world zero-shot deployments at customer sites, vastly outperforming the 46% pass rate of its VLA predecessor, Dyna-1. The model uses a new one-step video generation distillation pipeline that drops inference latency by two orders of magnitude for downstream planning.


Published on

Monday, August 10, 2026

Dyna-2 Proves Scaling Laws for Robotics: 1 Million Hours of Human Video Unlocks Zero-Shot Dexterity

Humanoids Daily

Written byHumanoids Daily

Advertisement

Advertisement

Key Takeaways

Hide

  • Dyna Robotics has unveiled Dyna-2, a world-action model (WAM) pre-trained on over one million hours of human video data.
  • The model demonstrates a first-of-its-kind human-to-robot transfer scaling law, proving that scaling human video data predictably improves zero-shot performance on unseen robot hardware.
  • Dyna-2’s architecture relies on video co-training (predicting future video states), which the company claims is essential for cross-embodiment generalization, directly challenging the industry standard of Vision-Language-Action (VLA) models.
  • In real-world zero-shot deployments at customer sites, Dyna-2 achieved an 87% pass rate, vastly outperforming the 46% pass rate of its VLA predecessor, Dyna-1.
  • The release also introduces a novel one-step video generation distillation pipeline, dropping inference latency by two orders of magnitude for downstream planning.

The robotics industry has spent the last year fiercely debating the architecture of physical intelligence, caught between fine-tuning existing language models and building native “world models” from scratch. Today, Dyna Robotics delivered what may be the strongest empirical evidence yet for the latter, unveiling Dyna-2—a world-action model (WAM) pre-trained on a staggering one million hours of human video data.

According to the company’s technical report and an accompanying social media thread, this massive scale has unlocked a holy grail of embodied AI: a human-to-robot transfer scaling law. In short, Dyna-2 proves that feeding a model more human video predictably improves its ability to control a robot it has never seen before.

Bridging the Embodiment Gap

The historical bottleneck in robot learning has been the data itself. While teleoperation yields high-quality, action-labeled data, it is slow and expensive to collect. The theoretical alternative is to learn from the boundless supply of human video on the internet, but translating a human hand’s movement into a robotic gripper’s action—the “embodiment gap”—has proven exceedingly difficult.

Dyna-2 attacks this problem purely through scale and objective design. The company curated nested subsets of egocentric human manipulation videos, scaling from 1,000 to 1,000,000 hours, keeping proportions from each source identical. When evaluated zero-shot on 39 distinct robot tasks across two stationary, bimanual platforms, Dyna-2’s performance improved monotonically as the human pre-training data increased. An inflection point emerged between 10,000 and 100,000 hours, suggesting that cross-embodiment knowledge transfer emerges naturally if the model simply sees enough human activity.

Furthermore, this zero-shot capability extended to post-training. With just a few hours of robot-specific data and zero human-robot alignment, post-trained Dyna-2 models solved tasks ranging from manipulating deformable objects to untwisting bottle caps.

[

Dyna Robotics

](https://x.com/DynaRobotics/status/2086856327150858298)

[

Dyna Robotics

](https://x.com/DynaRobotics/status/2086856327150858298)

@DynaRobotics

·Follow

Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws: • world-action models exhibit scaling law on human data across four orders of magnitude, from 1000

Watch on X

4:44 PM · Aug 10, 2026

[

3.1K](https://x.com/intent/like?tweet_id=2086856327150858298)[

Reply](https://x.com/intent/tweet?in_reply_to=2086856327150858298)

Copy link

Read 188 replies

Highlights & notes

    Notes