Dyna Robotics on Monday introduced a new robot foundation model DYNA-2, which it says was trained on over a million hours of human video and no robot data.
The US-based company believes this approach could help solve the cost issues that have limited general-purpose robots.
DYNA-2 aims to solve a common problem in building general-purpose robots, which is getting enough data. Most robots learn through teleoperation, where a person guides the machine through tasks for hours.
The issue here is that the method is slow, costly, and hard to scale, which limits how advanced these systems can get.
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000… pic.twitter.com/wZamR0axzS
— Dyna Robotics (@DynaRobotics) August 10, 2026
Dyna’s solution is to leave out the robot during training. DYNA-2, instead, was trained only on egocentric human video, which is footage from a person’s point of view as they interact with objects.
The company says this adds up to about 170 years of continuous experience. Co-founder Jason Ma explained, “Action data is scarce, but video is everywhere,” arguing that physical intuition “can be learned directly from human video” instead of using millions of hours on a robot arm.
On high-precision manufacturing tasks, the success rates rose from 20% to between 80% and 90%. Dyna credits this improvement to the scale of pre-training.
Across 15 benchmark tasks, models trained on more human video consistently outperformed those trained on less human video.
Dyna stands out because of its unique architecture. Instead of using a vision-language model, DYNA-2 is a World-Action Model based on video generation.
During pre-training, it predicts both the next video frame and the next action at the same time.
Dyna says this dual goal gives the model an internal sense of contact physics and spatial reasoning. The company hopes that knowledge from human videos will transfer to different robots, like stationary arms, humanoid prototypes, and dexterous hands, even though these devices were not used in pre-training.
Adapting to a new platform then takes just hours of fine-tuning instead of weeks of collecting new data. For example, Dyna reports that only 13 minutes of data taught robotic hands to twist off a bottle cap.
Dyna already has robots in real-world settings, which sets it apart from other robotics firms.
DYNA-1 robots are being used in hotels, restaurants, and laundromats, including gyms, according to reports.
The company was founded by Lindon Gao, York Yang, and Jason Ma, who previously worked at DeepMind.
Dyna’s stated plan is to push the same approach to 10 million hours of video, which it frames as a collection problem rather than a fleet-building one.
The smartest crypto minds already read our newsletter. Want in? Join them.