Industry leaders at WAIC 2026 said robotics needs 100 million hours of real-world data — roughly 10,000 times more than available today — before general-purpose robots can work reliably outside factories.
Embodied artificial intelligence will require roughly 100 million hours of real-world interaction data before robots can understand open-ended instructions and operate autonomously in homes and workplaces, executives and researchers said at the World Artificial Intelligence Conference in Shanghai.
"Model determines the starting point, but data defines the endgame," Yao Maoqing, partner and president of embodied business at Agibot and chairman of MeeFeng Technology, said at the conference. "When the data flywheel truly starts turning, intelligence emergence is only a matter of time."
Yao estimated that current embodied AI datasets trail large language model training data by a factor of more than 10,000 times when measured by token count or duration. Agibot, which has produced 15,000 humanoid robots and generated 10 billion yuan ($1.4 billion) in revenue since its founding in 2023, demonstrated a 3C production line in Nanchang, Jiangxi, where robots completed nearly 65,000 operations over six days with a 99.99% success rate.
The gap between controlled industrial deployments and general-purpose home robots represents both the biggest challenge and the largest market opportunity in physical AI. If the industry achieves the "ChatGPT moment" within the 2-to-5-year window most panelists predicted, the addressable market for robotic manipulation alone could reach tens of billions of dollars, reshaping supply chains from manufacturing to logistics to elder care.
Three walls blocking physical AI
Yao identified three bottlenecks — data, representation and closed-loop feedback — that he said must be resolved before physical AI can scale beyond demonstrations. Real-world interaction data remains scarce and expensive to collect, unlike web text that can be crawled at near-zero marginal cost. A unified physical representation that works across tasks, environments and robot embodiments does not yet exist. And trial-and-error in the physical world carries high costs: a single failure can damage components or destroy an entire robot.
Physical Intelligence, the San Francisco-based robotics startup, is tackling the data challenge through what it calls "context expansion." Ren Zhiyi, a research scientist at PI, said the company's latest π0.7 model incorporates historical video encoding, fine-grained language and image guidance, and metadata such as data quality scores and error labels. The approach allows a single model checkpoint to handle multiple tasks without task-specific fine-tuning. PI also demonstrated cross-embodiment transfer, where skills learned on a low-cost static robotic arm transferred to an industrial UR5 arm without additional data collection.
Dyna Robotics, founded in September 2024, has accumulated more than 200,000 hours of pre-training data using a "data pyramid" strategy: internet and human first-person video at the base, multi-task real-robot data in the middle, and real deployment data — including failure and recovery trajectories — at the top. Ma Yecheng, co-founder and chief scientist, said the company achieved 24-hour autonomous napkin folding with more than 800 napkins completed without human intervention. In tasks such as celery cutting, high-precision cup stacking and thin bandage sorting, one to two hours of post-training data enabled repeatable operation.
The data pyramid and the human advantage
Researchers at the conference converged on a layered approach to data collection. Xu Danfei, assistant professor at the Georgia Institute of Technology, argued that human first-person data — captured through head-mounted cameras — could become a critical training source. His team has accumulated roughly 20,000 hours of egocentric data since 2023 and observed scaling laws emerging in complex manipulation tasks as data volume grew.
"Humans are just another embodiment of a robot," Xu said. His work with Meta on hand-tracking and first-person video aims to erase the boundary between human demonstration and robot training data.
Genesis AI co-founder Wang Zunxuan advocated for combining human glove data, first-person video, real-robot data and simulation. Glove data provides tactile information that video alone cannot capture, while simulation enables大规模 evaluation and reinforcement learning without hardware wear. The company's Genesis physics engine focuses on pushing simulation fidelity to the point where it can reliably replace real-world evaluation.
Timeline debate: 2 years vs 5 years
In a closing panel, six industry leaders gave their estimates for when robots would achieve "ChatGPT moment" — defined as the ability to understand open-ended natural language instructions and perform everyday tasks out of the box. Yao predicted two years, contingent on data progress. Ren of PI said four to five years, noting he had thought it would take 10 years before joining the company. Xu estimated five years. Ma of Dyna Robotics said four years. Zhang Zhengyou, chief scientist at Tencent Robotics X, said three to five years. Zhao Zihao, CEO of Sunday Robotics, said less than three years.
"These three to five years are still promising, but we need to work steadily," Zhang said. He cautioned against expecting a single model or data source to deliver a sudden breakthrough, emphasizing that progress depends on continued advances across data collection, real-world deployment, failure recovery, simulation evaluation and hardware systems.
Agibot's Yao said the company's 100-robot distributed reinforcement learning system, deployed in a facility in Jiaxing near Shanghai, represents a new paradigm for post-training. The system supports sub-second model updates across the fleet, compressing training cycles from days to minutes. "This is not single-robot demonstrations in a laboratory," he said. "This is physical verification at scale, in real environments, on real tasks."
This article is for informational purposes only and does not constitute investment advice.