The Physical Data Bottleneck: Why China is Standardizing Embodied AI
As robotics labs race to build AI systems that can interact with the physical world, the lack of standardized, real-world training data has become the industry's primary bottleneck. China's National Data Administration is now stepping in to standardize how physical AI data is collected, labeled, and shared.
By Lila Morgan
- Data Infrastructure Providers
- Argue that physical AI requires entirely new pipelines for teleoperation and synchronized multi-sensor capture.
- State Regulators
- View standardized data as the critical infrastructure needed to accelerate national commercialization of robotics.
- Enterprise Software Vendors
- Focus on integrating embodied agents into existing business processes and supply chains.
Why it matters now
Every previous wave of artificial intelligence relied on a shortcut: the internet already contained billions of text documents and images to train on. Embodied AI has no such shortcut, meaning the organizations that can standardize and scale physical data collection will control the next generation of robotics.
Inside the National Data Administration offices in Beijing this week, regulators drafted a framework that acknowledges a hard truth about the next generation of artificial intelligence: the internet is no longer enough. On September 13, China's data regulator announced a push to develop national standards for "embodied AI"—the branch of machine learning that governs physical robots. [1] The initiative aims to guide local authorities and companies in building the high-quality, large-scale datasets required to turn humanoid robots from lab curiosities into commercial products. [1][1]
Embodied AI represents the merging of digital intelligence with physical systems. Despite industry marketing that often portrays humanoid robots as fully autonomous, the actual capability relies heavily on structured environments. Unlike a large language model that processes text on a screen, an embodied AI agent—whether a warehouse logistics robot or a surgical arm—must perceive, reason, and act in three-dimensional space. [5] This requires a continuous, closed feedback loop where the machine uses multimodal perception to see its surroundings, predicts outcomes using a world model, and executes movements via physical actuators. [6][3][4]
The transition from digital to physical AI has exposed a massive infrastructure gap. Every previous wave of artificial intelligence relied on a shortcut: the training data already existed. Language models trained on billions of web pages, and computer vision models trained on hundreds of millions of uploaded photographs. [2] Embodied AI does not get that shortcut. Companies are currently attempting to build over 1 million hours of proprietary datasets specifically for physical systems, a volume gap of several orders of magnitude compared to text corpora. [1][2] There is no equivalent corpus of physical actions, because each example must be physically performed, recorded, and synchronized in the real world.[1][2]
Teaching a robot to fold laundry or navigate a factory floor requires data that maps physical interactions, not just words. [1] This involves capturing teleoperation demonstrations, egocentric video footage, and tightly synchronized multi-sensor data. [4] When a human operator remotely controls a robot to demonstrate a task, the system must record every joint angle, gripper state, force reading, and camera frame in perfect synchrony. [2] As data infrastructure provider Nferent AI notes, "Teleoperation gives you fidelity, simulation gives you scale, and human video gives you diversity." [2] This produces the exact action-state correspondence the model needs to learn the policy, stripping away the illusion that the robot is "learning" simply by watching video.[1][2]
Teaching a robot to fold laundry or navigate a factory floor requires data that maps physical interactions, not just words.
Because this data is so difficult and expensive to collect, the industry is currently fragmented. Different robotics labs use different labeling conventions, sensor formats, and sharing protocols. [1] China's National Data Administration is stepping in to solve this by developing formal standards governing how data for embodied AI is collected, labeled, stored, and shared. [1] By standardizing these formats, the regulator intends to make it easier for companies to pool resources and avoid duplicating the immense effort of physical data collection.[1]
The push for standardization is tied to aggressive industrial timelines, though much of it remains in the planning phase. Beijing has made humanoid robots a stated industrial priority, targeting notable advancements in commercialization by the end of 2026. [1] Uniform data standards lower the barrier to entry for smaller robotics firms that cannot afford to build massive proprietary datasets from scratch, allowing them to leverage shared, standardized physical data to train their models.[1]
The demand for this data is being driven by the rise of Vision Language Action (VLA) models. [3] These architectures combine visual perception, language understanding, and physical action generation into a single system. To function safely alongside humans in shared workspaces like hospitals or retail environments, VLAs require diverse training data that reflects the full range of human co-workers and physical edge cases the system will encounter. [3]
The development of physical AI remains a data problem as much as it is a hardware or model problem. [3] The systems that will eventually operate in warehouses, homes, and public spaces will only be as capable as the physical training data that shaped them. By treating embodied AI data as critical national infrastructure, regulators are attempting to build the foundation for the autonomous workforce of the next decade, moving beyond software demos to actual physical utility. [1][5][1][3]
Different angles
State Regulators
View standardized data as the critical infrastructure needed to accelerate national commercialization of robotics.
For government bodies like China's National Data Administration, the fragmentation of robotics data is a bottleneck to industrial policy. By mandating uniform standards for how physical AI data is collected, labeled, and shared, regulators aim to create a national pool of training resources. This approach treats embodied AI data not as proprietary corporate IP, but as foundational infrastructure that can lower the barrier to entry for smaller manufacturers and accelerate the deployment of autonomous systems in factories and logistics hubs.
Data Infrastructure Providers
Argue that physical AI requires entirely new pipelines for teleoperation and synchronized multi-sensor capture.
Specialized data vendors emphasize that collecting training data for robots is fundamentally different from scraping text for language models. They point out that physical AI requires 'ground truth'—exact measurements of torque, joint angles, and spatial depth captured in tight synchronization with video. From their perspective, the industry's biggest challenge is not algorithmic, but operational: building the teleoperation rigs, sourcing the human operators, and ensuring the physical safety of the collection environments.
Still unresolved
- It remains unclear how strictly the National Data Administration will enforce data sharing among fiercely competitive private robotics firms.
- The exact technical specifications for the standardized sensor formats and labeling conventions have not yet been published.
Sources
[1]BloombergState RegulatorsChina’s Data Regulator Plans Standards Push for Embodied AI
Read on Bloomberg →
[2]Nferent AIData Infrastructure ProvidersThe Three Pillars of Embodied AI Data Collection
Read on Nferent AI →
[3]SAPEnterprise Software VendorsWhat is Embodied AI?
Read on SAP →
[4]arXivEnterprise Software VendorsEmbodied Artificial Intelligence: A Comprehensive Survey
Read on arXiv →
[5]Factlen Editorial TeamEnterprise Software VendorsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Technology
See all →Memory Management
The Stop-the-World Pause: How Generational Garbage Collection Balances Throughput and Latency in Application Runtimes
7 sources
Smart Home
Google Launches $99 Gemini-Powered Home Speaker, Its First Smart Audio Hardware in Six Years
2 sources
AI Provenance
The Evidence on AI Watermarking: Does It Actually Prevent Disinformation?
5 sources
Post-Quantum Crypto
The Post-Quantum Migration: Evidence on the Race to Secure the Internet by 2030
2 sources
Every angle. Every day.
Get Technology stories with full source coverage and perspective breakdowns delivered to your inbox.




