Skip to main content
Zero-Shot RoboticsCapability Milestone· 3 min read· in Artificial Intelligence

Figure's Helix 2.5 Robot Achieves 56% Success Rate in Unseen Homes Without Prior Mapping

Figure AI's latest humanoid robot successfully completed household chores in 30 unfamiliar homes on its first attempt, marking a major milestone in zero-shot generalization for physical AI.

By Mateo Ramos

Physical AI Optimists 60%Robotics Skeptics 40%
Physical AI Optimists
Argue that the 56% success rate proves scaling laws apply to robotics.
Robotics Skeptics
Highlight the gap between a 56% success rate and commercial viability.

Perspectives this story doesn't cover

  • Homeowners who participated in the trials
  • Consumer safety regulators

Why it matters

Until now, humanoid robots required extensive pre-mapping and controlled environments to function reliably. By proving a machine can walk into a completely unfamiliar house and immediately fold laundry or make a bed with better than coin-flip odds, Figure has demonstrated that general-purpose physical AI is crossing the threshold from laboratory theory to real-world deployment.

The moment a humanoid robot crosses the threshold of a front door it has never seen, its ability to function is entirely determined by its spatial inference engine. If the machine cannot instantly map the unfamiliar geometry of a living room or calculate the height of a kitchen counter without prior training data for that specific layout, its physical hardware is useless. Figure AI has now proven that its spatial inference architecture can handle that exact transition, deploying its Helix 2.5 robot into 30 completely unfamiliar homes and achieving a 56% success rate on household chores on the very first attempt.[3][4][5]

The mechanism driving this capability is known as zero-shot generalization. Instead of relying on pre-programmed waypoints or 3D lidar scans accurate to within 2 millimeters, the Helix 2.5 uses a vision-language model paired with real-time physical AI to interpret its surroundings dynamically. When instructed to "make the bed," the robot must identify what a bed looks like in this specific lighting, locate the pillows, and calculate the tension required to pull the sheets—all without having ever practiced in that specific bedroom.[4][5]

A 56% success rate might sound low for a human housekeeper, but in the context of autonomous robotics, it represents a massive leap. According to Forbes, this marks a six-fold improvement over previous generations of physical AI attempting unmapped, zero-shot tasks, which historically hovered near a 9% success baseline. The robot successfully folded towels, made beds, and cleaned up clutter in environments it had never encountered before.[1][6]

The Helix 2.5 achieved a 56% success rate in zero-shot environments, a six-fold improvement over previous physical AI benchmarks.

To ensure the results were genuinely zero-shot, Figure AI sent the Helix 2.5 into 30 distinct homes belonging to strangers. The robot was not given floor plans, and the homeowners did not modify their spaces to accommodate the machine. It had to navigate around everyday obstacles, adjust to varied lighting conditions across different times of day, and manipulate objects of wildly different shapes, weights, and textures.[2][3][5]

To ensure the results were genuinely zero-shot, Figure AI sent the Helix 2.5 into 30 distinct homes belonging to strangers.

Figure's billion-dollar bet on physical AI is predicated on the idea that the world is already built for humans, and therefore, a general-purpose robot must be able to operate in human spaces without requiring those spaces to be retrofitted. This 30-home trial serves as a proof of concept that the software bottleneck—teaching a machine to generalize physical tasks across infinite variations—is beginning to crack.[1][5]

While the 56% success rate is a breakthrough for zero-shot execution, it also means the robot failed 44 times out of 100 on average. The failures often stem from edge cases in physical manipulation—a towel slipping from a mechanical gripper, a sheet snagging on a mattress corner, or a miscalculation of a fragile object's weight. The mechanism for recovering from these physical errors, rather than simply freezing or dropping the object, remains the next major hurdle for Figure and the broader robotics industry.[3][4]

Physical manipulation of soft objects, like bedsheets and towels, remains one of the most complex challenges for zero-shot robotics.

The push for general-purpose humanoids is accelerating, with companies attempting to bridge the gap between 100-billion-parameter language models and physical actuators with 40 or more degrees of freedom. Figure's demonstration proves that the same scaling laws that improved text generation over the last four years are now yielding tangible results in physical space, moving the industry closer to commercially viable household assistants.[1][2]

The true test of the Helix 2.5 architecture will not be whether it can reach a 90% success rate in these specific 30 homes, but whether it can maintain its current 56% success rate as the complexity of the tasks scales up. Figure's engineering teams are now focused on refining the robot's error-recovery loops, aiming to deploy the next iteration of the hardware into commercial and industrial environments by late 2027.[4][5]

What to know

  • Figure AI's Helix 2.5 humanoid robot completed household chores in 30 unfamiliar homes without prior mapping.
  • The robot achieved a 56% success rate on its first attempt at tasks like folding towels and making beds.
  • The trial represents a six-fold improvement in zero-shot generalization for physical AI systems.
  • Engineers attribute the leap to a new spatial inference engine paired with vision-language models.

Sources

Source coverage

6 outlets

2 viewpoints surfaced

Physical AI Optimists 60%Robotics Skeptics 40%
  1. [1]ForbesPhysical AI Optimists

    Figure's Billion-Dollar Physical AI Bet Delivered A 6X Jump In Robot Chore Success

    Read on Forbes →
  2. [2]HighTechDad™Robotics Skeptics

    Tech Brief: Figure's Helix 2.5 Robot Completed Household Tasks in 30 Homes It Had Never Entered - HighTechDad™

    Read on HighTechDad™ →
  3. [3]TechRepublicRobotics Skeptics

    Figure's Robot Entered 30 Unseen Homes — and Succeeded 56% of the Time

    Read on TechRepublic →
  4. [4]Unite.AIPhysical AI Optimists

    Figure Introduces Helix 2.5, Tested Zero-Shot in 30 Unseen Homes - Unite.AI

    Read on Unite.AI →
  5. [5]Figure AIPhysical AI Optimists

    Helix 2.5: Zero-Shot 30-Home Generalization

    Read on Figure AI →
  6. [6]The National

    Video: Robot walks into 30 stranger homes — then makes beds, folds towels and cleans up

    Read on The National →

Comments

Stay informed

Every angle. Every day.

Get Artificial Intelligence stories with full source coverage and perspective breakdowns delivered to your inbox.