Figure published Helix 2.5 on 17 September. The number they put on the page is 237 of 420 full-task trials — 56% — across 30 Bay Area homes the robot had never seen. That is more than half. It is also 183 chores that did not finish. I think both facts belong in the same sentence.

The company calls this the first demonstration of zero-shot whole-body generalization at this scope on a humanoid. I believe the test is real, and I also think 56% is the honest way to read it. A housekeeper who fails 44% of the work does not get a second week.

Key takeaways

  • Helix 2.5: 237 of 420 full-task wins, 56%, in 30 unseen homes.
  • Index pretraining lifted a from-scratch policy from 9% to 56% on the same chores.
  • Beds 67%, towels 62%, toys 40%. Success meant the whole job — no partial credit.
  • "Zero-shot" applies to the houses and the objects. The three behaviors were trained elsewhere.
  • You still cannot buy a Figure 03. Our catalog page still lists the brain as Helix 02.

237 of 420 — the score Figure published

Brett Adcock said Figure rented the homes. Three chores: make the bed, fold the towels, tidy the living room. Success meant the entire job. Every one of 13 to 15 scattered toys in the basket. Every towel folded and placed. Both pillows and the comforter corners at the top of the bed, comforter pulled smooth. A safety stop counted as a fail. Timeouts counted as a fail. No partial credit.

Adcock and Figure's director of AI, Corey Lynch, also said the robot recorded a success in every one of those 30 houses. That is reach. It is not a 56% house that always works. The pooled rate is still 237 wins and 183 losses, and Figure ran the grading.

"Zero-shot" here meant the house, not the chore

Figure's own post qualifies the word. Zero-shot refers to the evaluation environments and the objects being manipulated. The three behaviors were specified with fine-tuning data collected somewhere else. Each task used one fixed checkpoint across all 30 homes. No weights were adapted after the robot arrived. No evaluation rollout was used to pick the checkpoint.

So the robot did not walk in and invent a new chore. It brought a trained bed-making policy into a bed it had not trained on, and a tidy policy into a living room whose toys were set aside before the test began. That is still a hard transfer problem. Whole-body control, locomotion, bimanual work, deformable cloth. A VLA model — vision-language-action, the stack that maps what the cameras see onto the next motor command — has to do all of that in a house that was not in the fine-tune set. It is not the sentence "the robot showed up and learned housework on the spot."

Index moved the needle from 9% to 56%

Figure trained two policies on identical task data. One started from random weights. One started from Index, the human-behavior dataset they took out of stealth on 25 August. Architecture, optimization, downstream data, evaluation — held fixed. The from-scratch policy succeeded on 9% of the zero-shot trials. The Index-pretrained policy succeeded on 56%. That gap is the result they want remembered, and it is a company-run comparison.

Index is humans filming chores on a phone. The August launch post said 264,000 downloads, 44,000 weekly active users, 16 million videos, and $15 million paid to creators. The Helix 2.5 post says the pipeline now generates about 35 minutes of new human experience every second, and that Figure has committed $3.5 billion of compute to training Helix. Those are their numbers. They also report a scaling law: four nested subsets of Index, an 8× increase in pretraining data, model size held fixed, and next-action prediction loss falling smoothly enough to forecast the largest run to four decimal places. Forecasting error was 0.54% of the variation across that range. That is a training-loss result. It does not, by itself, say 56% becomes 90% on the next doubling.

Learn broadly in pretraining, specify a behavior once, generalize at deployment. Figure wrote that recipe down. 56% is how far it has traveled in a stranger's house.

Beds were easier than toys

TaskWinsRateWhat counted as a win
Bed making 94 / 140 67% Both pillows and comforter corners at the top third; comforter pulled smooth
Towel folding 87 / 140 62% Every towel folded and placed in the basket
Tidying toys 56 / 140 40% All 13–15 scattered toys picked and placed in the basket
All three 237 / 420 56% Full task only — no credit for a half-made bed

A comforter has two corners and a known top of the bed. Toys hide under couches. The tidy task is a search problem as much as a grasp problem, which is why 40% sits so far below the other two. Figure also describes whole-body self-correction — stepping back, changing stance, walking around the bed to fix a fold — as a qualitative jump from Helix 02. Useful when it works. The table is still the score.

The warehouse marathon was a different test

We already wrote about Figure 03 sorting 100,000 packages in 81 hours. That run, and the Helix 02 work Figure cites — dishwasher unloading, a logistics task they say ran autonomously for 200 hours — learned from data collected where the robots would operate. A warehouse is one building. You can instrument it. You can collect there again.

A home is 30 different buildings. Figure says Helix 2.5 matched a Helix 02 success rate on the same chore with half the adaptation data, and did it across unseen houses. That is the claim that matters if you care about a product that visits more than one address. It is also why this piece is a sequel to the warehouse demo, not a rewrite of it. The sim-to-real gap is one transfer problem. House-to-house is another. Helix 2.5 is aimed at the second.

You still cannot order a Figure 03

Figure 03 still has no public order page. Our catalog card still lists the brain as Helix 02, because we have not retagged a score we have not re-run. The hardware line from 01 to 03 is real. Brett Adcock has been saying the home is the product since the company started. None of that is a checkout button.

If you want a home humanoid you can actually put a deposit on, 1X's NEO is the one with a price: $20,000 early access or $499 a month, US deliveries still promised for late 2026. We have been writing when the home robot is actually ready for a while. Helix 2.5 moves the research line. It does not move the buy line.

A housekeeper needs the other 44%

56% is a research result I believe as a company-run test. It is not a housekeeper. The next number that matters is the same three tasks, run again, after someone else has been living in the house. A chair that moved overnight. A dog. A chore that gets interrupted halfway. Until that number is published, Helix 2.5 is the best evidence Figure has that Index transfers. It is not evidence that the robot is ready to live with you.

Back to What's Next The 81-hour package sort