Figure AI Unveils Index Platform with 16 Million Crowdsourced Videos for Robot Foundation Models

Humanoid robotics startup Figure AI has launched Index, a global crowdsourced data collection platform engineered to capture real-world human task demonstrations at scale. Operating in stealth for four months prior to its public unveiling, the platform has compiled 16 million video demonstrations from contributors across 108 countries, generating embodied training data for Figure's physical AI foundation models. The initiative directly targets the primary bottleneck in scaling embodied AI: the

2 min
Figure AI Unveils Index Platform with 16 Million Crowdsourced Videos for Robot Foundation Models

Humanoid robotics startup Figure AI has launched Index, a global crowdsourced data collection platform engineered to capture real-world human task demonstrations at scale. Operating in stealth for four months prior to its public unveiling, the platform has compiled 16 million video demonstrations from contributors across 108 countries, generating embodied training data for Figure's physical AI foundation models.

The initiative directly targets the primary bottleneck in scaling embodied AI: the absence of vast, standardized real-world physical interaction data. While language models scale on publicly available internet text, general-purpose robotics require fine-grained demonstrations of manipulation, spatial navigation, and household physics across diverse environments.

Figure Index Data Ingestion Pipeline

Platform Architecture and Economics

Index utilizes a distributed mobile application where registered participants record themselves executing specific physical actions, such as opening cupboards, sorting laundry, manipulating hand tools, and navigating cluttered indoor layouts.

Key metrics disclosed by Figure at launch include:

  • 16 Million Video Submissions: Over 16 million verified clips uploaded across 108 countries, providing geographic and architectural diversity that resists training bias.
  • User Scale: 264,000 app downloads during the four-month stealth phase, with 43,000 weekly active contributors.
  • Ingestion Velocity: The platform currently processes more than 30 minutes of uploaded task video every second.
  • Direct Payouts: Figure has distributed $15 million in direct payments to contributors for validated task submissions.

The crowdsourcing model complements Figure's existing data streams, which include enterprise site telemetry, teleoperation by specialized operators, and residential environment data collected through corporate partnerships like its agreement with Brookfield Properties.

Integration into the Helix Multimodal Stack

Figure feeds Index video data into Helix, its multimodal embodied AI architecture. Helix synthesizes three distinct layers of data to train general-purpose policy models:

  1. High-precision human teleoperation traces capturing joint torque and force feedback.
  2. In-situ operational logs from deployed humanoid fleets, including Figure 02 deployments at BMW's Spartanburg manufacturing plant.
  3. Scaled egocentric and third-person video of everyday human object manipulation from the Index network.

By mapping diverse human hand and body trajectories into robot kinematics, Figure aims to generalize zero-shot task execution without needing custom per-task programming or fragile simulation-only domain transfers.

Hardware Scaling and Market Deployment

The launch of Index follows Figure's $1 billion Series C round in late 2025 at a $39 billion valuation, alongside hardware manufacturing expansion at its BotQ facility, where the company produces Figure 03 humanoid units. Figure has stated plans to allocate over $1 billion toward data acquisition and compute infrastructure over the next 12 months as it prepares commercial robotics-as-a-service offerings for industrial and domestic environments.

Sources

Written by

More to read

  • Synthetic Data Pipelines in Production LLM Post-Training: Architecture, Prompt Evolution, Quality Filtering, and Contamination Control

    Synthetic Data Pipelines in Production LLM Post-Training: Architecture, Prompt Evolution, Quality Filtering, and Contamination Control Scaling supervised fine-tuning (SFT) and preference alignment (DPO, PPO, GRPO) through human annotation faces severe economic and operational constraints. Human annotation costs between $5.00 and $50.00 per complex instruction-response trajectory, exhibits significant variance across labeler cohorts, and scales linearly with dataset volume. The LIMA study by Zho

    1 min
  • Anthropic Outlines $30 Trillion Total Addressable Market in Pre-IPO Pitch

    Anthropic is preparing to pitch prospective initial public offering investors on a total addressable market exceeding $30 trillion, according to a report from The Wall Street Journal. The projection relies on estimating the total monetary value of human labor and enterprise workflows that advanced AI models and autonomous agents could potentially automate or augment across the global economy. If presented in formal registration filings, the $30 trillion figure would surpass the previous record

    1 min
  • Induction Heads: Mathematical Foundations, Two-Layer Circuit Composition, and the Emergence of In-Context Learning in Transformers

    Induction Heads: Mathematical Foundations, Two-Layer Circuit Composition, and the Emergence of In-Context Learning in Transformers One of the defining capabilities of modern autoregressive large language models is in-context learning: the ability to infer rules, adapt to task formats, and execute complex few-shot instructions purely from prompt context without updating neural network weights. For years, in-context learning was treated as an enigmatic, emergent property of large-scale autoregres

    1 min