Skild AI Introduces S1 Robotics Foundation Model with In-Context Video Prompting

Robotics foundation model startup Skild AI has unveiled S1, a foundation model capable of learning physical manipulation tasks unseen during pretraining directly from a single video demonstration prompt without fine-tuning. Traditional robotic adaptation typically requires extensive task-specific teleoperation data, domain randomization, and model fine-tuning before a system can reliably execute novel actions. S1 employs in-context prompting to translate visual demonstrations directly into real

1 min
Skild AI Introduces S1 Robotics Foundation Model with In-Context Video Prompting

Robotics foundation model startup Skild AI has unveiled S1, a foundation model capable of learning physical manipulation tasks unseen during pretraining directly from a single video demonstration prompt without fine-tuning.

Traditional robotic adaptation typically requires extensive task-specific teleoperation data, domain randomization, and model fine-tuning before a system can reliably execute novel actions. S1 employs in-context prompting to translate visual demonstrations directly into real-time control policies, enabling robots to follow multi-step physical instructions up to ten minutes in duration.

Skild AI S1 Hierarchical Architecture

Benchmark Performance and Sample Efficiency

In evaluations measuring zero-shot execution on novel manipulation sequences:

  • Step-success rate: S1 recorded a 66% step-success rate on unseen manipulation tasks when prompted with a video demonstration, compared to 9% achieved by a language-prompted baseline model.
  • Demonstration efficiency: Skild AI reported that a single video demonstration delivered behavioral guidance roughly equivalent to 380 post-training physical robot examples.
  • In-context execution: The model processes video inputs as contextual prompts during inference, bypassing the need to update underlying neural network weights.

Hierarchical Architecture and Multi-Embodiment Training

S1 builds on Skild AI's omni-bodied foundation architecture, designed to decouple general task planning from embodiment-specific motor execution:

  • Two-tier policy structure: The architecture splits control into a low-frequency, high-level manipulation and navigation policy that outputs strategic objectives, and a high-frequency, low-level action policy that converts those objectives into joint angles and motor torques.
  • Cross-embodiment generalization: The system is pre-trained across diverse robot morphologies, including quadrupeds, humanoid systems, and tabletop manipulators, leveraging simulation data alongside internet video.
  • Scalable physical AI: By separating physical actuation from visual task representation, the system aims to deploy on cost-effective commercial hardware without requiring custom model architectures for each physical configuration.

Sources

Written by

More to read

  • Continuous Batching and Request Scheduling in Production LLM Serving: Comparing Orca, FastServe, Sarathi-Serve, and vLLM Architecture, Preemption Policies, Chunked Prefill Interleaving, and TTFT-TBT Trade-Offs

    Autoregressive large language model serving exhibits a fundamental architectural tension between compute utilization and latency guarantees. Standard deep learning inference pipelines rely on static request-level batching, where incoming queries are grouped into a fixed tensor, executed across forward passes until all sequences finish, and evicted simultaneously. In transformer-based text generation, static batching collapses serving efficiency. Because sequence lengths vary widely and token gen

    1 min
  • Figure AI Unveils Index Platform with 16 Million Crowdsourced Videos for Robot Foundation Models

    Humanoid robotics startup Figure AI has launched Index, a global crowdsourced data collection platform engineered to capture real-world human task demonstrations at scale. Operating in stealth for four months prior to its public unveiling, the platform has compiled 16 million video demonstrations from contributors across 108 countries, generating embodied training data for Figure's physical AI foundation models. The initiative directly targets the primary bottleneck in scaling embodied AI: the

    1 min
  • Anthropic Claude Autonomously Designs Validated Protein Binders Across 14 Targets

    Anthropic has released experimental results demonstrating autonomous de novo protein binder design using its frontier Claude models, backed by physical wet-lab validation from two independent contract research organizations. In empirical testing against 15 target proteins, Claude-designed mini-binders successfully bound to 14 targets, delivering an overall hit rate of 26.8% and a 49% binding rate for its top-ranked candidates. The campaign evaluated Claude Opus 4.8 and a preview build of Claude

    1 min