# Skild AI Unveils S1: A General-Purpose Robot Foundation Model That Learns From Watching Videos
## The Birth of a Robot Brain
A robotics startup founded in 2023 has recently introduced its flagship artificial intelligence model designed to serve as a universal brain for robots. The company has accumulated nearly $1.7 billion in funding over roughly two years, with the singular mission of building a general-purpose intelligence system that can be deployed across a wide variety of robotic platforms and tasks.
The newly revealed model, called S1, represents a significant leap in what robots can learn without requiring extensive retraining. According to the company, S1 leverages a technique known as in-context learning, which allows robots to pick up new and complex tasks simply by watching a single video demonstration of a human performing that task.
## How In-Context Learning Changes the Game
Traditional AI models deployed on robots typically require post-training every time they encounter a new task. Engineers must collect task-specific data, fine-tune the model, and run extensive tests before the robot can reliably perform the new action. This process is time-consuming and creates a major bottleneck when trying to scale robotic systems across different environments and use cases.
The new model flips this paradigm on its head. Instead of requiring task-specific retraining, operators can simply provide a video of a human completing a task as part of the model’s input — referred to as the “context.” The system then interprets the video and translates the demonstrated actions into commands the robot can execute.
What makes this especially impressive is the complexity of the tasks demonstrated. These are not short, simple maneuvers that last only a few seconds. The demonstrations include long-horizon tasks that can extend up to 10 minutes, involving sequences of coordinated actions that require the robot to maintain awareness and adapt throughout the entire process.
## Training on Every Kind of Data
One of the most distinctive aspects of how S1 was developed is the company’s decision to train the model on all four major categories of robot training data, rather than specializing in just one. Each type of data has its own strengths and limitations, and the company argues that combining them creates a more robust and capable foundation model.
**Teleoperation Data** involves a human directly controlling a robot’s movements in real time. This method produces extremely high-quality data that comes straight from the robot’s own sensors and actuators, making it highly reliable. However, it is slow to collect and tends to capture a narrower range of behaviors compared to other methods.
**Human Video Data** allows a robot to learn by observing recordings of people performing tasks. This source is abundant, diverse, and relatively easy to gather at scale. The challenge lies in the gap between what a human does and how a robot of a different form factor can replicate those actions.
**Simulation-Based Training** takes place entirely in virtual environments. It is highly scalable — millions of iterations can run on servers — and offers enormous diversity in scenarios and conditions. However, simulations do not perfectly mirror the unpredictability of the physical world, which can create a domain gap when the model is deployed on real hardware.
**Data-Capture Glove Data** involves humans wearing specialized gloves equipped with sensors while they perform tasks. This method captures fine-grained hand and finger movements with greater detail than teleoperation, and it is somewhat more scalable. Still, translating glove-recorded movements to a robot body requires significant interpretation.
The company’s philosophy is straightforward: no single data source is perfect on its own. Each compensates for the weaknesses of the others. Teleoperation data offers precision and direct robot applicability, which can offset the domain gap in video data. Human videos offer diversity and breadth, which can compensate for the limited range of teleoperation captures. By blending all four during pre-training, S1 learns a richer and more generalizable representation of how to interact with the world.
## Generalization Across Tasks and Form Factors
Rather than targeting a single industry or a narrow set of functions, the company has designed S1 to be broadly applicable. The model has demonstrated capabilities across a wide range of tasks, and the company has emphasized that it is not limited to any one vertical.
Particular attention has been given to longer-duration tasks that involve multi-step workflows. Examples include repotting a plant, preparing a cup of coffee, and cooking pancakes — activities that require sustained attention, sequential decision-making, and the ability to recover from unexpected outcomes.
One striking example involved a pancake-flipping maneuver that the model executed successfully even though the training data contained no explicit examples of flipping. The robot appeared to have inferred the technique by observing how a spatula is used in related cooking motions, then generalizing that understanding to produce the flipping action independently.
The model is also described as “omni-bodied,” meaning it is not constrained to a single robotic platform. It can operate on quadrupedal robots, humanoid robots, stationary robotic arms, and potentially other form factors. While the company has expressed interest in improving performance specifically on humanoid platforms in the future, the current architecture is designed to be form-factor agnostic.
In demonstrations from earlier versions of the model, a humanoid robot was able to adapt its behavior in real time when one of its limbs was damaged, reconfiguring its movements to continue completing the task — a sign of the kind of robustness that general-purpose robot intelligence could eventually provide.
## Is Robotics Ready for Its Own ChatGPT Moment?
The introduction of large language models (LLMs) transformed the AI landscape by enabling systems to handle a vast array of tasks without requiring task-specific training. Before the transformer-based models that powered tools like ChatGPT, AI systems were typically fine-tuned for every single new benchmark, problem, or scenario. The shift to prompt-based interaction made AI dramatically more accessible and versatile.
The company behind S1 draws a parallel to this transformation, suggesting that robot foundation models could follow a similar trajectory. The ability to simply “prompt” a robot by showing it a video, rather than going through weeks of task-specific engineering, represents a fundamental shift in how robots can be deployed.
However, the company is careful not to overstate its current readiness. While the technology is a compelling proof of concept, it is not yet at a stage where it can be handed off for unsupervised deployment in homes or workplaces. There are still significant challenges related to reliability, safety, and scalability that need to be addressed before robots powered by this kind of model become a common sight outside of controlled environments.
Looking ahead, the company has signaled that additional model releases and updates are coming. Early reports suggest that S1 is already being used in production settings to accelerate workflows and support commercial deployment efforts. The company recently completed an acquisition of Fetch Robotics, a move widely seen as a strategic step to gain the talent and infrastructure needed to scale models from research into real-world applications.
## Frequently Asked Questions (FAQ)
**What is S1?**
S1 is a robot foundation model developed by Skild AI, designed to serve as a general-purpose intelligence system that can be deployed across different types of robots and tasks.
**How does S1 learn new tasks?**
S1 uses a technique called in-context learning. By providing the model with a video of a human performing a task, the system can interpret the demonstration and translate it into actions that a robot can execute — without requiring task-specific retraining.
**What kinds of data was S1 trained on?**
S1 was trained on a combination of four data types: teleoperation data, human video recordings, simulated environments, and data-capture glove recordings. The model blends all of these sources to build a comprehensive understanding of how to perform tasks.
**Can S1 work on different types of robots?**
Yes. The model is described as “omni-bodied” and is capable of operating on quadrupedal robots, humanoid robots, stationary arms, and other form factors.
**What kinds of tasks has S1 demonstrated?**
S1 has been shown handling long-horizon tasks such as repotting plants, making coffee, and cooking pancakes — activities that involve multiple steps and can last up to 10 minutes.
**Is robotics at its “ChatGPT moment” yet?**
The company believes the current breakthrough is an important step in that direction, but acknowledges that the technology is not yet ready for widespread, unsupervised deployment in homes or businesses. Significant work remains on reliability and safety before that milestone is reached.
**What is Skild AI’s acquisition of Fetch Robotics about?**
The acquisition is intended to bring in talent and operational expertise needed to move S1 from a research model into real-world production environments and large-scale commercial deployments.
**How much funding has Skild AI raised?**
The company has raised nearly $1.7 billion since its founding in 2023.
## Conclusion
The introduction of S1 marks a pivotal moment in the pursuit of general-purpose robotic intelligence. By combining multiple data modalities and leveraging in-context learning, the model demonstrates that robots can generalize across tasks and form factors without requiring extensive retraining for each new scenario. While significant engineering and safety challenges remain before this technology can be deployed broadly in unstructured environments, the progress represents a meaningful step toward a future in which robots can be instructed as naturally as giving someone a demonstration. The coming months and years will be critical in determining whether models like S1 can transition from impressive research demonstrations to reliable, everyday tools in homes, factories, and beyond.
Thank you for reading



