Helix is a proprietary Vision-Language-Action (VLA) model developed by Figure AI for generalist humanoid control. It unifies perception, language understanding, and motor action in a single model. Figure AI announced Helix on February 20, 2025, under the title "A Vision-Language-Action Model for Generalist Humanoid Control."
Architecture
VLA models extend language-model architectures to output not just text but motor actions. In the manner that GPT-4V unifies text and vision modalities, Helix treats actuation as a first-class output modality. Robotics tasks are framed as multimodal generation problems, taking a camera feed and a prompt as input and producing joint commands and reasoning as output, rather than as traditional robotics-controller problems.
Figure AI describes Helix as the first publicly announced VLA model specifically targeting generalist humanoid control.
Related models and approaches
Helix represents an integrated stack in which Figure AI builds both the humanoid hardware and the controlling model. This contrasts with the embodiment-flexible foundation-model approach pursued by Physical Intelligence, whose models are intended to generalize across robot bodies rather than a single integrated platform.
Relationships
- developer: Figure AI
- related: The Industrial Explosion (Davidson, Hadshar), Physical Intelligence (alternative approach), Clone Robotics, Sora and Veo (Video Generation Models) (another multimodal direction), Multimodality, planned AI Robotics.
Sources
- Figure AI / Helix: A Vision-Language-Action Model for Generalist Humanoid Control (2025-02-20)