GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
Abstract
GigaBrain-0.7 is a vision-language-action model that improves embodied generalization via a three-system architecture, large-scale heterogeneous pretraining, and joint alignment training.
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including Ο_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
Community
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π0.5, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios.
Highlights
- 1. Proposing System-3: Integrates a world model into the robotβs real-time decision loop, enabling it to simulate and evaluate before acting.
- 2. Ready Out of the Box: The dual-pyramid framework enables a single pretrained model to perform many tasks, demonstrating the "One Model, Many Tasks" capability that characterizes foundation models.
- 3. Decisive Benchmark Leadership: GigaBrain-0.7 leads task success rates by a wide margin on Maker H01, our purpose-built embodied robot, and ranks first in all four evaluations on RoboColiseum, a platform with strong sim-to-real alignment.
- 4. Real-World Tasks in One Continuous Take: An industry-first demonstration of over 20 minutes of long-horizon, complex precision operations spanning more than 10 tasks, all in one continuous take.
Resources
- π Project: https://gigaai.cc/blog/gigabrain07
- π Paper: https://arxiv.org/abs/2608.15875
- π» Code: https://github.com/open-gigaai/giga-brain-0
- π€ Model: https://huggingface.co/open-gigaai/GigaBrain-0.7-3.5B-Base
- π€ Data: https://huggingface.co/datasets/open-gigaai/GigaBrain-0.7-SampleData
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Data Pyramid for Embodied Manipulation: A Survey (2026)
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories (2026)
- Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision (2026)
- DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation (2026)
- Unleashing More Actions via Action Compositional Training for VLA Models (2026)
- From Foundation to Application: Improving VLA Models in Practice (2026)
- G0.5: One Autoregressive Stream for Robot Reasoning and Action (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.15875 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash