AI product Open source
τ₀-VLA is an open-source hierarchical robot foundation model and reference implementation for long-horizon manipulation. A memory-augmented high-level policy generates the next subtask and uses world-model-guided test-time computation to search over alternatives when additional reasoning is needed; a generalist low-level policy then executes the selected subtask across robot embodiments. The low-level policy combines a Qwen3.5 vision-language backbone with a Mixture-of-Transformers action expert trained through conditional flow matching, using a unified 40-dimensional state/action space and multimodal co-training on tens of thousands of hours of heterogeneous real-world robot data.
The repository provides LeRobot-format data loading, prompting, masking and normalization, embodiment-specific adapters, training and post-training recipes, deployment contracts, a policy server, and open-loop evaluation tools. It includes an AgiBot World example subset and templates for other datasets and robots. Public v1 serving supports joint-control checkpoints; native end-effector data can be used for training but end-effector serving is not supported in that release. The reference environment uses Python 3.11, CUDA 12.8, and PyTorch 2.7.1. Code and model weights are released under the Apache License 2.0.
1 use taken from transcripts — each links to the moment in the video.
A vision-language-action robotics system that splits long-horizon work into high-level subtask planning and low-level generalist control. Its world model searches alternative actions when plans become uncertain, and the release includes training recipes, robot adapters, and evaluation tools.
1 in the library.