AI product Open source

τ₀-VLA

τ₀-VLA is an open-source hierarchical robot foundation model and reference implementation for long-horizon manipulation. A memory-augmented high-level policy generates the next subtask and uses world-model-guided test-time computation to search over alternatives when additional reasoning is needed; a generalist low-level policy then executes the selected subtask across robot embodiments. The low-level policy combines a Qwen3.5 vision-language backbone with a Mixture-of-Transformers action expert trained through conditional flow matching, using a unified 40-dimensional state/action space and multimodal co-training on tens of thousands of hours of heterogeneous real-world robot data.

View repository Mentioned in 1 video ↓

Overview

The repository provides LeRobot-format data loading, prompting, masking and normalization, embodiment-specific adapters, training and post-training recipes, deployment contracts, a policy server, and open-loop evaluation tools. It includes an AgiBot World example subset and templates for other datasets and robots. Public v1 serving supports joint-control checkpoints; native end-effector data can be used for training but end-effector serving is not supported in that release. The reference environment uses Python 3.11, CUDA 12.8, and PyTorch 2.7.1. Code and model weights are released under the Apache License 2.0.

What τ₀-VLA is used for

1 use taken from transcripts — each links to the moment in the video.

  • A vision-language-action robotics system that splits long-horizon work into high-level subtask planning and low-level generalist control. Its world model searches alternative actions when plans become uncertain, and the release includes training recipes, robot adapters, and evaluation tools.

Videos mentioning τ₀-VLA

1 in the library.