AI product Open source

TurboVLA

TurboVLA is a compact vision-language-action model and official implementation for robotic manipulation, developed by researchers at Huazhong University of Science and Technology and Huawei Technologies. Instead of routing visual input through a large language model, it independently encodes visual observations and language instructions, exchanges information through lightweight bidirectional vision-language interaction, and uses a compact decoder to predict continuous action chunks. The repository includes training and evaluation code for LIBERO and RoboTwin, with model checkpoints released separately. The project reports 32 Hz inference on an RTX 4090 with less than 1 GB of VRAM; on LIBERO, it reports a 97.7% average success rate, 31.2 ms inference latency, and 0.2 billion parameters.

View repository Mentioned in 1 video ↓

What TurboVLA is used for

1 use taken from transcripts — each links to the moment in the video.

  • A compact robot-control model that separately encodes vision and instructions, exchanges information through a lightweight bidirectional module, and predicts continuous action chunks. The video highlights its reported performance on the Libero benchmark.

Videos mentioning TurboVLA

1 in the library.