A compact robot-control model that separately encodes vision and instructions, exchanges information through a lightweight bidirectional module, and predicts continuous action chunks. The video highlights its reported performance on the Libero benchmark.