AI product Open source

LingBot-Map

LingBot-Map is a feed-forward 3D foundation model from the Robbyant Team for streaming 3D reconstruction from image sequences and video. Its Geometric Context Transformer unifies coordinate grounding, dense geometric cues, and long-range drift correction through anchor context, a pose-reference window, and trajectory memory.

View repository Mentioned in 1 video ↓

Overview

The system uses paged KV-cache attention for streaming inference and supports interactive browser-based reconstruction with point-cloud visualization, keyframe intervals, windowed processing for long sequences, and sky masking. It also provides an offline rendering pipeline for image folders or video, along with evaluation pipelines for datasets including KITTI and Oxford Spires. The repository reports approximately 20 FPS inference at 518×378 resolution on sequences exceeding 10,000 frames, and documents PyTorch installation with FlashInfer as the recommended attention backend and native SDPA as a fallback.

What LingBot-Map is used for

1 use taken from transcripts — each links to the moment in the video.

  • An open-source 3D foundation model that reconstructs scenes from streaming video. It can produce point clouds or render long camera flights from an image folder or video.

Videos mentioning LingBot-Map

1 in the library.