AI product Open source
ABot-Recon is a streaming 3D-reconstruction model and inference tool for reconstructing long video streams from video frames alone. At each step, it caches key-value features from the preceding 11 frames, predicts a point map in the current camera coordinate system, estimates the adjacent relative camera pose, and composes those poses sequentially into a global trajectory and point cloud without persistent learned long-range memory. It can output camera poses, relative poses, local and world point maps, colors, confidence maps, and run metadata.
An optional loop-closure backend retrieves candidate revisited-frame pairs using DINOv2-SALAD descriptors, predicts relative-pose constraints with ABot-Recon, and refines the trajectory through sparse pose-graph optimization. The repository provides Python and command-line interfaces, a public model checkpoint, a demo, evaluation code, and visualization utilities. The released configuration targets Linux, Python 3.10 or later, PyTorch 2.5.1, and CUDA 12.1; source code is licensed under Apache License 2.0, while the model weights have separate terms in MODEL_LICENSE.md.
1 use taken from transcripts — each links to the moment in the video.
Creates a 3D reconstruction from video by estimating geometry and camera motion while retaining only a small local frame context, with optional loop-closure refinement.
1 in the library.