Ant Group Open-Sources LingBot-Map, a Streaming 3D Reconstruction Model for Monocular Video
Ant Group's embodied-AI division has open-sourced LingBot-Map, a streaming monocular 3D reconstruction model released under Apache 2.0 on April 16, 2026. The model processes ordinary RGB video in real time, estimating camera pose and scene geometry frame by frame without requiring LiDAR or offline batch processing. A key innovation is its Geometric Context Transformer, which uses a paged KV-cache to reduce per-frame memory growth by roughly 80 times compared to full causal attention, enabling inference beyond 10,000 frames. This allows the model to run at approximately 20 FPS with FlashInfer enabled, nearly double the 10.5 FPS baseline, making live scene reconstruction practically viable. Three model checkpoints are available on Hugging Face, and benchmark results from the authors show significantly lower pose error compared to competing methods on both sparse and dense sequences.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in