Geometry-grounded vision models, such as VGGT, have emerged as robust visual encoders, providing essential geometric priors for robotic manipulation. However, the high computational cost of these models often leads to slow inference, limiting their practical applications in real-world robotics. This paper introduces eVGGT, a lightweight geometry-aware vision encoder distilled from the high-performing VGGT. Our findings demonstrate two primary advantages: i) integrating eVGGT into imitation learning frameworks (including ACT and Diffusion Policy) yields up to a 6.3% improvement in success rate over standard 2D encoders across bimanual and single-arm tasks in both simulation and real-world settings with variable viewpoints; ii) eVGGT achieves a nearly 5× speedup and a 63% reduction in memory usage compared to state-of-the-art geometry-aware encoders while maintaining comparable task performance. These results suggest that eVGGT substantially alleviates the performance-latency bottleneck that has limited geometry-aware visuomotor policies in real-world deployment.
We present eVGGT, an efficient geometry-aware vision encoder for visuomotor policies.
This paper explores the replacement of conventional 2D vision encoders in robotic manipulation with a geometry-aware encoder to more effectively capture the 3D global context.
Incorporating eVGGT improves average success rates over the original RGB-based Diffusion Policy by 6.3 percentage points on RoboTwin and 6.2 percentage points on ManiSkill.
Figure: Geometry-aware representations for robotic manipulation.
eVGGT is our proposed lightweight geometry-aware encoder distilled from VGGT. When integrated into Diffusion Policy, it achieves nearly 5× faster inference and 63% lower GPU memory usage than VGGT, while maintaining comparable manipulation performance.
We apply knowledge distillation to train a compressed eVGGT from VGGT, reducing the number of transformer blocks from 24 to 4 and substituting DINOv2-ViT-S for DINOv2-ViT-L. Training uses geometric supervision from VGGT, a depth-gradient loss, and student-only data augmentation.
Figure: Efficiency comparison of eVGGT and VGGT.
This website template is adapted from HyperNeRF.