AI product Open source

FoundationPose

FoundationPose is an NVIDIA Research implementation of a unified foundation model for 6D object pose estimation and tracking of novel objects. It supports model-based operation from a CAD model and model-free operation from a small set of reference images, applying to a new object at test time without fine-tuning. A neural implicit representation provides novel-view synthesis while keeping the downstream pose-estimation modules shared across both setups; the system uses large-scale synthetic training, a transformer-based architecture, contrastive learning, and language-model assistance. The repository includes robotic and augmented-reality demos and provides Docker-based setup instructions and downloadable model weights.

View repository Mentioned in 1 video ↓

What FoundationPose is used for

1 use taken from transcripts — each links to the moment in the video.

  • Extracts and tracks object poses from a human video demonstration so the robot policy can follow a sequence of goal poses.

Videos mentioning FoundationPose

1 in the library.