AI Potluck
Back to Gap Map Model components / Multimodal models

Cambrian

NYU VisionX
open weights / Overall score: 2.4

Cambrian is a family of open multimodal models from New York University's VisionX lab, built around a vision-centric study of how image encoders shape multimodal models. Cambrian-S extends the line to video and spatial reasoning on Qwen2.5 language models, and Cambrian-P adds per-frame camera tokens so that one model answers spatial questions about a video and estimates the camera's pose.

Openness

3 medium confidence
3.0
weights
open(safetensors for five Cambrian-P-7B variants on the Hub, ungated)
data
partial(VSI-590K, Cambrian-S-3M and Cambrian-P-Data are published under Apache 2.0, but the pose supervision needs raw ScanNet, ScanNet++ and ARKitScenes copies that the scripts do not download)
code
open(the Cambrian-P and Cambrian-S training scripts, with per-stage configs, in their repositories)
license
Apache-2.0(OSI

Cambrian-P is Apache 2.0, and New York University publishes its training scripts and most of its data, including the Cambrian-S mixture it was fine-tuned from. Its camera-pose training also needs the original ScanNet, ScanNet++ and ARKitScenes scans, which the release does not supply, so it cannot be rebuilt from what is published alone.

Adoption

1 high confidence
1.0

Hugging Face downloads over the trailing 30 days, summed across the Cambrian-1, Cambrian-S and Cambrian-P checkpoints. The lab's datasets are downloaded far more often than its models.

Capability

3 high confidence
3.0

Cambrian-S understands video and reasons about space in it, and keeps document and chart reading close to general models of its size. It has no audio input and is not shown acting in a live interface, so it sits level with MiniCPM-V.

Verified 2026-09-27