NYU VisionX
lab · United StatesScores
1 product on the map — 1 open-ish.
Openness
3 medium confidence- weights
- open(safetensors for five Cambrian-P-7B variants on the Hub, ungated)
- data
- partial(VSI-590K, Cambrian-S-3M and Cambrian-P-Data are published under Apache 2.0, but the pose supervision needs raw ScanNet, ScanNet++ and ARKitScenes copies that the scripts do not download)
- code
- open(the Cambrian-P and Cambrian-S training scripts, with per-stage configs, in their repositories)
- license
- Apache-2.0(OSI
Cambrian-P is Apache 2.0, and New York University publishes its training scripts and most of its data, including the Cambrian-S mixture it was fine-tuned from. Its camera-pose training also needs the original ScanNet, ScanNet++ and ARKitScenes scans, which the release does not supply, so it cannot be rebuilt from what is published alone.
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-p@main/cambrianp/scripts/Cambrian-P-7B.sh recorded 2026-09-27
The Cambrian-P-7B training script, run with --load_rec_data True and --rec_data_mode cut3r, and --rec_model_path None
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-p@main/doc/data_preparation.md recorded 2026-09-27
"Reconstruction supervision requires the original datasets"; preprocessing scripts read /path/to/raw/scannet, scannetpp and arkitscenes copies
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-p@main/LICENSE recorded 2026-09-27
Apache License, Version 2.0, "Copyright 2026 Cambrian-P Authors"
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-p@main/README.md recorded 2026-09-27
"model checkpoints, the annotated pose Cambrian-P-Data, and full training/eval code are all released"; five Cambrian-P-7B variants that "finetune from Cambrian-S-7B stage 3"; "## License See LICENSE."
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-p@main/vggt/LICENSE.txt recorded 2026-09-27
The vendored VGGT code carries "Attribution-NonCommercial 4.0 International"
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-s@main/cambrian/scripts/cambrians_7b_s4.sh recorded 2026-09-27
The Cambrian-S-7B stage-4 training script, fine-tuning from Qwen2.5-7B-Instruct
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-s@main/README.md recorded 2026-09-27
Cambrian-S models are trained on Cambrian-Alignment, Cambrian-7M, Cambrian-S-3M and VSI-590K, each linked on the Hub; four per-stage training scripts
- https://huggingface.co/api/datasets?author=nyu-visionx&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=cardData&expand[]=gated&expand[]=lastModified recorded 2026-09-27
VSI-590K, Cambrian-S-3M, Cambrian-10M, Cambrian-Alignment and Cambrian-P-Data published by nyu-visionx, each tagged apache-2.0
- https://huggingface.co/api/models/nyu-visionx/Cambrian-P-7B?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for Cambrian-P-7B: gated false, safetensors shards, no card and no license tag
- https://huggingface.co/api/models/nyu-visionx/Cambrian-S-7B?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for Cambrian-S-7B: gated false, safetensors shards, card license apache-2.0
- https://huggingface.co/datasets/nyu-visionx/Cambrian-S-3M/raw/main/README.md recorded 2026-09-27
Cambrian-S-3M "combines three video instruction datasets" with step-by-step hf download commands for LLaVA-Video-178K and LLaVA-Hound; "This dataset is released under the Apache 2.0 license"
Adoption
1 high confidenceHugging Face downloads over the trailing 30 days, summed across the Cambrian-1, Cambrian-S and Cambrian-P checkpoints. The lab's datasets are downloaded far more often than its models.
- https://huggingface.co/api/models?author=nyu-visionx&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=cardData&expand[]=gated&expand[]=lastModified recorded 2026-09-27
Cambrian-1, Cambrian-S and Cambrian-P checkpoints with 1,769 downloads in the trailing 30 days; Cambrian-S-7B 709, Cambrian-P-7B 229
Capability
3 high confidenceCambrian-S understands video and reasons about space in it, and keeps document and chart reading close to general models of its size. It has no audio input and is not shown acting in a live interface, so it sits level with MiniCPM-V.
- https://arxiv.org/html/2511.04670v1 recorded 2026-09-27
Table 13, Cambrian-S-7B stage 4: MMMU 48.0, ChartQA 74.7, OCRBench 64.8, TextVQA 76.6, DocVQA 84.8
- https://cdn.jsdelivr.net/gh/cambrian-mllm/cambrian-p@main/README.md recorded 2026-09-27
"Cambrian-P-7B ... achieves 73.7 average accuracy" on VSI-Bench, "+4.5%" over Cambrian-S-7B
- https://huggingface.co/nyu-visionx/Cambrian-S-7B/raw/main/README.md recorded 2026-09-27
"excels at spatial reasoning in video understanding"; VSI-Bench 67.5, VideoMME 63.4, EgoSchema 76.8