Cosmos
Website Framework Agent Skills Models Introduction Cosmos 3 Key Capabilities Model Architecture Model Family Supported Generation Settings Input and Output Use Cases Generator Reasoner Quickstart Generator with Diffusers Generator with vLLM-Omni Generator with NIM Generator with SGLang Reasoner with Transformers Reasoner with vLLM Reasoner with TensorRT-LLM Reasoner with NIM Troubleshooting Which CUDA version should I use? Which base container should I use? torch.cuda.isavailable() is False Import fails with libxcb.so.1: cannot open shared object file uv errors on install or sync Choosing an Integration Examples Inference Benchmarks Finetune Export and Convert Checkpoints Distill Limitations Ecosystem News License and Contact NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more. Surface Inputs Outputs Use Cases ---------------------------------------- Reasoner Text, vision Text World understanding, grounding, physical reasoning, task planning, action forecasting, embodied agent reasoning, and autonomous system decision making Generator Text, vision, sound, action Vision, sound, action World generation, world simulation, future prediction, synthetic data generation, policy learning, and robot training World understanding: Analyze videos and images for captions, temporal events, next actions, spatial grounding, physical plausibility, and causal outcomes. World generation: Produce images, videos, synchronized sound, and action-conditioned rollouts from text, image, video, or action inputs. Action modeling: Predict policy actions, inverse dynamics, and forward dynamics for robotics, camera motion, egocentric motion, and autonomous-driving settings. Research and production paths: Use Diffusers and Transformers for Python-first development, vLLM-Omni, vLLM, TensorRT-LLM, or SGLang for OpenAI-compatible serving, and NIM containers for turnkey Reasoner serving or Generator deployment for text-to-video and image-to-video generation. Post-training recipes: Adapt vision, action, and reasoner workflows with Cosmos Framework training recipes and task-specific evaluation [Coming Soon].
View Cosmos on GitHub
Website Framework Agent Skills Models Introduction Cosmos 3 Key Capabilities Model Architecture Model Family Supported Generation Settings Input and Output Use Cases Generator Reasoner Quickstart Generator with Diffusers Generator with vLLM-Omni Generator with NIM Generator with SGLang Reasoner with Transformers Reasoner with vLLM Reasoner with TensorRT-LLM Reasoner with NIM Troubleshooting Which CUDA version should I use? Which base container should I use? torch.cuda.isavailable() is False Import fails with libxcb.so.1: cannot open shared object file uv errors on install or sync Choosing an Integration Examples Inference Benchmarks Finetune Export and Convert Checkpoints Distill Limitations Ecosystem News License and Contact
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
Surface Inputs Outputs Use Cases ---------------------------------------- Reasoner Text, vision Text World understanding, grounding, physical reasoning, task planning, action forecasting, embodied agent reasoning, and autonomous system decision making Generator Text, vision, sound, action Vision, sound, action World generation, world simulation, future prediction, synthetic data generation, policy learning, and robot training World understanding: Analyze videos and images for captions, temporal events, next actions, spatial grounding, physical plausibility, and causal outcomes. World generation: Produce images, videos, synchronized sound, and action-conditioned rollouts from text, image, video, or action inputs. Action modeling: Predict policy actions, inverse dynamics, and forward dynamics for robotics, camera motion, egocentric motion, and autonomous-driving settings. Research and production paths: Use Diffusers and Transformers for Python-first development, vLLM-Omni, vLLM, TensorRT-LLM, or SGLang for OpenAI-compatible serving, and NIM containers for turnkey Reasoner serving or Generator deployment for text-to-video and image-to-video generation. Post-training recipes: Adapt vision, action, and reasoner workflows with Cosmos Framework training recipes and task-specific evaluation [Coming Soon].
Cosmos at a glance
| Stars | 12k |
|---|---|
| Forks | 892 |
| Language | Jupyter Notebook |
| License | Other |
| Last update | 2026-09-23 |
| Contributors | 47 |
How to install Cosmos
bash uvx hf@latest auth login
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
More repo topics
More free tools
Related MCP servers & CLIs
The most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.



























