Navid MoazzezNavid Moazzez

Cosmos

Website Framework Agent Skills Models Introduction Cosmos 3 Key Capabilities Model Architecture Model Family Supported Generation Settings Input and Output Use Cases Generator Reasoner Quickstart Generator with Diffusers Generator with vLLM-Omni Generator with NIM Generator with SGLang Reasoner with Transformers Reasoner with vLLM Reasoner with TensorRT-LLM Reasoner with NIM Troubleshooting Which CUDA version should I use? Which base container should I use? torch.cuda.isavailable() is False Import fails with libxcb.so.1: cannot open shared object file uv errors on install or sync Choosing an Integration Examples Inference Benchmarks Finetune Export and Convert Checkpoints Distill Limitations Ecosystem News License and Contact NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more. Surface Inputs Outputs Use Cases ---------------------------------------- Reasoner Text, vision Text World understanding, grounding, physical reasoning, task planning, action forecasting, embodied agent reasoning, and autonomous system decision making Generator Text, vision, sound, action Vision, sound, action World generation, world simulation, future prediction, synthetic data generation, policy learning, and robot training World understanding: Analyze videos and images for captions, temporal events, next actions, spatial grounding, physical plausibility, and causal outcomes. World generation: Produce images, videos, synchronized sound, and action-conditioned rollouts from text, image, video, or action inputs. Action modeling: Predict policy actions, inverse dynamics, and forward dynamics for robotics, camera motion, egocentric motion, and autonomous-driving settings. Research and production paths: Use Diffusers and Transformers for Python-first development, vLLM-Omni, vLLM, TensorRT-LLM, or SGLang for OpenAI-compatible serving, and NIM containers for turnkey Reasoner serving or Generator deployment for text-to-video and image-to-video generation. Post-training recipes: Adapt vision, action, and reasoner workflows with Cosmos Framework training recipes and task-specific evaluation [Coming Soon].

View Cosmos on GitHub
Navid Moazzezby Navid Moazzez·Updated Sept 30, 2026·2 min read
Cosmos

Website Framework Agent Skills Models Introduction Cosmos 3 Key Capabilities Model Architecture Model Family Supported Generation Settings Input and Output Use Cases Generator Reasoner Quickstart Generator with Diffusers Generator with vLLM-Omni Generator with NIM Generator with SGLang Reasoner with Transformers Reasoner with vLLM Reasoner with TensorRT-LLM Reasoner with NIM Troubleshooting Which CUDA version should I use? Which base container should I use? torch.cuda.isavailable() is False Import fails with libxcb.so.1: cannot open shared object file uv errors on install or sync Choosing an Integration Examples Inference Benchmarks Finetune Export and Convert Checkpoints Distill Limitations Ecosystem News License and Contact

NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.

Surface Inputs Outputs Use Cases ---------------------------------------- Reasoner Text, vision Text World understanding, grounding, physical reasoning, task planning, action forecasting, embodied agent reasoning, and autonomous system decision making Generator Text, vision, sound, action Vision, sound, action World generation, world simulation, future prediction, synthetic data generation, policy learning, and robot training World understanding: Analyze videos and images for captions, temporal events, next actions, spatial grounding, physical plausibility, and causal outcomes. World generation: Produce images, videos, synchronized sound, and action-conditioned rollouts from text, image, video, or action inputs. Action modeling: Predict policy actions, inverse dynamics, and forward dynamics for robotics, camera motion, egocentric motion, and autonomous-driving settings. Research and production paths: Use Diffusers and Transformers for Python-first development, vLLM-Omni, vLLM, TensorRT-LLM, or SGLang for OpenAI-compatible serving, and NIM containers for turnkey Reasoner serving or Generator deployment for text-to-video and image-to-video generation. Post-training recipes: Adapt vision, action, and reasoner workflows with Cosmos Framework training recipes and task-specific evaluation [Coming Soon].

Cosmos at a glance

Stars12k
Forks892
LanguageJupyter Notebook
LicenseOther
Last update2026-09-23
Contributors47

How to install Cosmos

bash uvx hf@latest auth login 

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers