Hi, and thanks for releasing CMD.
I built an unofficial ComfyUI adapter that calls an official nv-tlabs/cmd checkout (the official code is not vendored, and weights are not redistributed):
https://github.com/hiroki-abe-58/ComfyUI-NVIDIA-CMD
While getting it to run, I had to patch a few things at runtime. I'm reporting them here in case they're useful upstream, and I'm happy to send small PRs if you're open to them.
Tested environment
- Windows 11 Native (no WSL), RTX 5090 32GB (sm_120)
- PyTorch 2.9.1+cu130, no
flash-attn installed
- Checkpoints:
chunk1_short, chunk1_long, chunk1_camera_control
1. Attention without flash-attn
I replaced the attention call in cosmos.runtime.attention with torch.nn.functional.scaled_dot_product_attention, and inference ran fine on Blackwell.
Proposal: fall back to SDPA when flash_attn cannot be imported. That would make CMD usable on Windows and on setups where building flash-attn is difficult.
2. KV cache memory in long rollout
With chunk1_long (126 latent frames), the KV cache keeps entries for every generated frame and runs out of memory on a 32GB card. Capping the cache to local_attn_size (21) as a circular buffer gives a measured peak of 23,852 MiB and about 267 s per run.
Proposal: an option to bound the cache to local_attn_size. I haven't compared the outputs against the uncapped cache, since the uncapped run doesn't fit on my GPU. Could you confirm whether capping at local_attn_size is expected to be lossless, given that attention only looks back that far?
3. Importing CMD as a library
The top-level module names (utils, pipeline, wan, inference) collide easily with other packages once CMD is imported from another application. Right now I work around this with sys.path / sys.modules manipulation and by adding __init__.py files.
Question: would you consider namespacing these under a single package? This is the larger change of the three, so I'd only attempt it with your guidance.
I can open PRs for (1) and (2) as separate, minimal changes. Please let me know if there is a contribution process I should follow (CLA, DCO sign-off, etc.).
The adapter README clearly states that it is unofficial, and that the CMD code and weights remain under the NVIDIA OneWay Noncommercial License. If the "NVIDIA" in the repository name is a concern, I'm happy to rename it.
Hi, and thanks for releasing CMD.
I built an unofficial ComfyUI adapter that calls an official
nv-tlabs/cmdcheckout (the official code is not vendored, and weights are not redistributed):https://github.com/hiroki-abe-58/ComfyUI-NVIDIA-CMD
While getting it to run, I had to patch a few things at runtime. I'm reporting them here in case they're useful upstream, and I'm happy to send small PRs if you're open to them.
Tested environment
flash-attninstalledchunk1_short,chunk1_long,chunk1_camera_control1. Attention without flash-attn
I replaced the attention call in
cosmos.runtime.attentionwithtorch.nn.functional.scaled_dot_product_attention, and inference ran fine on Blackwell.Proposal: fall back to SDPA when
flash_attncannot be imported. That would make CMD usable on Windows and on setups where building flash-attn is difficult.2. KV cache memory in long rollout
With
chunk1_long(126 latent frames), the KV cache keeps entries for every generated frame and runs out of memory on a 32GB card. Capping the cache tolocal_attn_size(21) as a circular buffer gives a measured peak of 23,852 MiB and about 267 s per run.Proposal: an option to bound the cache to
local_attn_size. I haven't compared the outputs against the uncapped cache, since the uncapped run doesn't fit on my GPU. Could you confirm whether capping atlocal_attn_sizeis expected to be lossless, given that attention only looks back that far?3. Importing CMD as a library
The top-level module names (
utils,pipeline,wan,inference) collide easily with other packages once CMD is imported from another application. Right now I work around this withsys.path/sys.modulesmanipulation and by adding__init__.pyfiles.Question: would you consider namespacing these under a single package? This is the larger change of the three, so I'd only attempt it with your guidance.
I can open PRs for (1) and (2) as separate, minimal changes. Please let me know if there is a contribution process I should follow (CLA, DCO sign-off, etc.).
The adapter README clearly states that it is unofficial, and that the CMD code and weights remain under the NVIDIA OneWay Noncommercial License. If the "NVIDIA" in the repository name is a concern, I'm happy to rename it.