Hi team,
First of all, thank you for releasing this amazing project and making the 16B weights available!
However, during the installation and first run on a high-end Windows environment (RTX 4090, 24GB VRAM, 96GB System RAM, CUDA 12.4), I encountered several critical bugs related to hardcoded paths, as well as a severe memory bottleneck that prevents the model from generating anything.
Here is a detailed list of the issues and feedback for your next commit:
-
Severe VRAM Bottleneck on RTX 4090 (Request for Quantization)
Running the model in its default bfloat16 precision is practically impossible on a 24GB VRAM consumer GPU. The model loads the weights successfully, but once the diffusion process starts, it hangs indefinitely at 0/30 steps. It appears the 24GB VRAM limit is exceeded during the forward pass, causing the NVIDIA driver to silently swap to shared system RAM, slowing the generation to an absolute crawl (effectively frozen).
Request: Are there any plans to release official quantized versions of the model (e.g., 8-bit, 4-bit NF4/AWQ, or GGUF)? This would make the project actually usable for the open-source community with 24GB cards.
-
Incorrect versions in requirements.txt
diffusers==0.36.2 does not exist on PyPI. Changing it to the latest stable version (diffusers==0.37.1) fixes the installation.
transformers==4.54.0 also causes conflicts, especially since the code uses Qwen3. I had to install the dev version directly from the main branch (pip install git+https://github.com/huggingface/transformers.git --upgrade) to avoid an ImportError for the Qwen3 model.
- Deprecated Class in transformers
In the file src/models/mmdit/text_encoder/init.py (line 7), the code attempts to import AutoModelForVision2Seq. This class has been deprecated and renamed in recent transformers updates.
Fix: Change AutoModelForVision2Seq to AutoModelForImageTextToText.
-
Hardcoded absolute paths in configs/spatialedit_base_config.py
The base config file contains placeholder paths (your_base_path/model/...) for the VAE and the Text Encoder (Qwen3). This causes a FileNotFoundError during load_pipeline.
-
Hardcoded absolute path inside the source code (wanvae.py)
Even after fixing the configs file, the script still crashes with:
FileNotFoundError: [Errno 2] No such file or directory: '/pfs/yichengxiao/huggingface/model/Wan2.1-T2V-1.3B/Wan2.1_VAE.pth'
Bug location: In src/models/mmdit/vae/wanvae.py (around line 630 in the init function of WanxVAE), the pretrained argument has a hardcoded absolute path to a local server (/pfs/yichengxiao/...).
Fix: This default argument should be removed or changed to a relative path, so it correctly inherits the path defined in the config file.
- Missing VAE and Qwen downloads in the README
The README instructs users to download SpatialEdit-16B, VGGT, and YOLO26x, but misses the clear instruction to download Wan2.1_VAE.pth and Qwen3-VL-8B-Instruct. Without these two, the spatialedit_demo.py script cannot run.
I hope this feedback helps in cleaning up the repo and making it accessible for users outside your lab! Looking forward to your thoughts on releasing quantized weights.
Hi team,
First of all, thank you for releasing this amazing project and making the 16B weights available!
However, during the installation and first run on a high-end Windows environment (RTX 4090, 24GB VRAM, 96GB System RAM, CUDA 12.4), I encountered several critical bugs related to hardcoded paths, as well as a severe memory bottleneck that prevents the model from generating anything.
Here is a detailed list of the issues and feedback for your next commit:
Severe VRAM Bottleneck on RTX 4090 (Request for Quantization)
Running the model in its default bfloat16 precision is practically impossible on a 24GB VRAM consumer GPU. The model loads the weights successfully, but once the diffusion process starts, it hangs indefinitely at 0/30 steps. It appears the 24GB VRAM limit is exceeded during the forward pass, causing the NVIDIA driver to silently swap to shared system RAM, slowing the generation to an absolute crawl (effectively frozen).
Request: Are there any plans to release official quantized versions of the model (e.g., 8-bit, 4-bit NF4/AWQ, or GGUF)? This would make the project actually usable for the open-source community with 24GB cards.
Incorrect versions in requirements.txt
diffusers==0.36.2 does not exist on PyPI. Changing it to the latest stable version (diffusers==0.37.1) fixes the installation.
transformers==4.54.0 also causes conflicts, especially since the code uses Qwen3. I had to install the dev version directly from the main branch (pip install git+https://github.com/huggingface/transformers.git --upgrade) to avoid an ImportError for the Qwen3 model.
In the file src/models/mmdit/text_encoder/init.py (line 7), the code attempts to import AutoModelForVision2Seq. This class has been deprecated and renamed in recent transformers updates.
Fix: Change AutoModelForVision2Seq to AutoModelForImageTextToText.
Hardcoded absolute paths in configs/spatialedit_base_config.py
The base config file contains placeholder paths (your_base_path/model/...) for the VAE and the Text Encoder (Qwen3). This causes a FileNotFoundError during load_pipeline.
Hardcoded absolute path inside the source code (wanvae.py)
Even after fixing the configs file, the script still crashes with:
FileNotFoundError: [Errno 2] No such file or directory: '/pfs/yichengxiao/huggingface/model/Wan2.1-T2V-1.3B/Wan2.1_VAE.pth'
Bug location: In src/models/mmdit/vae/wanvae.py (around line 630 in the init function of WanxVAE), the pretrained argument has a hardcoded absolute path to a local server (/pfs/yichengxiao/...).
Fix: This default argument should be removed or changed to a relative path, so it correctly inherits the path defined in the config file.
The README instructs users to download SpatialEdit-16B, VGGT, and YOLO26x, but misses the clear instruction to download Wan2.1_VAE.pth and Qwen3-VL-8B-Instruct. Without these two, the spatialedit_demo.py script cannot run.
I hope this feedback helps in cleaning up the repo and making it accessible for users outside your lab! Looking forward to your thoughts on releasing quantized weights.