Wan-VACE-based Video Object and Shadow Removal
This Huawei collaboration project addresses the shortage of high-quality paired data for video object removal, especially when shadows and reflections must also be removed.
I independently built a data generation and fine-tuning pipeline using Wan2.2-5B, SAM2, and a baseline removal model. From roughly 4,000 generated groups, I selected 1,070 high-quality training pairs and retained 2,500 difficult cases for evaluation. I then fine-tuned the VACE module of Wan2.1-VACE-1.3B with DiffSynth-Studio, Accelerate, and DeepSpeed on six NVIDIA A800 GPUs.
All videos are displayed at 480 × 276, matching the output resolution of our model.
Condition Video vs. Baseline
Drag the slider to compare both videos frame by frame.
Condition Video vs. Ours
Drag the slider to inspect object and shadow removal.
Human evaluation found 2,446 of the 2,500 difficult cases usable, a 97.84% success rate, with improved object-shadow removal and background reconstruction.