Wan-VACE-based Video Object and Shadow Removal

Jan 2026 · 1 min read

This Huawei collaboration project addresses the shortage of high-quality paired data for video object removal, especially when shadows and reflections must also be removed.

I independently built a data generation and fine-tuning pipeline using Wan2.2-5B, SAM2, and a baseline removal model. From roughly 4,000 generated groups, I selected 1,070 high-quality training pairs and retained 2,500 difficult cases for evaluation. I then fine-tuned the VACE module of Wan2.1-VACE-1.3B with DiffSynth-Studio, Accelerate, and DeepSpeed on six NVIDIA A800 GPUs.

All videos are displayed at 480 × 276, matching the output resolution of our model.

Condition Video vs. Baseline

Condition Video Baseline ↔

Drag the slider to compare both videos frame by frame.

Condition Video vs. Ours

Condition Video Ours ↔

Drag the slider to inspect object and shadow removal.

Human evaluation found 2,446 of the 2,500 difficult cases usable, a 97.84% success rate, with improved object-shadow removal and background reconstruction.