<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Video Editing | Yufei Liu's Homepages</title><link>https://liu-yufei.github.io/tags/video-editing/</link><atom:link href="https://liu-yufei.github.io/tags/video-editing/index.xml" rel="self" type="application/rss+xml"/><description>Video Editing</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Wed, 01 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>https://liu-yufei.github.io/media/icon_hu17033853472409137288.png</url><title>Video Editing</title><link>https://liu-yufei.github.io/tags/video-editing/</link></image><item><title>OmniPersona: Holistic Identity-Preserving Video Creation and Editing via Spatiotemporal Decoupled Persona Injection in Unified DiT</title><link>https://liu-yufei.github.io/publication/omnipersona/</link><pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate><guid>https://liu-yufei.github.io/publication/omnipersona/</guid><description>&lt;p>&lt;strong>Role: Co-first Author (Equal Contribution)&lt;/strong>&lt;/p>
&lt;style>
.omni-visuals { margin: 2.75rem auto 3rem; max-width: 1180px; }
.omni-visuals h2 { margin: 0 0 1.25rem; font-size: 1.75rem; }
.omni-overview { display: block; width: 100%; height: auto; }
.omni-example-title { margin-top: 3.5rem !important; }
.omni-example { margin: 0; }
.omni-example video { display: block; width: 100%; height: auto; border-radius: 0.85rem; background: #fff; box-shadow: 0 1rem 2.5rem rgba(15, 23, 42, 0.14); }
.omni-example figcaption { max-width: 58rem; margin: 0.85rem auto 0; color: #64748b; text-align: center; font-size: 0.9rem; line-height: 1.55; }
@media (max-width: 720px) { .omni-example figcaption { text-align: left; font-size: 0.8rem; } }
&lt;/style>
&lt;section class="omni-visuals" aria-labelledby="omni-overview-title">
&lt;h2 id="omni-overview-title">Overview&lt;/h2>
&lt;img class="omni-overview" src="https://liu-yufei.github.io/publication/omnipersona/overview.jpg" alt="OmniPersona framework overview">
&lt;h2 class="omni-example-title">Generation Example&lt;/h2>
&lt;figure class="omni-example">
&lt;video autoplay muted loop playsinline preload="metadata" poster="/publication/omnipersona/generation-example-poster.jpg" aria-label="OmniPersona generation example with static input conditions on the left and the generated video on the right">
&lt;source src="https://liu-yufei.github.io/publication/omnipersona/generation-example.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>The static input conditions and the animated output are combined into one synchronized video, following the presentation layout.&lt;/figcaption>
&lt;/figure>
&lt;/section>
&lt;p>OmniPersona introduces Multimodal Holistic Persona Representation and spatiotemporally decoupled persona injection for unified video creation and editing. I co-led the work and contributed to the holistic identity representation module, combining face identity features with DINOv2 and CLIP features for clothing, hairstyle, and body characteristics.&lt;/p>
&lt;p>On reference-to-video generation, OmniPersona reaches 79.3% face similarity and 68.7% body similarity, improving over the strongest baseline by 5.2 and 34.2 percentage points, respectively.&lt;/p></description></item><item><title>Wan-VACE-based Video Object and Shadow Removal</title><link>https://liu-yufei.github.io/project/wan-vace-object-removal/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://liu-yufei.github.io/project/wan-vace-object-removal/</guid><description>&lt;p>This Huawei collaboration project addresses the shortage of high-quality paired data for video object removal, especially when shadows and reflections must also be removed.&lt;/p>
&lt;p>I independently built a data generation and fine-tuning pipeline using Wan2.2-5B, SAM2, and a baseline removal model. From roughly 4,000 generated groups, I selected 1,070 high-quality training pairs and retained 2,500 difficult cases for evaluation. I then fine-tuned the VACE module of Wan2.1-VACE-1.3B with DiffSynth-Studio, Accelerate, and DeepSpeed on six NVIDIA A800 GPUs.&lt;/p>
&lt;style>
.wan-comparisons { margin: 3.25rem auto; max-width: 1120px; }
.wan-comparisons h2 { margin: 0 0 1.5rem; font-size: 1.75rem; }
.wan-comparison { --split: 50%; position: relative; width: 100%; aspect-ratio: 40 / 23; overflow: hidden; border-radius: 0.8rem; background: #111; }
.wan-comparison + h2 { margin-top: 3.5rem; }
.wan-comparison video { position: absolute; inset: 0; width: 100%; height: 100%; object-fit: contain; }
.wan-compare-after { clip-path: inset(0 0 0 var(--split)); }
.wan-comparison-label { position: absolute; top: 1rem; z-index: 4; padding: 0.35rem 0.7rem; border-radius: 999px; color: #fff; background: rgba(0, 0, 0, 0.62); font-size: 0.85rem; font-weight: 700; backdrop-filter: blur(6px); }
.wan-comparison-label.left { left: 1rem; }
.wan-comparison-label.right { right: 1rem; }
.wan-comparison-line { position: absolute; z-index: 3; top: 0; bottom: 0; left: var(--split); width: 2px; background: #fff; transform: translateX(-1px); pointer-events: none; }
.wan-comparison-handle { position: absolute; z-index: 4; top: 50%; left: var(--split); width: 2.65rem; height: 2.65rem; border: 2px solid #fff; border-radius: 50%; color: #fff; background: rgba(0, 0, 0, 0.55); transform: translate(-50%, -50%); display: grid; place-items: center; font-size: 1rem; pointer-events: none; }
.wan-comparison-range { position: absolute; z-index: 5; inset: 0; width: 100%; height: 100%; margin: 0; opacity: 0; cursor: ew-resize; }
.wan-comparison-note { margin: 0.75rem 0 0; text-align: center; font-size: 0.88rem; opacity: 0.68; }
@media (max-width: 700px) {
.wan-comparison-label { top: 0.6rem; padding: 0.25rem 0.55rem; font-size: 0.72rem; }
.wan-comparison-label.left { left: 0.6rem; }
.wan-comparison-label.right { right: 0.6rem; }
.wan-comparison-handle { width: 2.2rem; height: 2.2rem; }
}
&lt;/style>
&lt;section class="wan-comparisons" aria-labelledby="wan-comparison-title">
&lt;p class="wan-comparison-note">All videos are displayed at 480 × 276, matching the output resolution of our model.&lt;/p>
&lt;h2 id="wan-comparison-title">Condition Video vs. Baseline&lt;/h2>
&lt;div class="wan-comparison" data-video-compare data-video-sync data-sync-duration="3.0625">
&lt;video data-sync-master muted playsinline preload="auto" poster="/project/wan-vace-object-removal/condition-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/condition.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;video class="wan-compare-after" muted playsinline preload="auto" poster="/project/wan-vace-object-removal/baseline-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/baseline.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;span class="wan-comparison-label left">Condition Video&lt;/span>
&lt;span class="wan-comparison-label right">Baseline&lt;/span>
&lt;span class="wan-comparison-line">&lt;/span>&lt;span class="wan-comparison-handle">↔&lt;/span>
&lt;input class="wan-comparison-range" type="range" min="0" max="100" value="50" aria-label="Compare the condition video with the baseline result">
&lt;/div>
&lt;p class="wan-comparison-note">Drag the slider to compare both videos frame by frame.&lt;/p>
&lt;h2>Condition Video vs. Ours&lt;/h2>
&lt;div class="wan-comparison" data-video-compare data-video-sync data-sync-duration="3.0625">
&lt;video data-sync-master muted playsinline preload="auto" poster="/project/wan-vace-object-removal/condition-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/condition.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;video class="wan-compare-after" muted playsinline preload="auto" poster="/project/wan-vace-object-removal/ours-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/ours.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;span class="wan-comparison-label left">Condition Video&lt;/span>
&lt;span class="wan-comparison-label right">Ours&lt;/span>
&lt;span class="wan-comparison-line">&lt;/span>&lt;span class="wan-comparison-handle">↔&lt;/span>
&lt;input class="wan-comparison-range" type="range" min="0" max="100" value="50" aria-label="Compare the condition video with our result">
&lt;/div>
&lt;p class="wan-comparison-note">Drag the slider to inspect object and shadow removal.&lt;/p>
&lt;/section>
&lt;p>Human evaluation found 2,446 of the 2,500 difficult cases usable, a 97.84% success rate, with improved object-shadow removal and background reconstruction.&lt;/p></description></item></channel></rss>