<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AIGC | Yufei Liu's Homepages</title><link>https://liu-yufei.github.io/tags/aigc/</link><atom:link href="https://liu-yufei.github.io/tags/aigc/index.xml" rel="self" type="application/rss+xml"/><description>AIGC</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://liu-yufei.github.io/media/icon_hu17033853472409137288.png</url><title>AIGC</title><link>https://liu-yufei.github.io/tags/aigc/</link></image><item><title>FLUX.2-based Image Identity Consistency Enhancement</title><link>https://liu-yufei.github.io/project/flux2-identity-consistency/</link><pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate><guid>https://liu-yufei.github.io/project/flux2-identity-consistency/</guid><description>&lt;p>Developed during an algorithm internship with the face team at Meitu MT-Lab, this project addresses identity drift in AIGC portrait editing. The module takes an edited portrait and a user reference image as inputs, then refines identity while retaining the intended edit.&lt;/p>
&lt;p>I built the data pipeline for face detection, identity clustering, mask generation, canonical alignment, reference matching, and quality filtering. I also designed joint perceptual, identity, and contour objectives for multi-step FLUX.2 LoRA training and developed evaluation tools for identity similarity, facial contour change, and denoising trajectories.&lt;/p>
&lt;p>The resulting model improved reference identity similarity across internal test sets and was deployed in the telephoto portrait feature of AI Camera.&lt;/p>
&lt;style>
.flux-results { margin: 3rem auto 1rem; max-width: 1120px; }
.flux-results h2 { margin: 0 0 1.5rem; font-size: 1.75rem; }
.flux-case { padding: 1.75rem 0; border-bottom: 1px solid rgba(128, 128, 128, 0.28); }
.flux-case:last-child { border-bottom: 0; }
.flux-case-head { display: flex; align-items: baseline; justify-content: space-between; gap: 1rem; margin-bottom: 1rem; }
.flux-case-title { margin: 0; font-size: 1.15rem; font-weight: 700; }
.flux-gain { color: #13866f; font-size: 1.05rem; font-weight: 700; white-space: nowrap; }
.flux-images { display: grid; grid-template-columns: repeat(3, minmax(0, 1fr)); gap: 1rem; }
.flux-image img { width: 100%; aspect-ratio: 3 / 4; object-fit: cover; object-position: center top; border-radius: 0.65rem; }
.flux-image-label { display: block; margin-top: 0.45rem; text-align: center; font-size: 0.92rem; font-weight: 650; }
.flux-sim { margin-top: 0.9rem; text-align: right; font-variant-numeric: tabular-nums; opacity: 0.78; }
.flux-summary { margin-top: 4rem; }
.flux-chart { display: grid; grid-template-columns: repeat(3, minmax(0, 1fr)); gap: 2rem; align-items: end; margin-top: 2rem; }
.flux-chart-group { text-align: center; }
.flux-bars { height: 230px; display: flex; align-items: end; justify-content: center; gap: 0.8rem; border-bottom: 1px solid rgba(128, 128, 128, 0.4); }
.flux-bar { position: relative; width: min(36%, 72px); min-height: 2rem; border-radius: 0.45rem 0.45rem 0 0; }
.flux-bar.before { background: #8f278c; }
.flux-bar.after { background: #bc84b7; }
.flux-bar-value { position: absolute; top: -1.75rem; left: 50%; transform: translateX(-50%); font-weight: 700; font-variant-numeric: tabular-nums; }
.flux-group-name { margin: 0.8rem 0 0.25rem; font-weight: 700; }
.flux-group-gain { color: #13866f; font-size: 1.35rem; font-weight: 750; }
.flux-group-meta { margin-top: 0.35rem; font-size: 0.85rem; line-height: 1.55; opacity: 0.72; }
.flux-legend { display: flex; justify-content: center; gap: 1.5rem; margin-top: 1.5rem; font-size: 0.9rem; }
.flux-legend span::before { content: ""; display: inline-block; width: 0.8rem; height: 0.8rem; margin-right: 0.4rem; border-radius: 0.18rem; vertical-align: -0.05rem; }
.flux-legend .before::before { background: #8f278c; }
.flux-legend .after::before { background: #bc84b7; }
.flux-conclusion { margin-top: 1.75rem; padding-top: 1.25rem; border-top: 1px solid rgba(128, 128, 128, 0.28); line-height: 1.8; }
@media (max-width: 700px) {
.flux-case-head { display: block; }
.flux-gain { display: block; margin-top: 0.35rem; }
.flux-images { gap: 0.5rem; }
.flux-image-label { font-size: 0.78rem; }
.flux-chart { grid-template-columns: 1fr; gap: 3rem; }
.flux-bars { height: 180px; gap: 0.35rem; }
.flux-bar { width: 38%; }
.flux-bar-value { font-size: 0.78rem; }
.flux-group-meta { font-size: 0.85rem; }
}
&lt;/style>
&lt;section class="flux-results" aria-labelledby="flux-qualitative-title">
&lt;h2 id="flux-qualitative-title">Qualitative Results&lt;/h2>
&lt;article class="flux-case">
&lt;div class="flux-case-head">
&lt;h3 class="flux-case-title">Frontal Face Reshaping × JJ Lin&lt;/h3>
&lt;span class="flux-gain">Identity Similarity &amp;#43;0.186&lt;/span>
&lt;/div>
&lt;div class="flux-images">
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case1-input_hu16733012294050469293.webp" alt="Input image for Frontal Face Reshaping × JJ Lin">&lt;figcaption class="flux-image-label">Input&lt;/figcaption>&lt;/figure>
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case1-reference_hu4679323974125261232.webp" alt="Reference image for Frontal Face Reshaping × JJ Lin">&lt;figcaption class="flux-image-label">Reference&lt;/figcaption>&lt;/figure>
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case1-output_hu3685900372443741842.webp" alt="Refined output for Frontal Face Reshaping × JJ Lin">&lt;figcaption class="flux-image-label">Output&lt;/figcaption>&lt;/figure>
&lt;/div>
&lt;div class="flux-sim">Input–Reference 0.410　→　Output–Reference 0.596&lt;/div>
&lt;/article>
&lt;article class="flux-case">
&lt;div class="flux-case-head">
&lt;h3 class="flux-case-title">Half-Profile Face Reshaping × Jackson Yee&lt;/h3>
&lt;span class="flux-gain">Identity Similarity &amp;#43;0.298&lt;/span>
&lt;/div>
&lt;div class="flux-images">
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case2-input_hu12950896630273560113.webp" alt="Input image for Half-Profile Face Reshaping × Jackson Yee">&lt;figcaption class="flux-image-label">Input&lt;/figcaption>&lt;/figure>
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case2-reference_hu10496544986031870863.webp" alt="Reference image for Half-Profile Face Reshaping × Jackson Yee">&lt;figcaption class="flux-image-label">Reference&lt;/figcaption>&lt;/figure>
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case2-output_hu11499135139916125719.webp" alt="Refined output for Half-Profile Face Reshaping × Jackson Yee">&lt;figcaption class="flux-image-label">Output&lt;/figcaption>&lt;/figure>
&lt;/div>
&lt;div class="flux-sim">Input–Reference 0.561　→　Output–Reference 0.859&lt;/div>
&lt;/article>
&lt;article class="flux-case">
&lt;div class="flux-case-head">
&lt;h3 class="flux-case-title">Strong Stage Lighting × G-Dragon&lt;/h3>
&lt;span class="flux-gain">Identity Similarity &amp;#43;0.389&lt;/span>
&lt;/div>
&lt;div class="flux-images">
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case3-input_hu14981943814846579746.webp" alt="Input image for Strong Stage Lighting × G-Dragon">&lt;figcaption class="flux-image-label">Input&lt;/figcaption>&lt;/figure>
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case3-reference_hu9276370965852099679.webp" alt="Reference image for Strong Stage Lighting × G-Dragon">&lt;figcaption class="flux-image-label">Reference&lt;/figcaption>&lt;/figure>
&lt;figure class="flux-image">&lt;img src="https://liu-yufei.github.io/project/flux2-identity-consistency/case3-output_hu10029804751481071616.webp" alt="Refined output for Strong Stage Lighting × G-Dragon">&lt;figcaption class="flux-image-label">Output&lt;/figcaption>&lt;/figure>
&lt;/div>
&lt;div class="flux-sim">Input–Reference 0.286　→　Output–Reference 0.675&lt;/div>
&lt;/article>
&lt;/section>
&lt;section class="flux-results flux-summary" aria-labelledby="flux-quantitative-title">
&lt;h2 id="flux-quantitative-title">Overall Evaluation&lt;/h2>
&lt;div class="flux-chart" role="img" aria-label="Average identity similarity before and after refinement on three test sets">
&lt;div class="flux-chart-group">
&lt;div class="flux-bars">&lt;div class="flux-bar before" style="height:50.75%">&lt;span class="flux-bar-value">0.406&lt;/span>&lt;/div>&lt;div class="flux-bar after" style="height:85%">&lt;span class="flux-bar-value">0.680&lt;/span>&lt;/div>&lt;/div>
&lt;div class="flux-group-name">Celeb-new&lt;/div>&lt;div class="flux-group-gain">+0.274&lt;/div>
&lt;div class="flux-group-meta">115 cases, 100.0% improved&lt;br>Median pose gap 1.8°, non-face MAE 0.060&lt;/div>
&lt;/div>
&lt;div class="flux-chart-group">
&lt;div class="flux-bars">&lt;div class="flux-bar before" style="height:34%">&lt;span class="flux-bar-value">0.272&lt;/span>&lt;/div>&lt;div class="flux-bar after" style="height:75%">&lt;span class="flux-bar-value">0.600&lt;/span>&lt;/div>&lt;/div>
&lt;div class="flux-group-name">Celeb-old&lt;/div>&lt;div class="flux-group-gain">+0.328&lt;/div>
&lt;div class="flux-group-meta">28 cases, 96.4% improved&lt;br>Median pose gap 1.4°, non-face MAE 0.043&lt;/div>
&lt;/div>
&lt;div class="flux-chart-group">
&lt;div class="flux-bars">&lt;div class="flux-bar before" style="height:42.25%">&lt;span class="flux-bar-value">0.338&lt;/span>&lt;/div>&lt;div class="flux-bar after" style="height:90.13%">&lt;span class="flux-bar-value">0.721&lt;/span>&lt;/div>&lt;/div>
&lt;div class="flux-group-name">Colleague&lt;/div>&lt;div class="flux-group-gain">+0.383&lt;/div>
&lt;div class="flux-group-meta">440 cases, 99.8% improved&lt;br>Median pose gap 3.1°, non-face MAE 0.227&lt;/div>
&lt;/div>
&lt;/div>
&lt;div class="flux-legend">&lt;span class="before">Input → Reference&lt;/span>&lt;span class="after">Output → Reference&lt;/span>&lt;/div>
&lt;p class="flux-conclusion">Identity similarity improves consistently across all three test sets, while pose changes and non-face-region differences remain limited.&lt;/p>
&lt;/section></description></item><item><title>OmniPersona: Holistic Identity-Preserving Video Creation and Editing via Spatiotemporal Decoupled Persona Injection in Unified DiT</title><link>https://liu-yufei.github.io/publication/omnipersona/</link><pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate><guid>https://liu-yufei.github.io/publication/omnipersona/</guid><description>&lt;p>&lt;strong>Role: Co-first Author (Equal Contribution)&lt;/strong>&lt;/p>
&lt;style>
.omni-visuals { margin: 2.75rem auto 3rem; max-width: 1180px; }
.omni-visuals h2 { margin: 0 0 1.25rem; font-size: 1.75rem; }
.omni-overview { display: block; width: 100%; height: auto; }
.omni-example-title { margin-top: 3.5rem !important; }
.omni-example { margin: 0; }
.omni-example video { display: block; width: 100%; height: auto; border-radius: 0.85rem; background: #fff; box-shadow: 0 1rem 2.5rem rgba(15, 23, 42, 0.14); }
.omni-example figcaption { max-width: 58rem; margin: 0.85rem auto 0; color: #64748b; text-align: center; font-size: 0.9rem; line-height: 1.55; }
@media (max-width: 720px) { .omni-example figcaption { text-align: left; font-size: 0.8rem; } }
&lt;/style>
&lt;section class="omni-visuals" aria-labelledby="omni-overview-title">
&lt;h2 id="omni-overview-title">Overview&lt;/h2>
&lt;img class="omni-overview" src="https://liu-yufei.github.io/publication/omnipersona/overview.jpg" alt="OmniPersona framework overview">
&lt;h2 class="omni-example-title">Generation Example&lt;/h2>
&lt;figure class="omni-example">
&lt;video autoplay muted loop playsinline preload="metadata" poster="/publication/omnipersona/generation-example-poster.jpg" aria-label="OmniPersona generation example with static input conditions on the left and the generated video on the right">
&lt;source src="https://liu-yufei.github.io/publication/omnipersona/generation-example.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>The static input conditions and the animated output are combined into one synchronized video, following the presentation layout.&lt;/figcaption>
&lt;/figure>
&lt;/section>
&lt;p>OmniPersona introduces Multimodal Holistic Persona Representation and spatiotemporally decoupled persona injection for unified video creation and editing. I co-led the work and contributed to the holistic identity representation module, combining face identity features with DINOv2 and CLIP features for clothing, hairstyle, and body characteristics.&lt;/p>
&lt;p>On reference-to-video generation, OmniPersona reaches 79.3% face similarity and 68.7% body similarity, improving over the strongest baseline by 5.2 and 34.2 percentage points, respectively.&lt;/p></description></item><item><title>Wan-VACE-based Video Object and Shadow Removal</title><link>https://liu-yufei.github.io/project/wan-vace-object-removal/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://liu-yufei.github.io/project/wan-vace-object-removal/</guid><description>&lt;p>This Huawei collaboration project addresses the shortage of high-quality paired data for video object removal, especially when shadows and reflections must also be removed.&lt;/p>
&lt;p>I independently built a data generation and fine-tuning pipeline using Wan2.2-5B, SAM2, and a baseline removal model. From roughly 4,000 generated groups, I selected 1,070 high-quality training pairs and retained 2,500 difficult cases for evaluation. I then fine-tuned the VACE module of Wan2.1-VACE-1.3B with DiffSynth-Studio, Accelerate, and DeepSpeed on six NVIDIA A800 GPUs.&lt;/p>
&lt;style>
.wan-comparisons { margin: 3.25rem auto; max-width: 1120px; }
.wan-comparisons h2 { margin: 0 0 1.5rem; font-size: 1.75rem; }
.wan-comparison { --split: 50%; position: relative; width: 100%; aspect-ratio: 40 / 23; overflow: hidden; border-radius: 0.8rem; background: #111; }
.wan-comparison + h2 { margin-top: 3.5rem; }
.wan-comparison video { position: absolute; inset: 0; width: 100%; height: 100%; object-fit: contain; }
.wan-compare-after { clip-path: inset(0 0 0 var(--split)); }
.wan-comparison-label { position: absolute; top: 1rem; z-index: 4; padding: 0.35rem 0.7rem; border-radius: 999px; color: #fff; background: rgba(0, 0, 0, 0.62); font-size: 0.85rem; font-weight: 700; backdrop-filter: blur(6px); }
.wan-comparison-label.left { left: 1rem; }
.wan-comparison-label.right { right: 1rem; }
.wan-comparison-line { position: absolute; z-index: 3; top: 0; bottom: 0; left: var(--split); width: 2px; background: #fff; transform: translateX(-1px); pointer-events: none; }
.wan-comparison-handle { position: absolute; z-index: 4; top: 50%; left: var(--split); width: 2.65rem; height: 2.65rem; border: 2px solid #fff; border-radius: 50%; color: #fff; background: rgba(0, 0, 0, 0.55); transform: translate(-50%, -50%); display: grid; place-items: center; font-size: 1rem; pointer-events: none; }
.wan-comparison-range { position: absolute; z-index: 5; inset: 0; width: 100%; height: 100%; margin: 0; opacity: 0; cursor: ew-resize; }
.wan-comparison-note { margin: 0.75rem 0 0; text-align: center; font-size: 0.88rem; opacity: 0.68; }
@media (max-width: 700px) {
.wan-comparison-label { top: 0.6rem; padding: 0.25rem 0.55rem; font-size: 0.72rem; }
.wan-comparison-label.left { left: 0.6rem; }
.wan-comparison-label.right { right: 0.6rem; }
.wan-comparison-handle { width: 2.2rem; height: 2.2rem; }
}
&lt;/style>
&lt;section class="wan-comparisons" aria-labelledby="wan-comparison-title">
&lt;p class="wan-comparison-note">All videos are displayed at 480 × 276, matching the output resolution of our model.&lt;/p>
&lt;h2 id="wan-comparison-title">Condition Video vs. Baseline&lt;/h2>
&lt;div class="wan-comparison" data-video-compare data-video-sync data-sync-duration="3.0625">
&lt;video data-sync-master muted playsinline preload="auto" poster="/project/wan-vace-object-removal/condition-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/condition.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;video class="wan-compare-after" muted playsinline preload="auto" poster="/project/wan-vace-object-removal/baseline-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/baseline.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;span class="wan-comparison-label left">Condition Video&lt;/span>
&lt;span class="wan-comparison-label right">Baseline&lt;/span>
&lt;span class="wan-comparison-line">&lt;/span>&lt;span class="wan-comparison-handle">↔&lt;/span>
&lt;input class="wan-comparison-range" type="range" min="0" max="100" value="50" aria-label="Compare the condition video with the baseline result">
&lt;/div>
&lt;p class="wan-comparison-note">Drag the slider to compare both videos frame by frame.&lt;/p>
&lt;h2>Condition Video vs. Ours&lt;/h2>
&lt;div class="wan-comparison" data-video-compare data-video-sync data-sync-duration="3.0625">
&lt;video data-sync-master muted playsinline preload="auto" poster="/project/wan-vace-object-removal/condition-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/condition.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;video class="wan-compare-after" muted playsinline preload="auto" poster="/project/wan-vace-object-removal/ours-poster.png?v=resolution-matched-20261007">&lt;source src="https://liu-yufei.github.io/project/wan-vace-object-removal/ours.mp4?v=resolution-matched-20261007" type="video/mp4">&lt;/video>
&lt;span class="wan-comparison-label left">Condition Video&lt;/span>
&lt;span class="wan-comparison-label right">Ours&lt;/span>
&lt;span class="wan-comparison-line">&lt;/span>&lt;span class="wan-comparison-handle">↔&lt;/span>
&lt;input class="wan-comparison-range" type="range" min="0" max="100" value="50" aria-label="Compare the condition video with our result">
&lt;/div>
&lt;p class="wan-comparison-note">Drag the slider to inspect object and shadow removal.&lt;/p>
&lt;/section>
&lt;p>Human evaluation found 2,446 of the 2,500 difficult cases usable, a 97.84% success rate, with improved object-shadow removal and background reconstruction.&lt;/p></description></item></channel></rss>