OmniPersona: Holistic Identity-Preserving Video Creation and Editing via Spatiotemporal Decoupled Persona Injection in Unified DiT
Jul 2026·
,,·
1 min read
Zijie Meng
Yufei Liu
Bingcai Wei
Xixin Cao
Role: Co-first Author (Equal Contribution)
Overview

Generation Example
OmniPersona introduces Multimodal Holistic Persona Representation and spatiotemporally decoupled persona injection for unified video creation and editing. I co-led the work and contributed to the holistic identity representation module, combining face identity features with DINOv2 and CLIP features for clothing, hairstyle, and body characteristics.
On reference-to-video generation, OmniPersona reaches 79.3% face similarity and 68.7% body similarity, improving over the strongest baseline by 5.2 and 34.2 percentage points, respectively.