OmniPersona: Holistic Identity-Preserving Video Creation and Editing via Spatiotemporal Decoupled Persona Injection in Unified DiT

Jul 2026·
Zijie Meng
Yufei Liu
Yufei Liu
,
Bingcai Wei
,
Xixin Cao
· 1 min read
Type
Publication
In ACM Multimedia 2026 (MM ‘26)

Role: Co-first Author (Equal Contribution)

Overview

OmniPersona framework overview

Generation Example

The static input conditions and the animated output are combined into one synchronized video, following the presentation layout.

OmniPersona introduces Multimodal Holistic Persona Representation and spatiotemporally decoupled persona injection for unified video creation and editing. I co-led the work and contributed to the holistic identity representation module, combining face identity features with DINOv2 and CLIP features for clothing, hairstyle, and body characteristics.

On reference-to-video generation, OmniPersona reaches 79.3% face similarity and 68.7% body similarity, improving over the strongest baseline by 5.2 and 34.2 percentage points, respectively.