Vera: Identity-Faithful Human Subject-to-Video Generation
Summary
Vera is a new unified human-centric Subject-to-Video (S2V) generation framework designed for both single- and multi-person scenarios, addressing the critical issue of identity drift in human subjects across video frames, poses, and interactions. Existing S2V methods often struggle with maintaining identity-critical human details, leading to subject confusion, attribute swapping, and excessive copying of reference appearance cues, particularly in multi-person contexts. To overcome this, Vera leverages a newly constructed million-pair identity-aligned human image-video dataset, built through person-level cross-clip retrieval, which provides explicit identity correspondence. The framework integrates two complementary designs: Identity-Focal Masked Supervision (IFMS) for strengthened identity-aware learning with spatially focused supervision, and Reference-Aware Layer-wise Attention (RALA) to regulate how video tokens interact with reference identity cues within the DiT backbone. Experiments show Vera significantly improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.
Key takeaway
For Computer Vision Engineers developing human-centric video generation systems, Vera offers a robust solution to persistent identity drift and confusion issues. If your current S2V models struggle with maintaining consistent human identity across frames or in multi-person scenarios, consider integrating identity-aligned datasets and focused attention mechanisms. This approach significantly enhances identity consistency, improves multi-person subject binding, and reduces unwanted reference-image copying, leading to more natural and reliable video outputs.
Key insights
Identity-faithful human video generation requires explicit identity correspondence and focused attention mechanisms.
Principles
- Explicit identity correspondence is crucial for human S2V.
- Spatially focused supervision enhances identity-aware learning.
- Regulating reference interaction stabilizes identity anchors.
Method
Vera constructs a million-pair identity-aligned dataset, then applies Identity-Focal Masked Supervision (IFMS) for identity-aware learning and Reference-Aware Layer-wise Attention (RALA) within a DiT backbone to manage reference identity cues.
Topics
- Subject-to-Video Generation
- Human Identity Consistency
- Multi-person Video Synthesis
- Deep Implicit Templates
- Computer Vision
- Generative AI
Best for: Research Scientist, AI Scientist, Computer Vision Engineer, Machine Learning Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computer Vision and Pattern Recognition.