Vera: Identity-Faithful Human Subject-to-Video Generation

· Source: Computer Vision and Pattern Recognition · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Computer Vision · Depth: Expert, quick

Summary

Vera is a new unified human-centric Subject-to-Video (S2V) generation framework designed for both single- and multi-person scenarios, addressing the critical issue of identity drift in human subjects across video frames, poses, and interactions. Existing S2V methods often struggle with maintaining identity-critical human details, leading to subject confusion, attribute swapping, and excessive copying of reference appearance cues, particularly in multi-person contexts. To overcome this, Vera leverages a newly constructed million-pair identity-aligned human image-video dataset, built through person-level cross-clip retrieval, which provides explicit identity correspondence. The framework integrates two complementary designs: Identity-Focal Masked Supervision (IFMS) for strengthened identity-aware learning with spatially focused supervision, and Reference-Aware Layer-wise Attention (RALA) to regulate how video tokens interact with reference identity cues within the DiT backbone. Experiments show Vera significantly improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.

Key takeaway

For Computer Vision Engineers developing human-centric video generation systems, Vera offers a robust solution to persistent identity drift and confusion issues. If your current S2V models struggle with maintaining consistent human identity across frames or in multi-person scenarios, consider integrating identity-aligned datasets and focused attention mechanisms. This approach significantly enhances identity consistency, improves multi-person subject binding, and reduces unwanted reference-image copying, leading to more natural and reliable video outputs.

Key insights

Identity-faithful human video generation requires explicit identity correspondence and focused attention mechanisms.

Principles

Method

Vera constructs a million-pair identity-aligned dataset, then applies Identity-Focal Masked Supervision (IFMS) for identity-aware learning and Reference-Aware Layer-wise Attention (RALA) within a DiT backbone to manage reference identity cues.

Topics

Best for: Research Scientist, AI Scientist, Computer Vision Engineer, Machine Learning Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computer Vision and Pattern Recognition.