fMRI2Face: A Full-HD fMRI-Video Dataset and Geometry-Guided Neural Decoding Framework for Dynamic Human Face Reconstruction

· Source: cs.CV updates on arXiv.org · Field: Science & Research — Life Sciences & Biology, Research Methodology & Innovation · Depth: Expert, extended

Summary

fMRI-Face is introduced as the first large-scale fMRI dataset featuring controllable, full-HD (1920x1080) digital human facial videos. This dataset comprises 62,856 paired fMRI-video samples, collected while participants viewed photorealistic, background-free facial videos with precisely controlled identity, expression, and head pose. Building on this, the fMRI2Face framework is proposed, a geometry-guided neural video decoding system designed to reconstruct dynamic facial videos directly from fMRI signals. fMRI2Face utilizes Brain-derived Appearance Context for global identity attributes and Morphable 3D Facial Control for explicit geometry-aware guidance of pose and expression. These controls are integrated through Neural-Controlled Video Diffusion with auxiliary latent completion. The framework significantly outperforms existing neural decoding methods, improving PSNR from 15.7091 to 18.3238, reducing Fréchet Video Distance from 361.3160 to 82.7316, and boosting identity cosine similarity from 0.2281 to 0.3458.

Key takeaway

For research scientists developing fMRI-based visual decoding systems, you should consider integrating explicit 3D facial geometry priors with global appearance context. This approach, exemplified by fMRI2Face, significantly enhances the fidelity and temporal coherence of reconstructed dynamic faces. Your models will achieve more accurate identity preservation and motion consistency by leveraging controlled digital human datasets and structured facial controls, moving closer to a "digital window" into human perception.

Key insights

Reconstructing dynamic faces from fMRI requires combining global appearance context with explicit 3D geometry-aware motion control.

Principles

Method

fMRI2Face uses Brain-derived Appearance Context and Morphable 3D Facial Control, integrated via Neural-Controlled Video Diffusion with auxiliary latent completion, to reconstruct facial videos from fMRI signals.

In practice

Topics

Best for: AI Scientist, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CV updates on arXiv.org.