fMRI2Face: A Full-HD fMRI-Video Dataset and Geometry-Guided Neural Decoding Framework for Dynamic Human Face Reconstruction
Summary
fMRI-Face is introduced as the first large-scale fMRI dataset featuring controllable, full-HD (1920x1080) digital human facial videos. This dataset comprises 62,856 paired fMRI-video samples, collected while participants viewed photorealistic, background-free facial videos with precisely controlled identity, expression, and head pose. Building on this, the fMRI2Face framework is proposed, a geometry-guided neural video decoding system designed to reconstruct dynamic facial videos directly from fMRI signals. fMRI2Face utilizes Brain-derived Appearance Context for global identity attributes and Morphable 3D Facial Control for explicit geometry-aware guidance of pose and expression. These controls are integrated through Neural-Controlled Video Diffusion with auxiliary latent completion. The framework significantly outperforms existing neural decoding methods, improving PSNR from 15.7091 to 18.3238, reducing Fréchet Video Distance from 361.3160 to 82.7316, and boosting identity cosine similarity from 0.2281 to 0.3458.
Key takeaway
For research scientists developing fMRI-based visual decoding systems, you should consider integrating explicit 3D facial geometry priors with global appearance context. This approach, exemplified by fMRI2Face, significantly enhances the fidelity and temporal coherence of reconstructed dynamic faces. Your models will achieve more accurate identity preservation and motion consistency by leveraging controlled digital human datasets and structured facial controls, moving closer to a "digital window" into human perception.
Key insights
Reconstructing dynamic faces from fMRI requires combining global appearance context with explicit 3D geometry-aware motion control.
Principles
- Controlled digital human stimuli enhance fMRI decoding accuracy.
- Explicit 3D facial geometry improves dynamic face reconstruction.
- Complementary neural controls yield high-fidelity video synthesis.
Method
fMRI2Face uses Brain-derived Appearance Context and Morphable 3D Facial Control, integrated via Neural-Controlled Video Diffusion with auxiliary latent completion, to reconstruct facial videos from fMRI signals.
In practice
- Use full-HD, background-free digital human stimuli for fMRI studies.
- Employ 3D parametric models for geometry-consistent facial guidance.
- Integrate appearance and motion controls in video diffusion models.
Topics
- Neural Decoding
- fMRI
- Face Reconstruction
- Video Diffusion Models
- Digital Humans
- Morphable 3D Models
Best for: AI Scientist, Research Scientist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CV updates on arXiv.org.