MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Summary
MultiRef-Compass is a new unified benchmark for multi-reference-to-audio-video (MR2AV) generation, a complex task requiring models to jointly reason over multiple references and generate synchronized audio-visual content. Existing benchmarks primarily focus on text-driven generation or single-reference preservation, leaving MR2AV largely unevaluated. MultiRef-Compass features 350 carefully curated samples constructed through a scalable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. It defines a four-dimension evaluation protocol with 14 sub-metrics, integrating automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework for scalable and auditable assessment. Experiments on eight representative MR2AV systems reveal substantial room for improvement.
Key takeaway
For AI Scientists and Machine Learning Engineers developing or evaluating multi-reference-to-audio-video (MR2AV) generation systems, you should integrate MultiRef-Compass into your workflow. This benchmark provides a comprehensive, auditable framework to rigorously assess model performance across critical dimensions like reference consistency and audio-visual coherence, revealing specific areas for improvement in your MR2AV models.
Key insights
Comprehensive evaluation of multi-reference-to-audio-video generation requires specialized benchmarks beyond existing methods.
Principles
- MR2AV models must jointly reason over multiple references.
- Models must faithfully preserve each reference and correctly bind/compose entities.
- Existing benchmarks are insufficient for MR2AV's unique complexities.
Method
MultiRef-Compass uses a scalable asset-composition pipeline and a rejudging-enhanced MLLM-as-a-Judge framework, evaluating across four dimensions and 14 sub-metrics.
In practice
- Covers multi-view subject preservation and multi-entity binding.
- Assesses human-object-scene composition.
- Evaluates perceptual fidelity and reference-conditioned composition.
Topics
- MR2AV Generation
- MultiRef-Compass
- Audio-Video Generation
- Multi-Reference Evaluation
- MLLM-as-a-Judge
- Benchmark
Best for: Research Scientist, AI Scientist, Machine Learning Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computer Vision and Pattern Recognition.