MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

· Source: Computer Vision and Pattern Recognition · Field: Technology & Digital — Artificial Intelligence & Machine Learning · Depth: Expert, quick

Summary

MultiRef-Compass is a new unified benchmark for multi-reference-to-audio-video (MR2AV) generation, a complex task requiring models to jointly reason over multiple references and generate synchronized audio-visual content. Existing benchmarks primarily focus on text-driven generation or single-reference preservation, leaving MR2AV largely unevaluated. MultiRef-Compass features 350 carefully curated samples constructed through a scalable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. It defines a four-dimension evaluation protocol with 14 sub-metrics, integrating automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework for scalable and auditable assessment. Experiments on eight representative MR2AV systems reveal substantial room for improvement.

Key takeaway

For AI Scientists and Machine Learning Engineers developing or evaluating multi-reference-to-audio-video (MR2AV) generation systems, you should integrate MultiRef-Compass into your workflow. This benchmark provides a comprehensive, auditable framework to rigorously assess model performance across critical dimensions like reference consistency and audio-visual coherence, revealing specific areas for improvement in your MR2AV models.

Key insights

Comprehensive evaluation of multi-reference-to-audio-video generation requires specialized benchmarks beyond existing methods.

Principles

Method

MultiRef-Compass uses a scalable asset-composition pipeline and a rejudging-enhanced MLLM-as-a-Judge framework, evaluating across four dimensions and 14 sub-metrics.

In practice

Topics

Best for: Research Scientist, AI Scientist, Machine Learning Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computer Vision and Pattern Recognition.