3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Computer Vision & 3D Generative Models · Depth: Expert, quick

Summary

3D-DefectBench is introduced as a new benchmark and framework designed for systematically analyzing vision-language model (VLM)-based 3D defect detection pipelines. It addresses the challenge of scaling generative 3D systems by providing nine fine-grained binary defects across geometry, texture, and prompt adherence, offering actionable diagnostics for generator development. The study employed a balanced factorial design, varying VLM, camera protocol, visual input, and prompt schema across 84 inference designs, resulting in approximately 3.2 million scored defect decisions. Findings indicate that while VLM choice is the primary determinant of agreement with human labels, other pipeline factors significantly impact performance and interact with model selection. A cost-effective six-view RGB protocol performed comparably to denser inputs. Critically, even the best of 12 VLM judges lagged trained human labelers, and texture agreement suffered with noisier "silver labels" versus expert consensus. The research emphasizes evaluating automated judges as complete pipelines, not just standalone models.

Key takeaway

For Machine Learning Engineers developing or deploying generative 3D systems, you must evaluate automated defect judges as complete pipelines, not just standalone models. Your evaluation strategy should account for rendering protocols, visual evidence, and prompt schemas, as these factors significantly impact performance and interact with VLM selection. Prioritize expert-consensus human labels for robust texture agreement and consider a cost-effective six-view RGB protocol for visual input.

Key insights

VLM-based 3D defect detection requires evaluating the entire pipeline, not just the model, for reliable automation.

Principles

Method

3D-DefectBench uses a factorial design varying VLM, camera protocol, visual input, and prompt schema across 84 inference designs, generating ~3.2 million defect decisions.

In practice

Topics

Best for: Research Scientist, AI Scientist, Machine Learning Engineer, Computer Vision Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.