3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
Summary
3D-DefectBench is introduced as a new benchmark and framework designed for systematically analyzing vision-language model (VLM)-based 3D defect detection pipelines. It addresses the challenge of scaling generative 3D systems by providing nine fine-grained binary defects across geometry, texture, and prompt adherence, offering actionable diagnostics for generator development. The study employed a balanced factorial design, varying VLM, camera protocol, visual input, and prompt schema across 84 inference designs, resulting in approximately 3.2 million scored defect decisions. Findings indicate that while VLM choice is the primary determinant of agreement with human labels, other pipeline factors significantly impact performance and interact with model selection. A cost-effective six-view RGB protocol performed comparably to denser inputs. Critically, even the best of 12 VLM judges lagged trained human labelers, and texture agreement suffered with noisier "silver labels" versus expert consensus. The research emphasizes evaluating automated judges as complete pipelines, not just standalone models.
Key takeaway
For Machine Learning Engineers developing or deploying generative 3D systems, you must evaluate automated defect judges as complete pipelines, not just standalone models. Your evaluation strategy should account for rendering protocols, visual evidence, and prompt schemas, as these factors significantly impact performance and interact with VLM selection. Prioritize expert-consensus human labels for robust texture agreement and consider a cost-effective six-view RGB protocol for visual input.
Key insights
VLM-based 3D defect detection requires evaluating the entire pipeline, not just the model, for reliable automation.
Principles
- VLM choice dominates pipeline performance.
- Pipeline factors interact with model selection.
- Human reference labels critically affect agreement.
Method
3D-DefectBench uses a factorial design varying VLM, camera protocol, visual input, and prompt schema across 84 inference designs, generating ~3.2 million defect decisions.
In practice
- Use a compact six-view RGB protocol.
- Calibrate judges across human reference regimes.
- Provide expert-consensus human labels.
Topics
- 3D Generative AI
- Vision-Language Models
- Defect Detection
- Automated Evaluation
- 3D-DefectBench
- Pipeline Optimization
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, Computer Vision Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.