ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting
Summary
ZeroSplat is a novel training-free and zero-feature framework designed for Generalized Referring 3D Gaussian Splatting Segmentation (GR3DGS). This framework addresses fundamental limitations of existing Referring 3D Gaussian Splatting (R3DGS) methods, which are restricted to single-target queries, lack intrinsic 3D point-level understanding, and incur high computational overhead due to per-scene optimization. ZeroSplat overcomes these bottlenecks by lifting 2D Vision-Language Model (VLM) priors into 3D space using robust multi-view geometric constraints, enabling point-level understanding without additional feature storage. The new GR3DGS task, requiring dynamic segmentation of 0, 1, or $N$ targets, is supported by two new benchmarks, GR-LERF and GR-ScanNet. Experiments show ZeroSplat significantly outperforms state-of-the-art methods in both generalized and single-target scenarios while maintaining exceptional efficiency.
Key takeaway
For AI Scientists and Computer Vision Engineers developing 3D scene understanding applications, ZeroSplat offers a compelling alternative to current Referring 3D Gaussian Splatting (R3DGS) methods. If your projects involve complex, multi-target referring segmentation or demand high efficiency without extensive per-scene optimization, you should investigate ZeroSplat. Its training-free, zero-feature approach, leveraging 2D VLM priors, can significantly improve performance and reduce computational overhead for generalized 3D segmentation tasks.
Key insights
ZeroSplat enables generalized 3D referring segmentation by lifting 2D VLM priors to 3D without per-scene optimization or feature storage.
Principles
- Existing R3DGS methods are bottlenecked by 2D pixel operations and per-scene optimization.
- Multi-view geometric constraints can lift 2D VLM priors for intrinsic 3D point-level understanding.
- Generalized referring segmentation requires handling 0, 1, or $N$ targets dynamically.
Method
ZeroSplat lifts 2D Vision-Language Model (VLM) priors into 3D space through robust multi-view geometric constraints, achieving intrinsic point-level understanding without additional feature storage or per-scene optimization.
In practice
- Perform training-free, zero-feature 3D referring segmentation.
- Segment an arbitrary number of targets (0, 1, or $N$) in 3D scenes.
- Evaluate generalized 3D segmentation using GR-LERF and GR-ScanNet benchmarks.
Topics
- 3D Gaussian Splatting
- Referring Segmentation
- Vision-Language Models
- Multi-view Geometry
- 3D Scene Understanding
- Zero-feature Framework
Best for: Research Scientist, AI Scientist, Computer Vision Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Takara TLDR - Daily AI Papers.