Engine-Native Editable 3D World Reconstruction with Objects and Lighting

· Source: Takara TLDR - Daily AI Papers · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems, Gaming & Interactive Media · Depth: Expert, medium

Summary

Lumera introduces a benchmark and reference pipeline for engine-native, light-aware 3D scene parsing from a single image, enabling editable 3D world reconstruction with objects and lighting. The Lumera-2K dataset, built from 2,513 UE5 projects, provides 3.73M components, 63M object instances, 102.6K engine-native parametric lights, and 95.1K camera views. Lumera-Box and Lumera-Light adapt Visual Language Models (VLM) to parse object boxes and parametric light tuples (x,y,z,r,g,b,I), integrating per-object mesh reconstruction, HDR environment estimation, and a bounded agentic refinement loop. Lumera-Box achieves strong overall detection, geometry, semantic, and layout scores (merged mAP 0.1141, IoU-B 0.2472, F-score 0.2762) against competitors like DetAny3D and WildDet3D. Lumera-Light recovers almost all non-empty scenes (recall 0.998) but shows limitations in individual-light localization (F1 0.209 at 0.5 m), with median position error 0.261 m, median ΔE2000 4.59, and intensity Pearson r=0.628.

Key takeaway

For computer vision engineers developing 3D scene reconstruction systems, Lumera demonstrates a robust approach to generating engine-native, editable 3D worlds from single images, including parametric lights. You should consider integrating VLM-based parsing for objects and lights, and explore agentic refinement loops to improve scene fidelity and editability. Be aware of current limitations in individual light localization and cross-engine generalization when planning your implementation.

Key insights

Lumera enables engine-native 3D scene reconstruction from a single image, parsing objects and parametric lights for editable environments.

Principles

Method

Lumera's pipeline adapts VLM for object box and parametric light parsing, integrating per-object mesh reconstruction, HDR environment estimation, and a bounded agentic refinement loop.

In practice

Topics

Best for: Research Scientist, AI Scientist, Computer Vision Engineer, Robotics Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Takara TLDR - Daily AI Papers.