KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding

· Source: cs.CV updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems · Depth: Expert, extended

Summary

KineBench is a novel IDM-free closed-loop benchmark designed to evaluate the physical consistency and actionable potential of Embodied World Models (EWMs). It resolves the attribution ambiguity inherent in existing Inverse Dynamics Model (IDM)-based evaluation frameworks by employing an explicit kinematic grounding pipeline. KineBench extracts 6D end-effector poses directly from generated video frames using cascaded visual foundation models, then executes these poses in a physics simulator for validation. The benchmark integrates two classical 3D kinematic metrics, Spectral Arc Length (SPARC) for trajectory smoothness and the Maruyama Manipulability Index for kinematic feasibility, providing complementary diagnostic signals. Comprising 20 diverse manipulation tasks within ManiSkill3, KineBench is structured into four progressive suites: basic execution, task transfer, visual out-of-distribution generalization, and complexity-conditioned scaling. Initial evaluations on frontier models like Wan 2.1 (1.3B) and CogVideoX (2B) reveal task-complexity-bounded nonlinear scaling in embodied video generation, offering empirical guidance for future data-scaling strategies.

Key takeaway

For Robotics Engineers developing or evaluating Embodied World Models, you should adopt IDM-free benchmarks like KineBench to ensure accurate physical plausibility assessment. Relying on Inverse Dynamics Models introduces confounding factors, making it difficult to diagnose true model failures. Integrate 3D kinematic metrics such as SPARC and the Maruyama Manipulability Index into your evaluation pipeline to gain deeper insights into motion fluency and embodiment awareness, guiding more robust model development.

Key insights

Evaluating EWMs requires IDM-free 3D kinematic grounding to accurately assess physical plausibility and actionability.

Principles

Method

KineBench uses YOLO for 2D mask, MoGeV2 for depth, and FoundationPose for 6D end-effector pose extraction from generated frames, then executes poses in ManiSkill3.

In practice

Topics

Code references

Best for: Research Scientist, AI Scientist, Robotics Engineer, Machine Learning Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CV updates on arXiv.org.