Score: 0

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?

Published: June 5, 2025 | arXiv ID: 2506.05287v1

By: Yuqian Yuan , Ronghao Dang , Long Li and more

Potential Business Impact:

Helps robots understand how things change when used.

Business Areas:

Image Recognition Data and Analytics, Software

The emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions. To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios. Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types. To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation framework with four types of questions and design a novel multi-scale temporal accuracy metric for open-ended temporal evaluation. Based on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems.

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark

CV and Pattern Recognition

Helps robots understand and act in the world.

9 Jan 2025 2

90%

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

CV and Pattern Recognition

Teaches computers to understand different viewpoints.

24 Jul 2025 2

90%

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

CV and Pattern Recognition

Tests computers on understanding moving 3D things.

22 Mar 2025 1

View PDF Login to Bookmark

Page Count

32 pages

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?

Helps robots understand how things change when used.

Technical Abstract

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding