Score: 1

GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions

Published: September 25, 2025 | arXiv ID: 2509.21050v1

By: Bing Liu , Wenqiang Yv , Xuzheng Yang and more

Potential Business Impact:

Helps computers understand math drawings from words.

Business Areas:

Geospatial Data and Analytics, Navigation and Mapping

AI-driven geometric problem solving is a complex vision-language task that requires accurate diagram interpretation, mathematical reasoning, and robust cross-modal grounding. A foundational yet underexplored capability for this task is the ability to identify and interpret geometric elements based on natural language queries. To address this, we introduce the task of Referring Expression Comprehension (REC) for geometric problems, which evaluates whether models can localize points, shapes, and spatial relations in diagrams in response to textual prompts. We present GeoRef, a benchmark dataset constructed from existing geometric problem corpora, featuring diverse, high-quality annotations and queries. Due to the lack of annotated data for this task, we generate a large-scale synthetic training dataset using a structured geometric formal language, enabling broad coverage of geometric concepts and facilitating model adaptation. We explore two fine-tuning approaches: Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). Our results show that GRPO significantly outperforms SFT by better aligning model behavior with task-specific rewards. Furthermore, we propose a verify-and-regenerate mechanism that detects incorrect predictions and re-infers answers using contextual reasoning history, further boosting accuracy. Notably, even state-of-the-art Multimodal Large Language Models (MLLMs) struggle with this task, underscoring the necessity of explicitly evaluating and strengthening geometric grounding as a prerequisite for robust geometric problem solving. Moreover, models trained on GeoRef demonstrate measurable improvements on downstream geometric reasoning tasks, highlighting the broader value of REC as a foundation for multimodal mathematical understanding.

GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language

Artificial Intelligence

Creates better math problems for computers to learn.

31 Oct 2025 1

89%

GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes

CV and Pattern Recognition

Teaches computers to understand maps without human help.

27 Nov 2025 1

88%

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

Robotics

Robots understand and go to exact spots.

4 Jun 2025 1

View PDF Login to Bookmark

Page Count

14 pages

GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions

Helps computers understand math drawings from words.

Technical Abstract

GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language

GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics