Score: 0

Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism

Published: December 29, 2025 | arXiv ID: 2512.23243v1

By: Siyu Zhang , Ying Chen , Lianlei Shan and more

Potential Business Impact:

Makes maps understand details better and faster.

Business Areas:

Image Recognition Data and Analytics, Software

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields such as environmental monitoring and urban planning. To address the deficiencies of existing methods, including the failure of fixed resolutions to balance efficiency and detail, as well as the lack of semantic hierarchy in single-scale alignment, this study proposes a Vision-language Model (VLM) framework integrated with two key innovations: the Dynamic Resolution Input Strategy (DRIS) and the Multi-scale Vision-language Alignment Mechanism (MS-VLAM).Specifically, the DRIS adopts a coarse-to-fine approach to adaptively allocate computational resources according to the complexity of image content, thereby preserving key fine-grained features while reducing redundant computational overhead. The MS-VLAM constructs a three-tier alignment mechanism covering object, local-region and global levels, which systematically captures cross-modal semantic consistency and alleviates issues of semantic misalignment and granularity imbalance.Experimental results on the RS-GPT4V dataset demonstrate that the proposed framework significantly improves the accuracy of semantic understanding and computational efficiency in tasks including image captioning and cross-modal retrieval. Compared with conventional methods, it achieves superior performance in evaluation metrics such as BLEU-4 and CIDEr for image captioning, as well as R@10 for cross-modal retrieval. This technical framework provides a novel approach for constructing efficient and robust multimodal remote sensing systems, laying a theoretical foundation and offering technical guidance for the engineering application of intelligent remote sensing interpretation.

Co-Training Vision Language Models for Remote Sensing Multi-task Learning

CV and Pattern Recognition

Lets computers understand many satellite picture jobs.

26 Nov 2025 2

90%

ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing

CV and Pattern Recognition

Finds real changes in satellite pictures.

29 Dec 2025 1

90%

SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model

CV and Pattern Recognition

Makes blurry satellite pictures clear and detailed.

29 May 2025 2

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Page Count

30 pages

Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism

Makes maps understand details better and faster.

Technical Abstract

Co-Training Vision Language Models for Remote Sensing Multi-task Learning

ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing

SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model