Score: 3

Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation

Published: March 25, 2025 | arXiv ID: 2503.19647v1

By: Niccolo Avogaro , Thomas Frick , Mattia Rigotti and more

BigTech Affiliations: IBM

Potential Business Impact:

Helps computers understand pictures better by combining words and images.

Business Areas:

Visual Search Internet Services

Large Vision-Language Models (VLMs) are increasingly being regarded as foundation models that can be instructed to solve diverse tasks by prompting, without task-specific training. We examine the seemingly obvious question: how to effectively prompt VLMs for semantic segmentation. To that end, we systematically evaluate the segmentation performance of several recent models guided by either text or visual prompts on the out-of-distribution MESS dataset collection. We introduce a scalable prompting scheme, few-shot prompted semantic segmentation, inspired by open-vocabulary segmentation and few-shot learning. It turns out that VLMs lag far behind specialist models trained for a specific segmentation task, by about 30% on average on the Intersection-over-Union metric. Moreover, we find that text prompts and visual prompts are complementary: each one of the two modes fails on many examples that the other one can solve. Our analysis suggests that being able to anticipate the most effective prompt modality can lead to a 11% improvement in performance. Motivated by our findings, we propose PromptMatcher, a remarkably simple training-free baseline that combines both text and visual prompts, achieving state-of-the-art results outperforming the best text-prompted VLM by 2.5%, and the top visual-prompted VLM by 3.5% on few-shot prompted semantic segmentation.

Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation

CV and Pattern Recognition

Helps computers understand pictures better with words or examples.

6 May 2025 1

91%

Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

CV and Pattern Recognition

Helps computers understand new pictures they haven't seen.

31 Aug 2025 1

90%

Exploring Multimodal Prompt for Visualization Authoring with Large Language Models

Human-Computer Interaction

Draw pictures to help computers make charts.

18 Apr 2025 0

View PDF Login to Bookmark

Country of Origin

🇺🇸 🇨🇭 Switzerland, United States

Page Count

27 pages

Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation

Helps computers understand pictures better by combining words and images.

Technical Abstract

Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation

Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

Exploring Multimodal Prompt for Visualization Authoring with Large Language Models