Score: 2

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation

Published: October 7, 2025 | arXiv ID: 2510.06139v1

By: Zanyi Wang , Dengyang Jiang , Liuzhuozheng Li and more

Potential Business Impact:

Helps computers find and track objects in videos.

Business Areas:

Image Recognition Data and Analytics, Software

Referring Video Object Segmentation (RVOS) requires segmenting specific objects in a video guided by a natural language description. The core challenge of RVOS is to anchor abstract linguistic concepts onto a specific set of pixels and continuously segment them through the complex dynamics of a video. Faced with this difficulty, prior work has often decomposed the task into a pragmatic `locate-then-segment' pipeline. However, this cascaded design creates an information bottleneck by simplifying semantics into coarse geometric prompts (e.g, point), and struggles to maintain temporal consistency as the segmenting process is often decoupled from the initial language grounding. To overcome these fundamental limitations, we propose FlowRVS, a novel framework that reconceptualizes RVOS as a conditional continuous flow problem. This allows us to harness the inherent strengths of pretrained T2V models, fine-grained pixel control, text-video semantic alignment, and temporal coherence. Instead of conventional generating from noise to mask or directly predicting mask, we reformulate the task by learning a direct, language-guided deformation from a video's holistic representation to its target mask. Our one-stage, generative approach achieves new state-of-the-art results across all major RVOS benchmarks. Specifically, achieving a $\mathcal{J}\&\mathcal{F}$ of 51.1 in MeViS (+1.6 over prior SOTA) and 73.3 in the zero shot Ref-DAVIS17 (+2.7), demonstrating the significant potential of modeling video understanding tasks as continuous deformation processes.

Few-Shot Referring Video Single- and Multi-Object Segmentation via Cross-Modal Affinity with Instance Sequence Matching

CV and Pattern Recognition

Lets computers find and track many things in videos.

18 Apr 2025 0

91%

Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model

CV and Pattern Recognition

Helps computers find objects in videos by text.

19 Aug 2025 2

91%

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation

CV and Pattern Recognition

Lets computers find any object in videos using words.

6 Sep 2025 1

View PDF Login to Bookmark

Repos / Data Links

github.com

Page Count

15 pages

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation

Helps computers find and track objects in videos.

Technical Abstract

Few-Shot Referring Video Single- and Multi-Object Segmentation via Cross-Modal Affinity with Instance Sequence Matching

Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation