Score: 0

Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance

Published: January 12, 2026 | arXiv ID: 2601.07221v1

By: Jongwon Ryu , Joonhyung Park , Jaeho Han and more

Multi-domain image-to-image translation re quires grounding semantic differences ex pressed in natural language prompts into corresponding visual transformations, while preserving unrelated structural and seman tic content. Existing methods struggle to maintain structural integrity and provide fine grained, attribute-specific control, especially when multiple domains are involved. We propose LACE (Language-grounded Attribute Controllable Translation), built on two compo nents: (1) a GLIP-Adapter that fuses global semantics with local structural features to pre serve consistency, and (2) a Multi-Domain Control Guidance mechanism that explicitly grounds the semantic delta between source and target prompts into per-attribute translation vec tors, aligning linguistic semantics with domain level visual changes. Together, these modules enable compositional multi-domain control with independent strength modulation for each attribute. Experiments on CelebA(Dialog) and BDD100K demonstrate that LACE achieves high visual fidelity, structural preservation, and interpretable domain-specific control, surpass ing prior baselines. This positions LACE as a cross-modal content generation framework bridging language semantics and controllable visual translation.

Language as an Anchor: Preserving Relative Visual Geometry for Domain Incremental Learning

CV and Pattern Recognition

Keeps AI smart when learning new things.

18 Nov 2025 0

88%

Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation

CV and Pattern Recognition

Helps computers see objects in new places.

3 Dec 2025 1

87%

Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models

CV and Pattern Recognition

Helps computers understand pictures and words better.

24 Aug 2025 1

View PDF Login to Bookmark

Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance

Technical Abstract

Language as an Anchor: Preserving Relative Visual Geometry for Domain Incremental Learning

Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation

Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models