Score: 1

Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning

Published: August 5, 2025 | arXiv ID: 2508.03102v1

By: Tianjiao Jiang , Zhen Zhang , Yuhang Liu and more

Potential Business Impact:

Teaches computers to learn new things with fewer examples.

Few-shot learning (FSL) often requires effective adaptation of models using limited labeled data. However, most existing FSL methods rely on entangled representations, requiring the model to implicitly recover the unmixing process to obtain disentangled representations using only limited supervision, which hinders effective adaptation. Recent theoretical studies show that multimodal contrastive learning methods, such as CLIP, can disentangle latent representations up to linear transformations. In light of this, we propose the Causal CLIP Adapter (CCA), a novel framework that explicitly disentangles visual features extracted from CLIP using unsupervised Independent Component Analysis (ICA). This removes the need to learn the unmixing process from the labeled data, thereby reducing the number of trainable parameters and mitigating overfitting. Taking a step further, while ICA can obtain visual disentangled representations, it may also disrupt CLIP's intra- and inter-modal alignment. To counteract this, CCA further leverages CLIP's inherent cross-modal alignment by enhancing it in two ways: unidirectionally, through fine-tuning a CLIP-based text classifier, and bidirectionally, via a cross-attention mechanism that enriches visual and textual representations through mutual interaction. Both unimodal and cross-modal classification outputs can be effectively combined linearly to improve classification accuracy. Extensive experiments on 11 benchmark datasets demonstrate that our method consistently outperforms state-of-the-art approaches in terms of few-shot performance and robustness to distributional shifts, while maintaining computational efficiency. Code will be available at https://github.com/tianjiao-j/CCA.

CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images

CV and Pattern Recognition

Finds fake pictures even if made by new tools.

15 Dec 2025 1

89%

CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images

CV and Pattern Recognition

Finds fake pictures even from new AI.

15 Dec 2025 1

88%

Language-Guided Visual Perception Disentanglement for Image Quality Assessment and Conditional Image Generation

CV and Pattern Recognition

Helps computers see images better, not just understand them.

4 Mar 2025 1

View PDF Login to Bookmark

Country of Origin

🇦🇺 Australia

Repos / Data Links

github.com

Page Count

11 pages

Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning

Teaches computers to learn new things with fewer examples.

Technical Abstract

CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images

CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images

Language-Guided Visual Perception Disentanglement for Image Quality Assessment and Conditional Image Generation