Score: 2

Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models

Published: May 8, 2025 | arXiv ID: 2505.05163v2

By: Aishwarya Venkataramanan, Paul Bodesheim, Joachim Denzler

Potential Business Impact:

Shows how sure computers are about pictures and words.

Business Areas:

Image Recognition Data and Analytics, Software

Vision-Language Models (VLMs) learn joint representations by mapping images and text into a shared latent space. However, recent research highlights that deterministic embeddings from standard VLMs often struggle to capture the uncertainties arising from the ambiguities in visual and textual descriptions and the multiple possible correspondences between images and texts. Existing approaches tackle this by learning probabilistic embeddings during VLM training, which demands large datasets and does not leverage the powerful representations already learned by large-scale VLMs like CLIP. In this paper, we propose GroVE, a post-hoc approach to obtaining probabilistic embeddings from frozen VLMs. GroVE builds on Gaussian Process Latent Variable Model (GPLVM) to learn a shared low-dimensional latent space where image and text inputs are mapped to a unified representation, optimized through single-modal embedding reconstruction and cross-modal alignment objectives. Once trained, the Gaussian Process model generates uncertainty-aware probabilistic embeddings. Evaluation shows that GroVE achieves state-of-the-art uncertainty calibration across multiple downstream tasks, including cross-modal retrieval, visual question answering, and active learning.

Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes

CV and Pattern Recognition

Helps AI know when it's unsure.

16 Dec 2025 1

89%

Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere

Machine Learning (CS)

Helps computers understand pictures and words better.

16 May 2025 0

88%

Are vision language models robust to uncertain inputs?

CV and Pattern Recognition

Makes AI admit when it doesn't know.

17 May 2025 1

View PDF Login to Bookmark

Repos / Data Links

github.com

Page Count

22 pages

Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models

Shows how sure computers are about pictures and words.

Technical Abstract

Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes

Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere

Are vision language models robust to uncertain inputs?