Score: 1

Hyperspectral Variational Autoencoders for Joint Data Compression and Component Extraction

Published: November 23, 2025 | arXiv ID: 2511.18521v1

By: Core Francisco Park , Manuel Perez-Carrasco , Caroline Nowlan and more

Potential Business Impact:

Shrinks huge satellite pictures to share them faster.

Business Areas:

Image Recognition Data and Analytics, Software

Geostationary hyperspectral satellites generate terabytes of data daily, creating critical challenges for storage, transmission, and distribution to the scientific community. We present a variational autoencoder (VAE) approach that achieves x514 compression of NASA's TEMPO satellite hyperspectral observations (1028 channels, 290-490nm) with reconstruction errors 1-2 orders of magnitude below the signal across all wavelengths. This dramatic data volume reduction enables efficient archival and sharing of satellite observations while preserving spectral fidelity. Beyond compression, we investigate to what extent atmospheric information is retained in the compressed latent space by training linear and nonlinear probes to extract Level-2 products (NO2, O3, HCHO, cloud fraction). Cloud fraction and total ozone achieve strong extraction performance (R^2 = 0.93 and 0.81 respectively), though these represent relatively straightforward retrievals given their distinct spectral signatures. In contrast, tropospheric trace gases pose genuine challenges for extraction (NO2 R^2 = 0.20, HCHO R^2 = 0.51) reflecting their weaker signals and complex atmospheric interactions. Critically, we find the VAE encodes atmospheric information in a semi-linear manner - nonlinear probes substantially outperform linear ones - and that explicit latent supervision during training provides minimal improvement, revealing fundamental encoding challenges for certain products. This work demonstrates that neural compression can dramatically reduce hyperspectral data volumes while preserving key atmospheric signals, addressing a critical bottleneck for next-generation Earth observation systems. Code - https://github.com/cfpark00/Hyperspectral-VAE

Variational Autoencoder Framework for Hyperspectral Retrievals (Hyper-VAE) of Phytoplankton Absorption and Chlorophyll a in Coastal Waters for NASA's EMIT and PACE Missions

Machine Learning (CS)

Helps satellites see tiny ocean life changes.

18 Apr 2025 0

89%

Denoising Vision Transformer Autoencoder with Spectral Self-Regularization

CV and Pattern Recognition

Makes AI create better pictures faster.

16 Nov 2025 1

89%

Physically Interpretable Representation Learning with Gaussian Mixture Variational AutoEncoder (GM-VAE)

Machine Learning (CS)

Finds hidden patterns in messy science data.

26 Nov 2025 1

View PDF Login to Bookmark

Country of Origin

🇺🇸 United States

Repos / Data Links

github.com

Page Count

16 pages

Hyperspectral Variational Autoencoders for Joint Data Compression and Component Extraction

Shrinks huge satellite pictures to share them faster.

Technical Abstract

Variational Autoencoder Framework for Hyperspectral Retrievals (Hyper-VAE) of Phytoplankton Absorption and Chlorophyll a in Coastal Waters for NASA's EMIT and PACE Missions

Denoising Vision Transformer Autoencoder with Spectral Self-Regularization

Physically Interpretable Representation Learning with Gaussian Mixture Variational AutoEncoder (GM-VAE)