Score: 0

The aftermath of compounds: Investigating Compounds and their Semantic Representations

Published: October 31, 2025 | arXiv ID: 2510.27477v1

By: Swarang Joshi

Potential Business Impact:

Helps computers understand word meanings better.

Business Areas:
Semantic Search Internet Services

This study investigates how well computational embeddings align with human semantic judgments in the processing of English compound words. We compare static word vectors (GloVe) and contextualized embeddings (BERT) against human ratings of lexeme meaning dominance (LMD) and semantic transparency (ST) drawn from a psycholinguistic dataset. Using measures of association strength (Edinburgh Associative Thesaurus), frequency (BNC), and predictability (LaDEC), we compute embedding-derived LMD and ST metrics and assess their relationships with human judgments via Spearmans correlation and regression analyses. Our results show that BERT embeddings better capture compositional semantics than GloVe, and that predictability ratings are strong predictors of semantic transparency in both human and model data. These findings advance computational psycholinguistics by clarifying the factors that drive compound word processing and offering insights into embedding-based semantic modeling.

Page Count
7 pages

Category
Computer Science:
Computation and Language