Score: 2

Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

Published: January 12, 2026 | arXiv ID: 2601.07533v1

By: Julian Schelb , Michael Wittweiler , Marie Revellio and more

Potential Business Impact:

Finds old book ideas hidden in new writing.

Business Areas:

Semantic Search Internet Services

Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusions and paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising of a curated dataset of ~172k text segments containing 545 expert-verified parallels linking Late Antique authors to a corpus of classical authors. Using this data, we establish baselines for retrieval and classification of intertextualities with state-of-the-art LLMs.

Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

Computation and Language

Finds old Latin words in mixed-language papers.

22 Oct 2025 0

87%

Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

Computation and Language

Finds old Latin words in messy old papers.

22 Oct 2025 0

86%

The Mediomatix Corpus: Parallel Data for Romansh Idioms via Comparable Schoolbooks

Computation and Language

Helps computers translate between different Romansh languages.

22 Aug 2025 1

View PDF Login to Bookmark

Repos / Data Links

github.com github.com

Page Count

19 pages

Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

Finds old book ideas hidden in new writing.

Technical Abstract

Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

The Mediomatix Corpus: Parallel Data for Romansh Idioms via Comparable Schoolbooks