Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
By: Ekaterina Artemova , Laurie Burchell , Daryna Dementieva and more
Potential Business Impact:
Builds helpful language tools for rare languages.
This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful language technologies. Participants will walk away with a practical toolkit for building end-to-end NLP pipelines for underrepresented languages -- from data collection and web crawling to parallel sentence mining, machine translation, and downstream applications such as text classification and multimodal reasoning. The tutorial presents strategies for tackling the challenges of data scarcity and cultural variance, offering hands-on methods and modeling frameworks. We will focus on fair, reproducible, and community-informed development approaches, grounded in real-world scenarios. We will showcase a diverse set of use cases covering over 10 languages from different language families and geopolitical contexts, including both digitally resource-rich and severely underrepresented languages.
Similar Papers
Dealing with the Hard Facts of Low-Resource African NLP
Computation and Language
Helps computers understand a rare language.
Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review
Computation and Language
Helps computers talk in less common languages.
Exploring NLP Benchmarks in an Extremely Low-Resource Setting
Computation and Language
Helps computers understand rare languages better.