Score: 0

An Information Theoretic Perspective on Agentic System Design

Published: December 25, 2025 | arXiv ID: 2512.21720v1

By: Shizhe He , Avanika Narayan , Ishan S. Khare and more

Potential Business Impact:

Makes AI smarter by teaching it to remember more.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

Agentic language model (LM) systems power modern applications like "Deep Research" and "Claude Code," and leverage multi-LM architectures to overcome context limitations. Beneath their apparent diversity lies a recurring pattern: smaller "compressor" LMs (that can even run locally) distill raw context into compact text that is then consumed by larger "predictor" LMs. Despite their popularity, the design of compressor-predictor systems remains largely ad hoc, with little guidance on how compressor and predictor choices shape downstream performance. In practice, attributing gains to compression versus prediction requires costly, task-specific pairwise sweeps. We argue that these agentic system design questions are, at root, information-theoretic. Viewing the compressor LM as a noisy channel, we introduce a simple estimator of mutual information between the context and its compression to quantify compression quality in a task-independent way. We show that mutual information strongly predicts downstream performance, independent of any specific task. Through an information-theoretic framework, we perform a comprehensive empirical analysis across five datasets and three model families. Results reveal that larger compressors not only are more accurate, but also more token-efficient, conveying more bits of information per token. A 7B Qwen-2.5 compressor, for instance, is $1.6\times$ more accurate, $4.6\times$ more concise, and conveys $5.5\times$ more bits of mutual information per token than its 1.5B sibling. Across datasets, scaling compressors is substantially more effective than scaling predictors, enabling larger on-device compressors to pair with smaller cloud predictors. Applied to a Deep Research system, these principles enable local compressors as small as 3B parameters to recover $99\%$ of frontier-LM accuracy at $26\%$ of API costs.

Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation

Computation and Language

Makes computers understand writing better for searching.

21 Nov 2025 0

88%

Know Your Limits: Entropy Estimation Modeling for Compression and Generalization

Computation and Language

Makes computers understand and write language better.

13 Nov 2025 2

88%

Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression

Artificial Intelligence

Measures how well AI understands words, saving computer power.

11 Nov 2025 2

View PDF Login to Bookmark

Page Count

43 pages

An Information Theoretic Perspective on Agentic System Design

Makes AI smarter by teaching it to remember more.

Technical Abstract

Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation

Know Your Limits: Entropy Estimation Modeling for Compression and Generalization

Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression