Score: 0

From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models

Published: October 20, 2025 | arXiv ID: 2510.18030v1

By: Ziyan Wang , Enmao Diao , Qi Le and more

Potential Business Impact:

Makes smart computer programs smaller and faster.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectures. However, the dominant local paradigm is task-agnostic: by optimizing layer-wise reconstruction rather than task objectives, it tends to preserve perplexity or generic zero-shot behavior but fails to capitalize on modest task-specific calibration signals, often yielding limited downstream gains. We revisit global structured pruning and present GISP-Global Iterative Structured Pruning-a post-training method that removes attention heads and MLP channels using first-order, loss-based important weights aggregated at the structure level with block-wise normalization. An iterative schedule, rather than one-shot pruning, stabilizes accuracy at higher sparsity and mitigates perplexity collapse without requiring intermediate fine-tuning; the pruning trajectory also forms nested subnetworks that support a "prune-once, deploy-many" workflow. Furthermore, because importance is defined by a model-level loss, GISP naturally supports task-specific objectives; we instantiate perplexity for language modeling and a margin-based objective for decision-style tasks. Extensive experiments show that across Llama2-7B/13B, Llama3-8B, and Mistral-0.3-7B, GISP consistently lowers WikiText-2 perplexity and improves downstream accuracy, with especially strong gains at 40-50% sparsity; on DeepSeek-R1-Distill-Llama-3-8B with GSM8K, task-aligned calibration substantially boosts exact-match accuracy.

Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration

Computation and Language

Shrinks big computer brains to work faster.

6 Jan 2026 0

89%

Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models

Computation and Language

Makes smart AI smaller without losing its thinking.

1 Dec 2025 1

89%

Two-Stage Regularization-Based Structured Pruning for LLMs

Machine Learning (CS)

Shrinks big AI models without losing smarts.

23 May 2025 0

View PDF Login to Bookmark

Page Count

16 pages

From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models

Makes smart computer programs smaller and faster.

Technical Abstract

Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration

Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models

Two-Stage Regularization-Based Structured Pruning for LLMs