Score: 1

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs

Published: April 28, 2025 | arXiv ID: 2504.19746v1

By: Xilong Xie , Liang Wang , Limin Xiao and more

Potential Business Impact:

Makes smart computer programs use less power.

Business Areas:

Quantum Computing Science and Engineering

Large language models (LLMs) have significantly advanced the natural language processing paradigm but impose substantial demands on memory and computational resources. Quantization is one of the most effective ways to reduce memory consumption of LLMs. However, advanced single-precision quantization methods experience significant accuracy degradation when quantizing to ultra-low bits. Existing mixed-precision quantization methods are quantized by groups with coarse granularity. Employing high precision for group data leads to substantial memory overhead, whereas low precision severely impacts model accuracy. To address this issue, we propose FineQ, software-hardware co-design for low-bit fine-grained mixed-precision quantization of LLMs. First, FineQ partitions the weights into finer-grained clusters and considers the distribution of outliers within these clusters, thus achieving a balance between model accuracy and memory overhead. Then, we propose an outlier protection mechanism within clusters that uses 3 bits to represent outliers and introduce an encoding scheme for index and data concatenation to enable aligned memory access. Finally, we introduce an accelerator utilizing temporal coding that effectively supports the quantization algorithm while simplifying the multipliers in the systolic array. FineQ achieves higher model accuracy compared to the SOTA mixed-precision quantization algorithm at a close average bit-width. Meanwhile, the accelerator achieves up to 1.79x energy efficiency and reduces the area of the systolic array by 61.2%.

FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference

Hardware Architecture

Makes smart computer programs run faster, use less power.

19 Apr 2025 0

90%

FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference

Machine Learning (CS)

Makes big AI models run faster on less power.

31 Dec 2025 0

90%

InfiJanice: Joint Analysis and In-situ Correction Engine for Quantization-Induced Math Degradation in Large Language Models

Machine Learning (CS)

Fixes AI math mistakes after shrinking it.

16 May 2025 0

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Page Count

7 pages

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs

Makes smart computer programs use less power.

Technical Abstract

FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference

FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference

InfiJanice: Joint Analysis and In-situ Correction Engine for Quantization-Induced Math Degradation in Large Language Models