Score: 1

Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth

Published: January 6, 2026 | arXiv ID: 2601.02609v1

By: Arjun S. Nair

Potential Business Impact:

Trains AI models much faster and uses less memory.

Business Areas:

Text Analytics Data and Analytics, Software

Large language model fine-tuning is bottlenecked by memory: a 7B parameter model requires 84GB--14GB for weights, 14GB for gradients, and 56GB for FP32 optimizer states--exceeding even A100-40GB capacity. We present Chronicals, an open-source training framework achieving 3.51x speedup over Unsloth through four synergistic optimizations: (1) fused Triton kernels eliminating 75% of memory traffic via RMSNorm (7x), SwiGLU (5x), and QK-RoPE (2.3x) fusion; (2) Cut Cross-Entropy reducing logit memory from 5GB to 135MB through online softmax computation; (3) LoRA+ with theoretically-derived 16x differential learning rates between adapter matrices; and (4) Best-Fit Decreasing sequence packing recovering 60-75% of compute wasted on padding. On Qwen2.5-0.5B with A100-40GB, Chronicals achieves 41,184 tokens/second for full fine-tuning versus Unsloth's 11,736 tokens/second (3.51x). For LoRA at rank 32, we reach 11,699 tokens/second versus Unsloth MAX's 2,857 tokens/second (4.10x). Critically, we discovered that Unsloth's reported 46,000 tokens/second benchmark exhibited zero gradient norms--the model was not training. We provide complete mathematical foundations: online softmax correctness proofs, FlashAttention IO complexity bounds O(N^2 d^2 M^{-1}), LoRA+ learning rate derivations from gradient magnitude analysis, and bin-packing approximation guarantees. All implementations, benchmarks, and proofs are available at https://github.com/Ajwebdevs/Chronicals with pip installation via https://pypi.org/project/chronicals/.

Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism

Distributed, Parallel, and Cluster Computing

Lets computers train bigger AI brains with less memory.

5 Mar 2025 0

85%

Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study

Machine Learning (CS)

Trains smart computer programs on less powerful computers.

7 Sep 2025 0

85%

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers

Machine Learning (CS)

Makes huge AI models train faster, safer.

12 Feb 2025 0

View PDF Login to Bookmark

Repos / Data Links

github.com

Page Count

61 pages

Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth

Trains AI models much faster and uses less memory.

Technical Abstract

Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism

Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers