Score: 1

A High-Level Compiler Integration Approach for Deep Learning Accelerators Supporting Abstraction and Optimization

Published: July 7, 2025 | arXiv ID: 2507.04828v1

By: Samira Ahmadifarsani, Daniel Mueller-Gritschneder, Ulf Schlichtmann

Potential Business Impact:

Lets computers use new chips faster.

Business Areas:

Field-Programmable Gate Array (FPGA) Hardware

The growing adoption of domain-specific architectures in edge computing platforms for deep learning has highlighted the efficiency of hardware accelerators. However, integrating custom accelerators into modern machine learning (ML) compilers remains a complex challenge due to the need for significant modifications in compilation layers and specialized scheduling techniques. Existing frameworks offer partial solutions and require users to navigate intricate compiler internals. In this paper, we introduce a TVM-based compilation integration approach that targets GEMM-based deep learning accelerators. Our approach abstracts the complexities of compiler integration, enabling seamless integration of accelerators without requiring in-depth knowledge of the underlying compiler. Furthermore, we extend and incorporate design space exploration tools, specifically CoSA, to automate efficient tensor scheduling, accounting for factors such as uneven mapping and double buffering. Our framework is benchmarked on the Gemmini accelerator, demonstrating performance comparable to its specialized manually implemented toolchain.

A Multi-level Compiler Backend for Accelerated Micro-kernels Targeting RISC-V ISA Extensions

Programming Languages

Makes AI run much faster on new chips.

6 Feb 2025 1

88%

Autocomp: LLM-Driven Code Optimization for Tensor Accelerators

Programming Languages

Makes computer chips run programs much faster.

24 May 2025 2

88%

Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems

Distributed, Parallel, and Cluster Computing

Makes AI run faster on different computers.

28 Apr 2025 0

View PDF Login to Bookmark

Country of Origin

🇦🇹 🇩🇪 Germany, Austria

Page Count

8 pages

A High-Level Compiler Integration Approach for Deep Learning Accelerators Supporting Abstraction and Optimization

Lets computers use new chips faster.

Technical Abstract

A Multi-level Compiler Backend for Accelerated Micro-kernels Targeting RISC-V ISA Extensions

Autocomp: LLM-Driven Code Optimization for Tensor Accelerators

Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems