Score: 0

Understanding the Failure Modes of Transformers through the Lens of Graph Neural Networks

Published: December 9, 2025 | arXiv ID: 2512.09182v1

By: Hunjae Lee

Transformers and more specifically decoder-only transformers dominate modern LLM architectures. While they have shown to work exceptionally well, they are not without issues, resulting in surprising failure modes and predictably asymmetric performance degradation. This article is a study of many of these observed failure modes of transformers through the lens of graph neural network (GNN) theory. We first make the case that much of deep learning, including transformers, is about learnable information mixing and propagation. This makes the study of model failure modes a study of bottlenecks in information propagation. This naturally leads to GNN theory, where there is already a rich literature on information propagation bottlenecks and theoretical failure modes of models. We then make the case that many issues faced by GNNs are also experienced by transformers. In addition, we analyze how the causal nature of decoder-only transformers create interesting geometric properties in information propagation, resulting in predictable and potentially devastating failure modes. Finally, we observe that existing solutions in transformer research tend to be ad-hoc and driven by intuition rather than grounded theoretical motivation. As such, we unify many such solutions under a more theoretical perspective, providing insight into why they work, what problem they are actually solving, and how they can be further improved to target specific failure modes of transformers. Overall, this article is an attempt to bridge the gap between observed failure modes in transformers and a general lack of theoretical understanding of them in this space.

Are GNNs Actually Effective for Multimodal Fault Diagnosis in Microservice Systems?

Software Engineering

Finds computer problems without complex maps.

6 Jan 2025 1

87%

Why does your graph neural network fail on some graphs? Insights from exact generalisation error

Machine Learning (Stat)

Explains why computer "brains" learn from connected data.

12 Sep 2025 1

87%

Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting

Machine Learning (CS)

Makes computer predictions of future events better.

25 Sep 2025 0

View PDF Login to Bookmark

Understanding the Failure Modes of Transformers through the Lens of Graph Neural Networks

Technical Abstract

Are GNNs Actually Effective for Multimodal Fault Diagnosis in Microservice Systems?

Why does your graph neural network fail on some graphs? Insights from exact generalisation error

Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting