Abstract
The Transformer architecture drives modern Artificial Intelligence (AI), yet the physical principles that may constrain self-attention training remain poorly characterized. We develop a thermodynamic framework for attention training, drawing on the established Boltzmann correspondence between softmax attention and equilibrium statistical mechanics, and we propose a First Law analogue that decomposes the training energy budget into a heat term (the entropic cost of ordering attention) and a work term (the gain in mutual information about the target). From this framework we derive a Landauer-type bound on learning, which states that the loss reduction during training is bounded below by the entropic cost of structuring attention against thermal noise. The bound is satisfied across all configurations tested: 625 grid points spanning three datasets on a compact Vision Transformer trained from scratch (MNIST, CIFAR-10, and OrganAMNIST), and ten temperatures on a pretrained ViT-Small fine-tuned on Food-101. Reusing the same physical principles at inference time, we show that the thermodynamic work performed by each input patch provides a quantitative, energy-based measure of feature importance that outperforms standard attention weights and Integrated Gradients on ImageNet across pretrained ViT-Small, ViT-Base, and ViT-Large (22M to 304M parameters). The result is an integrated diagnostic framework that links phase structure, training-time bounds, and inference-time attribution within a single empirically falsifiable thermodynamic apparatus.
| Original language | English |
|---|---|
| Article number | 194 |
| Pages (from-to) | 1-28 |
| Number of pages | 28 |
| Journal | AI |
| Volume | 7 |
| Issue number | 6 |
| Early online date | 25 May 2026 |
| DOIs | |
| Publication status | Published (in print/issue) - 30 Jun 2026 |
Bibliographical note
© 2026 by the authors. Licensee MDPI, Basel, Switzerland.Data Availability Statement
All datasets used in this study are publicly available. MNIST [10],CIFAR-10 [11], and Food-101 [14] were obtained via the torchvision library. OrganAMNIST [12] was obtained via the medmnist library. The ImageNet ILSVRC-2012 validation set [15] was downloaded from https://image-net.org (accessed on 21 May 2026). Pretrained Vision Transformer weights were sourced from the timm library [13]. Python (version 3.11.13) scripts implementing the methods can be found at https://github.com/rcsotero/thermo-attention (accessed on 21 May 2026).
Funding
This work was supported by grant RGPIN-2022-03042 from the Natural Sciences and Engineering Council of Canada.
Keywords
- transformers
- attention
- statistical physics
- hysteresis
- landauer limit
- explianable AI
- Landauer limit
- explainable AI
Fingerprint
Dive into the research topics of 'The Thermodynamics of Attention: First Law and Landauer Limit Analogues for Learning and Explainability'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver