A

Speed Limits for Deep Learning

arXiv (Cornell University)

Abstract

State-of-the-art neural networks require extreme computational power to train. It is therefore natural to wonder whether they are optimally trained. Here we apply a recent advancement in stochastic thermodynamics which allows bounding the speed at which one can go from the initial weight distribution to the final distribution of the fully trained network, based on the ratio of their Wasserstein-2 distance and the entropy production rate of the dynamical process connecting them. Considering both gradient-flow and Langevin training dynamics, we provide analytical expressions for these speed limits for linear and linearizable neural networks e.g. Neural Tangent Kernel (NTK). Remarkably, given some plausible scaling assumptions on the NTK spectra and spectral decomposition of the labels -- learning is optimal in a scaling sense. Our results are consistent with small-scale experiments with Convolutional Neural Networks (CNNs) and Fully Connected Neural networks (FCNs) on CIFAR-10, showing a short highly non-optimal regime followed by a longer optimal regime.

Authors 4

  1. Tel Aviv University

    Affiliation as printed

    Department of Applied Mathematics , School of Mathematical Sciences , Tel Aviv Univer- sity , Tel Aviv 69978 , Israel

  2. Google (United States) · Hebrew University of Jerusalem

    Affiliation as printed

    Google Research

    Hebrew University , Racah Institute of Physics , Jerusalem , 9190401 , Israel

  3. Forschungszentrum Jülich · Jülich Aachen Research Alliance

    Affiliation as printed

    Institute of Neuroscience and Medicine (INM-6) , Jülich Research Centre , Jülich , Germany and

  4. RWTH Aachen University

    Affiliation as printed

    Faculty of Physics , RWTH Aachen , Aachen , Germany

Cited by 0 stored of 0

No patents citing this paper on Lens.org (checked 2026-10-06).

References 0