Data variability in neural network Bayesian inference
Physical Review Research, vol. 8
Abstract
Bayesian inference and kernel methods are well established in machine learning. The neural network Gaussian process in particular provides a concept to investigate neural networks in the limit of infinitely wide hidden layers by using kernel and inference methods. Here, we build upon this limit and provide a field theoretic formalism that covers the generalization properties of infinitely wide networks. We systematically compute generalization properties of linear, nonlinear, and deep nonlinear networks for kernel matrices with heterogeneous entries. In contrast to currently employed spectral methods, we derive the generalization properties from the statistical properties of the input, elucidating the interplay of input dimensionality, size of the training dataset, and variability of the data. We show that data variability leads to a non-Gaussian action reminiscent of a φ 3 + φ 4 theory. Using our formalism on a synthetic task and on MNIST, we obtain a homogeneous kernel matrix approximation for the learning curve as well as corrections due to data variability that allow the estimation of the generalization properties and exact results for the bounds of the learning curves in the case of infinitely many training data points.
Authors 0
- Author list not loaded yet.
Cited by 0 stored of 0
No patents citing this paper on Lens.org (checked 2026-10-06).
References 30
-
W3176240710details pending0citations
-
W3176723190details pending0citations
-
W2030450972details pending0citations
-
W2072555316details pending0citations
-
W2080792322details pending0citations
-
W3198964499details pending0citations
-
W3216843556details pending0citations