Published August 2021 | Version v1
Journal article

Landscape and training regimes in deep learning

  • 1. Institute of Physics, École Polytechnique Fédérale de Lausanne, Lausanne, 1015 (Switzerland)

Description

Deep learning algorithms are responsible for a technological revolution in a variety of tasks including image recognition or Go playing. Yet, why they work is not understood. Ultimately, they manage to classify data lying in high dimension – a feat generically impossible due to the geometry of high dimensional space and the associated curse of dimensionality. Understanding what kind of structure, symmetry or invariance makes data such as images learnable is a fundamental challenge. Other puzzles include that (i) learning corresponds to minimizing a loss in high dimension, which is in general not convex and could well get stuck bad minima. (ii) Deep learning predicting power increases with the number of fitting parameters, even in a regime where data are perfectly fitted. In this manuscript, we review recent results elucidating (i, ii) and the perspective they offer on the (still unexplained) curse of dimensionality paradox. We base our theoretical discussion on the (h,α) plane where h controls the number of parameters and α the scale of the output of the network at initialization, and provide new systematic measures of performance in that plane for two common image classification datasets. We argue that different learning regimes can be organized into a phase diagram. A line of critical points sharply delimits an under-parametrized phase from an over-parametrized one. In over-parametrized nets, learning can operate in two regimes separated by a smooth cross-over. At large initialization, it corresponds to a kernel method, whereas for small initializations features can be learnt, together with invariants in the data. We review the properties of these different phases, of the transition separating them and some open questions. Our treatment emphasizes analogies with physical systems, scaling arguments and the development of numerical observables to quantitatively test these results empirically. Practical implications are also discussed, including the benefit of averaging nets with distinct initial weights, or the choice of parameters (h,α) optimizing performance.

Availability note (English)

Available from http://dx.doi.org/10.1016/j.physrep.2021.04.001

Additional details

Identifiers

DOI
10.1016/j.physrep.2021.04.001;
PII
S0370157321001290;

Publishing Information

Journal Title
Physics Reports
Journal Volume
924
Journal Page Range
p. 1-18
ISSN
0370-1573
CODEN
PRPLCM

INIS

Country of Publication
Netherlands
Country of Input or Organization
International Atomic Energy Agency (IAEA)
INIS RN
54083574
Subject category
S97: MATHEMATICAL METHODS AND COMPUTING;
Descriptors DEI
CLASSIFICATION; GEOMETRY; MACHINE LEARNING; NEURAL NETWORKS; OPTIMIZATION; PERFORMANCE; PHASE DIAGRAMS; SYMMETRY
Descriptors DEC
ALGORITHMS; ARTIFICIAL INTELLIGENCE; DIAGRAMS; INFORMATION; LEARNING; MATHEMATICAL LOGIC; MATHEMATICS

Optional Information

Copyright
Copyright (c) 2021 The Author(s). Published by Elsevier B.V.