A Tutorial on Neural Networks and Gradient-free Training

November 26, 2022 · The Cartographer · 🏛 arXiv.org

"No code URL or promise found in abstract"
"Title-pattern auto-detect: A Tutorial on Neural Networks and Gradient-free Training"

Evidence collected by the PWNC Scanner

Authors Turibius Rozario, Arjun Trivedi, Ankit Goel arXiv ID 2211.17217 Category eess.SY: Systems & Control (EE) Cross-listed cs.LG Citations 1 Venue arXiv.org Last Checked 4 days ago

Abstract

This paper presents a compact, matrix-based representation of neural networks in a self-contained tutorial fashion. Specifically, we develop neural networks as a composition of several vector-valued functions. Although neural networks are well-understood pictorially in terms of interconnected neurons, neural networks are mathematical nonlinear functions constructed by composing several vector-valued functions. Using basic results from linear algebra, we represent a neural network as an alternating sequence of linear maps and scalar nonlinear functions, also known as activation functions. The training of neural networks requires the minimization of a cost function, which in turn requires the computation of a gradient. Using basic multivariable calculus results, the cost gradient is also shown to be a function composed of a sequence of linear maps and nonlinear functions. In addition to the analytical gradient computation, we consider two gradient-free training methods and compare the three training methods in terms of convergence rate and prediction accuracy.