DocumentCode
3431629
Title
A nonmonotone learning rate strategy for SGD training of deep neural networks
Author
Keskar, Nitish Shirish ; Saon, George
Author_Institution
Northwestern Univ., Evanston, IL, USA
fYear
2015
fDate
19-24 April 2015
Firstpage
4974
Lastpage
4978
Abstract
The algorithm of choice for cross-entropy training of deep neural network (DNN) acoustic models is mini-batch stochastic gradient descent (SGD). One of the important decisions for this algorithm is the learning rate strategy (also called stepsize selection). We investigate several existing schemes and propose a new learning rate strategy which is inspired by nonmonotone linesearch techniques in nonlinear optimization and the NewBob algorithm. This strategy was found to be relatively insensitive to poorly tuned parameters and resulted in lower word error rates compared to Newbob on two different LVCSR tasks (English broadcast news transcription 50 hours and Switchboard telephone conversations 300 hours). Further, we discuss some justifications for the method by briefly linking it to results in optimization theory.
Keywords
acoustic signal processing; entropy; gradient methods; learning (artificial intelligence); nonlinear programming; speech recognition; stochastic processes; (DNN) acoustic models; NewBob algorithm; automatic speech recognition system; deep neural network SGD training; mini-batch stochastic gradient descent; nonlinear optimization; nonmonotone learning rate strategy; nonmonotone linesearch technique; word error rate; Artificial neural networks; Training; deep learning; learning rate; nonmonotonicity; speech recognition; stepsize;
fLanguage
English
Publisher
ieee
Conference_Titel
Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on
Conference_Location
South Brisbane, QLD
Type
conf
DOI
10.1109/ICASSP.2015.7178917
Filename
7178917
Link To Document