• DocumentCode
    3431629
  • Title

    A nonmonotone learning rate strategy for SGD training of deep neural networks

  • Author

    Keskar, Nitish Shirish ; Saon, George

  • Author_Institution
    Northwestern Univ., Evanston, IL, USA
  • fYear
    2015
  • fDate
    19-24 April 2015
  • Firstpage
    4974
  • Lastpage
    4978
  • Abstract
    The algorithm of choice for cross-entropy training of deep neural network (DNN) acoustic models is mini-batch stochastic gradient descent (SGD). One of the important decisions for this algorithm is the learning rate strategy (also called stepsize selection). We investigate several existing schemes and propose a new learning rate strategy which is inspired by nonmonotone linesearch techniques in nonlinear optimization and the NewBob algorithm. This strategy was found to be relatively insensitive to poorly tuned parameters and resulted in lower word error rates compared to Newbob on two different LVCSR tasks (English broadcast news transcription 50 hours and Switchboard telephone conversations 300 hours). Further, we discuss some justifications for the method by briefly linking it to results in optimization theory.
  • Keywords
    acoustic signal processing; entropy; gradient methods; learning (artificial intelligence); nonlinear programming; speech recognition; stochastic processes; (DNN) acoustic models; NewBob algorithm; automatic speech recognition system; deep neural network SGD training; mini-batch stochastic gradient descent; nonlinear optimization; nonmonotone learning rate strategy; nonmonotone linesearch technique; word error rate; Artificial neural networks; Training; deep learning; learning rate; nonmonotonicity; speech recognition; stepsize;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on
  • Conference_Location
    South Brisbane, QLD
  • Type

    conf

  • DOI
    10.1109/ICASSP.2015.7178917
  • Filename
    7178917