• DocumentCode
    2706072
  • Title

    Rhythm and Tempo Analysis Toward Automatic Music Transcription

  • Author

    Takeda, Haruto ; Nishimoto, Takuya ; Sagayama, Shigeki

  • Author_Institution
    Graduate Sch. of Inf. Sci. & Technol., Tokyo Univ.
  • Volume
    4
  • fYear
    2007
  • fDate
    15-20 April 2007
  • Abstract
    This paper discusses model-based rhythm and tempo analysis of music data in the MIDI format. The data is assumed to be obtained from a module performing multi-pitch analysis of music acoustic signals inside an automatic transcription system. In performed music, observed note lengths and local tempo fluctuate from the nominal note lengths and long-term tempo. Applying the framework of continuous speech recognition to rhythm recognition, we take a probabilistic top-down approach on the joint estimation of rhythm and tempo from the performed onset events in MIDI data. Short-term rhythm patterns are extracted from existing music samples and form a "rhythm vocabulary." Local tempo is represented by a smooth curve. The entire problem is formulated as an integrated optimization problem to maximize a posterior probability, which can be solved by an iterative algorithm which alternately estimates rhythm and tempo. Evaluation of the algorithm through various experiments is also presented.
  • Keywords
    acoustic signal processing; iterative methods; maximum likelihood estimation; music; optimisation; MIDI format; a posterior probability; automatic music transcription; continuous speech recognition; integrated optimization problem; iterative algorithm; model-based rhythm analysis; model-based tempo analysis; multi-pitch analysis; music acoustic signals; music data; probabilistic top-down approach; rhythm recognition; rhythm vocabulary; Automatic speech recognition; Data analysis; Data mining; Iterative algorithms; Multiple signal classification; Music; Performance analysis; Rhythm; Signal analysis; Speech recognition; HMM; Rhythm recognition; Viterbi search; piecewise polynomial tempo curve; rhythm N-gram model; rhythm vocabulary;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing, 2007. ICASSP 2007. IEEE International Conference on
  • Conference_Location
    Honolulu, HI
  • ISSN
    1520-6149
  • Print_ISBN
    1-4244-0727-3
  • Electronic_ISBN
    1520-6149
  • Type

    conf

  • DOI
    10.1109/ICASSP.2007.367320
  • Filename
    4218351