论文信息 - HMM-based music retrieval using stereophonic feature information and framelength adaptation

HMM-based music retrieval using stereophonic feature information and framelength adaptation

Music retrieval methods are in the focus of recent interest due to the increasing size of music databases as e.g. in the Internet. Among different query methods content-based media retrieval analyzing intrinsic characteristics of the source seems to form the most intuitive access. The key-melody in a song can be regarded as the major characteristic in music and leads to a query by humming or singing. In this paper we turn our attention to both, the features and the algorithm of matching in audio music retrieval. Nowadays approaches propagate the use of dynamic time warping for the matching process. As reference mostly midi-data or humming itself is used. However, first attempts matching humming to polyphonic audio exist. In this contribution we introduce hidden Markov models as an alternative for humming queries matching humming itself, mobile phone ring tones and polyphonic audio. The second object of our research is the introduction of a new way of melody enhancement prior to a latter feature extraction by use of stereophonic information. Further an adaptation throughout the extraction process of the frame length to the tempo of a musical piece helps improving similarity matching performance. The paper addresses the design of a working recognition engine and results achieved with respect to the alluded methods. A test database consisting of polyphonic audio clips, ring tones, and sung user data is described in detail.

[1] J. Stephen Downie,et al. Workshop on the creation of standardized test collections, tasks and metrics for music information retrieval (MIR) and music digital library (MDL) evaluation , 2002, JCDL '02.

[2] Lawrence R. Rabiner,et al. A tutorial on hidden Markov models and selected applications in speech recognition , 1989, Proc. IEEE.

[3] Kyoungro Yoon,et al. Query by humming: matching humming query to polyphonic audio , 2002, Proceedings. IEEE International Conference on Multimedia and Expo.

[4] R. Dressler. Dolby Surround Pro Logic II Decoder Principles of Operation , 2000 .

[5] Kunio Kashino,et al. Fast music retrieval using polyphonic binary feature vectors , 2002, Proceedings. IEEE International Conference on Multimedia and Expo.

[6] J. Reiss,et al. Benchmarking Music Information Retrieval Systems , 2002 .

[7] Jeremy Pickens,et al. A Survey of Feature Selection Techniques for Music Information Retrieval , 2001 .