Music/speech classification using high-level features derived from fmri brain imaging

Xi Jiang; Tuo Zhang; Xintao Hu; Lie Lu; Junwei Han; Lei Guo; Tianming Liu

doi:10.1145/2393347.2396322

Music/speech classification using high-level features derived from fmri brain imaging

Xi Jiang, Tuo Zhang, Xintao Hu, Lie Lu, Junwei Han, Lei Guo, Tianming Liu

School of Automation

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution › peer-review

12 Scopus citations

Abstract

With the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features.

Original language	English
Title of host publication	MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia
Pages	825-828
Number of pages	4
DOIs	https://doi.org/10.1145/2393347.2396322
State	Published - 2012
Event	20th ACM International Conference on Multimedia, MM 2012 - Nara, Japan Duration: 29 Oct 2012 → 2 Nov 2012

Publication series

Name	MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia

Conference

Conference	20th ACM International Conference on Multimedia, MM 2012
Country/Territory	Japan
City	Nara
Period	29/10/12 → 2/11/12

Keywords

brain imaging space
functional magnetic resonance imaging
music/speech classification
semantic gap

Access to Document

10.1145/2393347.2396322

Cite this

Jiang, X., Zhang, T., Hu, X., Lu, L., Han, J., Guo, L., & Liu, T. (2012). Music/speech classification using high-level features derived from fmri brain imaging. In MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia (pp. 825-828). (MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia). https://doi.org/10.1145/2393347.2396322

@inproceedings{83d1b1d868e64dd4aee366ce80a7ec61,

title = "Music/speech classification using high-level features derived from fmri brain imaging",

abstract = "With the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features.",

keywords = "brain imaging space, functional magnetic resonance imaging, music/speech classification, semantic gap",

author = "Xi Jiang and Tuo Zhang and Xintao Hu and Lie Lu and Junwei Han and Lei Guo and Tianming Liu",

year = "2012",

doi = "10.1145/2393347.2396322",

language = "英语",

isbn = "9781450310895",

series = "MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia",

pages = "825--828",

booktitle = "MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia",

note = "20th ACM International Conference on Multimedia, MM 2012 ; Conference date: 29-10-2012 Through 02-11-2012",

}

Jiang, X, Zhang, T, Hu, X, Lu, L, Han, J , Guo, L & Liu, T 2012, Music/speech classification using high-level features derived from fmri brain imaging. in MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia. MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia, pp. 825-828, 20th ACM International Conference on Multimedia, MM 2012, Nara, Japan, 29/10/12. https://doi.org/10.1145/2393347.2396322

Music/speech classification using high-level features derived from fmri brain imaging. / Jiang, Xi; Zhang, Tuo; Hu, Xintao et al.
MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia. 2012. p. 825-828 (MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia).

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution › peer-review

TY - GEN

T1 - Music/speech classification using high-level features derived from fmri brain imaging

AU - Jiang, Xi

AU - Zhang, Tuo

AU - Hu, Xintao

AU - Lu, Lie

AU - Han, Junwei

AU - Guo, Lei

AU - Liu, Tianming

PY - 2012

Y1 - 2012

N2 - With the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features.

AB - With the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features.

KW - brain imaging space

KW - functional magnetic resonance imaging

KW - music/speech classification

KW - semantic gap

UR - http://www.scopus.com/inward/record.url?scp=84871373357&partnerID=8YFLogxK

U2 - 10.1145/2393347.2396322

DO - 10.1145/2393347.2396322

M3 - 会议稿件

AN - SCOPUS:84871373357

SN - 9781450310895

T3 - MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia

SP - 825

EP - 828

BT - MM 2012 - Proceedings of the 20th ACM International Conference on Multimedia

T2 - 20th ACM International Conference on Multimedia, MM 2012

Y2 - 29 October 2012 through 2 November 2012

ER -

Music/speech classification using high-level features derived from fmri brain imaging

Abstract

Publication series

Conference

Keywords

Access to Document

Other files and links

Fingerprint

Cite this