Laplacian eigenmaps for automatic story segmentation of broadcast news

Lei Xie, Lilei Zheng, Zihan Liu, Yanning Zhang

科研成果: 期刊稿件文章同行评审

34 引用 (Scopus)

摘要

We propose Laplacian Eigenmaps (LE)-based approaches to automatic story segmentation on speech recognition transcripts of broadcast news. We reinforce story boundaries by applying LE analysis to sentence connective strength matrix and reveal the intrinsic geometric structure of stories. Specifically, we construct a Euclidean space in which each sentence is mapped to a vector. As a result, the original inter-sentence connective strength is reflected by the Euclidean distances between the corresponding vectors and cohesive relations between sentences become geometrically evident. Taking advantage of LE, we present three story segmentation approaches: LE-TextTiling, spectral clustering and LE-DP. In LE-DP, we formalize story segmentation as a straightforward criterion minimization problem and give a fast dynamic programming solution to it. Extensive story segmentation experiments on three corpora demonstrate that the proposed LE-based approaches achieve superior performances and significantly outperform several state-of-the-art methods. For instance, LE-TextTiling obtains a relative F1-measure increase of 17.8% on CCTV Mandarin BN corpus as compared to conventional TextTiling; LE-DP achieves a high F1-measure of 0.7460, which significantly outperforms a recent CRF-prosody approach with an F1-measure of 0.6783 on TDT2 Mandarin BN corpus.

源语言英语
文章编号5934585
页(从-至)276-289
页数14
期刊IEEE Transactions on Audio, Speech and Language Processing
20
1
DOI
出版状态已出版 - 2012

指纹

探究 'Laplacian eigenmaps for automatic story segmentation of broadcast news' 的科研主题。它们共同构成独一无二的指纹。

引用此