TY - JOUR
T1 - LSTCM
T2 - Long-Term and Short-Term Transform of Convolutive Model in the STFT Domain
AU - Pan, Chao
AU - Chen, Jingdong
AU - Benesty, Jacob
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2026
Y1 - 2026
N2 - This paper investigates the convolutive transfer function (CTF) model in the short-time Fourier transform (STFT) domain. We introduce an interpolation process for the source signal in the STFT domain, expressing the source signal as an interpolation of multi-frame STFT-domain signals. Based on this interpolation, we derive a CTF model, where CTF coefficients are proportional to the Fourier transform of the windowed impulse response. Notably, the window is independent of the STFT analysis window. We propose the LSTCM approach, which transfers CTF coefficients between different window lengths and step sizes. The LSTCM consists of two parts: the decoding process and the recoding process. The decoding process converts CTF coefficients into time-domain impulse responses by utilizing frequency band results from the Fourier transform of the upsampled CTF coefficients, concatenating these results, and applying the inverse Fourier transform. The recoding process translates the time-domain impulse response back into CTF coefficients for the target window length and step size. Simulations indicate that an overlap rate greater than 75% between adjacent frames is necessary for an accurate model. To demonstrate the potential of the proposed LSTCM framework, we apply it to establish a connection between a long-term source separation approach and a short-term noise reduction method in the STFT domain. The long-term source separation generates estimates of impulse responses, while the LSTCM builds the accurate model in the short-term STFT domain, leading to a multiple-input/output-inverse-theorem (MINT) filter and a Wiener filter derived from the model parameters. The results illustrate the significant potential of the LSTCM.
AB - This paper investigates the convolutive transfer function (CTF) model in the short-time Fourier transform (STFT) domain. We introduce an interpolation process for the source signal in the STFT domain, expressing the source signal as an interpolation of multi-frame STFT-domain signals. Based on this interpolation, we derive a CTF model, where CTF coefficients are proportional to the Fourier transform of the windowed impulse response. Notably, the window is independent of the STFT analysis window. We propose the LSTCM approach, which transfers CTF coefficients between different window lengths and step sizes. The LSTCM consists of two parts: the decoding process and the recoding process. The decoding process converts CTF coefficients into time-domain impulse responses by utilizing frequency band results from the Fourier transform of the upsampled CTF coefficients, concatenating these results, and applying the inverse Fourier transform. The recoding process translates the time-domain impulse response back into CTF coefficients for the target window length and step size. Simulations indicate that an overlap rate greater than 75% between adjacent frames is necessary for an accurate model. To demonstrate the potential of the proposed LSTCM framework, we apply it to establish a connection between a long-term source separation approach and a short-term noise reduction method in the STFT domain. The long-term source separation generates estimates of impulse responses, while the LSTCM builds the accurate model in the short-term STFT domain, leading to a multiple-input/output-inverse-theorem (MINT) filter and a Wiener filter derived from the model parameters. The results illustrate the significant potential of the LSTCM.
KW - Convolutive transfer function (CTF) model
KW - MINT filter
KW - STFT domain
KW - microphone array processing
UR - https://www.scopus.com/pages/publications/105026087053
U2 - 10.1109/TASLPRO.2025.3646043
DO - 10.1109/TASLPRO.2025.3646043
M3 - 文章
AN - SCOPUS:105026087053
SN - 2998-4173
VL - 34
SP - 324
EP - 338
JO - IEEE Transactions on Audio, Speech and Language Processing
JF - IEEE Transactions on Audio, Speech and Language Processing
ER -