TY - JOUR
T1 - FASS
T2 - Discrete Fourier Transform Adapters for Speech Separation with Incremental Learning
AU - Yang, Ziye
AU - Song, Xiang
AU - Zhao, Min
AU - Chen, Jie
AU - Richard, Cedric
AU - Cohen, Israel
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2026
Y1 - 2026
N2 - Speech separation with incremental learning (SSIL), which facilitates the continuous adaptation of speech separation models to new languages, remains a critical yet underexplored area. In this paper, we propose a novel framework, termed discrete Fourier transform Adapters for Speech Separation with incremental learning (FASS), designed to mitigate catastrophic forgetting by training orthogonal Fourier-domain adapters using a parameter expansion-fusion strategy. This approach stems from our analysis of how acquiring new tasks interferes with retaining prior knowledge. Specifically, FASS converts sparse Fourier-domain parameters into dense parameter-domain weights via the inverse discrete Fourier transform (IDFT), which could preserve previous knowledge without necessitating task identifiers or additional storage. Furthermore, an extended variant, FASS-random (FASS-r), enhances scalability by randomly assigning indices to the learnable frequencies while maintaining theoretical performance. Comprehensive theoretical analyses and extensive experimental results substantiate the effectiveness of our approach.
AB - Speech separation with incremental learning (SSIL), which facilitates the continuous adaptation of speech separation models to new languages, remains a critical yet underexplored area. In this paper, we propose a novel framework, termed discrete Fourier transform Adapters for Speech Separation with incremental learning (FASS), designed to mitigate catastrophic forgetting by training orthogonal Fourier-domain adapters using a parameter expansion-fusion strategy. This approach stems from our analysis of how acquiring new tasks interferes with retaining prior knowledge. Specifically, FASS converts sparse Fourier-domain parameters into dense parameter-domain weights via the inverse discrete Fourier transform (IDFT), which could preserve previous knowledge without necessitating task identifiers or additional storage. Furthermore, an extended variant, FASS-random (FASS-r), enhances scalability by randomly assigning indices to the learnable frequencies while maintaining theoretical performance. Comprehensive theoretical analyses and extensive experimental results substantiate the effectiveness of our approach.
KW - Fourier transform adapters
KW - incremental learning
KW - orthogonal-space learning
KW - speech separation
UR - https://www.scopus.com/pages/publications/105040978219
U2 - 10.1109/TASLPRO.2026.3700032
DO - 10.1109/TASLPRO.2026.3700032
M3 - 文章
AN - SCOPUS:105040978219
SN - 1558-7916
VL - 34
SP - 3116
EP - 3131
JO - IEEE Transactions on Audio, Speech and Language Processing
JF - IEEE Transactions on Audio, Speech and Language Processing
ER -