Abstract
Speech separation with incremental learning (SSIL), which facilitates the continuous adaptation of speech separation models to new languages, remains a critical yet underexplored area. In this paper, we propose a novel framework, termed discrete Fourier transform Adapters for Speech Separation with incremental learning (FASS), designed to mitigate catastrophic forgetting by training orthogonal Fourier-domain adapters using a parameter expansion-fusion strategy. This approach stems from our analysis of how acquiring new tasks interferes with retaining prior knowledge. Specifically, FASS converts sparse Fourier-domain parameters into dense parameter-domain weights via the inverse discrete Fourier transform (IDFT), which could preserve previous knowledge without necessitating task identifiers or additional storage. Furthermore, an extended variant, FASS-random (FASS-r), enhances scalability by randomly assigning indices to the learnable frequencies while maintaining theoretical performance. Comprehensive theoretical analyses and extensive experimental results substantiate the effectiveness of our approach.
| Original language | English |
|---|---|
| Pages (from-to) | 3116-3131 |
| Number of pages | 16 |
| Journal | IEEE Transactions on Audio, Speech and Language Processing |
| Volume | 34 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Fourier transform adapters
- incremental learning
- orthogonal-space learning
- speech separation
Fingerprint
Dive into the research topics of 'FASS: Discrete Fourier Transform Adapters for Speech Separation with Incremental Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver