Skip to main navigation Skip to search Skip to main content

FASS: Discrete Fourier Transform Adapters for Speech Separation with Incremental Learning

  • Ziye Yang
  • , Xiang Song
  • , Min Zhao
  • , Jie Chen
  • , Cedric Richard
  • , Israel Cohen
  • Northwestern Polytechnical University Xian
  • Polytechnical University in Shenzhen
  • Xi'an Jiaotong University
  • The University of Hong Kong
  • Université Côte d'Azur
  • Technion-Israel Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Speech separation with incremental learning (SSIL), which facilitates the continuous adaptation of speech separation models to new languages, remains a critical yet underexplored area. In this paper, we propose a novel framework, termed discrete Fourier transform Adapters for Speech Separation with incremental learning (FASS), designed to mitigate catastrophic forgetting by training orthogonal Fourier-domain adapters using a parameter expansion-fusion strategy. This approach stems from our analysis of how acquiring new tasks interferes with retaining prior knowledge. Specifically, FASS converts sparse Fourier-domain parameters into dense parameter-domain weights via the inverse discrete Fourier transform (IDFT), which could preserve previous knowledge without necessitating task identifiers or additional storage. Furthermore, an extended variant, FASS-random (FASS-r), enhances scalability by randomly assigning indices to the learnable frequencies while maintaining theoretical performance. Comprehensive theoretical analyses and extensive experimental results substantiate the effectiveness of our approach.

Original languageEnglish
Pages (from-to)3116-3131
Number of pages16
JournalIEEE Transactions on Audio, Speech and Language Processing
Volume34
DOIs
StatePublished - 2026

Keywords

  • Fourier transform adapters
  • incremental learning
  • orthogonal-space learning
  • speech separation

Fingerprint

Dive into the research topics of 'FASS: Discrete Fourier Transform Adapters for Speech Separation with Incremental Learning'. Together they form a unique fingerprint.

Cite this