Skip to main navigation Skip to search Skip to main content

Meta-Learning with Pretrained Audio Representations Enables One-Shot Acoustic Signal Classification

  • Northwestern Polytechnical University Xian
  • Université du Quebec

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Few-shot acoustic signal classification remains a challenging problem due to the high diversity and variability of acoustic data and limited availability of labeled samples. While pretrained audio classification models have proven effective for various acoustic signal classification tasks, fine-tuning them can still lead to overfitting in low-resource settings. In this work, we proposes an attention-based meta-learning framework that operates on the hidden states of a pretrained audio classification model. Specifically, we introduce a trainable hierarchical additive attention module to extract meaningful features from the hidden states of a large-scale pre-trained Audio Spectrogram Transformer (AST). The attention mechanism is trained with a simple meta-learning paradigm, enabling effective adaptation to one-shot learning tasks. We evaluates the proposed model on multiple acoustic signal classification tasks, including acoustic scene classification, sound event recognition and underwater vessel noise classification. Experimental results demonstrate that our proposed framework substantially outperforms the existing methods such as CNN-based prototypical networks in terms of one-shot classification accuracy. This research not only provides an efficient solution for low data resource acoustic pattern recognition tasks but also demonstrate the strong potential of pre-trained audio classification models when combined with metalearning framework for few-shot learning.

Original languageEnglish
Title of host publication2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages172-176
Number of pages5
ISBN (Electronic)9798331572068
DOIs
StatePublished - 2025
Event17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025 - Singapore, Singapore
Duration: 22 Oct 202524 Oct 2025

Publication series

Name2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025

Conference

Conference17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Country/TerritorySingapore
CitySingapore
Period22/10/2524/10/25

Fingerprint

Dive into the research topics of 'Meta-Learning with Pretrained Audio Representations Enables One-Shot Acoustic Signal Classification'. Together they form a unique fingerprint.

Cite this