TY - GEN
T1 - Meta-Learning with Pretrained Audio Representations Enables One-Shot Acoustic Signal Classification
AU - Wu, Haoxiang
AU - Zhao, Zhengqiao
AU - Chen, Jingdong
AU - Benesty, Jacob
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Few-shot acoustic signal classification remains a challenging problem due to the high diversity and variability of acoustic data and limited availability of labeled samples. While pretrained audio classification models have proven effective for various acoustic signal classification tasks, fine-tuning them can still lead to overfitting in low-resource settings. In this work, we proposes an attention-based meta-learning framework that operates on the hidden states of a pretrained audio classification model. Specifically, we introduce a trainable hierarchical additive attention module to extract meaningful features from the hidden states of a large-scale pre-trained Audio Spectrogram Transformer (AST). The attention mechanism is trained with a simple meta-learning paradigm, enabling effective adaptation to one-shot learning tasks. We evaluates the proposed model on multiple acoustic signal classification tasks, including acoustic scene classification, sound event recognition and underwater vessel noise classification. Experimental results demonstrate that our proposed framework substantially outperforms the existing methods such as CNN-based prototypical networks in terms of one-shot classification accuracy. This research not only provides an efficient solution for low data resource acoustic pattern recognition tasks but also demonstrate the strong potential of pre-trained audio classification models when combined with metalearning framework for few-shot learning.
AB - Few-shot acoustic signal classification remains a challenging problem due to the high diversity and variability of acoustic data and limited availability of labeled samples. While pretrained audio classification models have proven effective for various acoustic signal classification tasks, fine-tuning them can still lead to overfitting in low-resource settings. In this work, we proposes an attention-based meta-learning framework that operates on the hidden states of a pretrained audio classification model. Specifically, we introduce a trainable hierarchical additive attention module to extract meaningful features from the hidden states of a large-scale pre-trained Audio Spectrogram Transformer (AST). The attention mechanism is trained with a simple meta-learning paradigm, enabling effective adaptation to one-shot learning tasks. We evaluates the proposed model on multiple acoustic signal classification tasks, including acoustic scene classification, sound event recognition and underwater vessel noise classification. Experimental results demonstrate that our proposed framework substantially outperforms the existing methods such as CNN-based prototypical networks in terms of one-shot classification accuracy. This research not only provides an efficient solution for low data resource acoustic pattern recognition tasks but also demonstrate the strong potential of pre-trained audio classification models when combined with metalearning framework for few-shot learning.
UR - https://www.scopus.com/pages/publications/105030454293
U2 - 10.1109/APSIPAASC65261.2025.11249268
DO - 10.1109/APSIPAASC65261.2025.11249268
M3 - 会议稿件
AN - SCOPUS:105030454293
T3 - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
SP - 172
EP - 176
BT - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Y2 - 22 October 2025 through 24 October 2025
ER -