A hybrid CNN-autoencoder-SVM/XGBoost model for polyphonic orchestral instrument classification

Indonesian Journal of Electrical Engineering and Computer Science

A hybrid CNN-autoencoder-SVM/XGBoost model for polyphonic orchestral instrument classification

Abstract

Polyphonic orchestral recordings pose significant challenges in music information retrieval (MIR) due to their overlapping frequency ranges and timbral similarities among instrument families, which complicate multi-label instrument classification. Prior studies have explored the integration of convolutional neural networks (CNN)-based feature extraction with classical machine learning (ML) classifiers, often on monophonic or simpler datasets like IRMAS. But the integration of deep learning (DL) feature extraction, Autoencoder (AE)-based dimensionality reduction, and ML classifiers for polyphonic orchestral instrument recognition remains underexplored. This study proposes a hybrid framework utilizing a pre-trained Inception V3 CNN for feature extraction from mel-spectrograms, followed by an optional 50% dimensionality reduction via AE, and finally, classification with support vector machines (SVM) or extreme gradient boosting (XGBoost). Experiments were run on two polyphonic datasets, OpenMIC-2018 and Orchset. The results demonstrate that non-AE configurations generally outperform AE variants. These results extend prior studies such using polyphonic datasets. The results highlight the practical value of hybrid CNN-ML pipelines and the trade-offs of feature compression in MIR.

Discover Our Library

Embark on a journey through our expansive collection of articles and let curiosity lead your path to innovation.

Explore Now
Library 3D Ilustration