Abstract
This report describes the 4 submissions for Task 1 (Audio scene
classification) of the DCASE-2016 challenge of the CP-JKU team.
We propose 4 different approaches for Audio Scene Classification
(ASC). First, we propose a novel i-vector extraction scheme for
ASC using both left and right audio channels. Second, we propose a
Deep Convolutional Neural Network (DCNN) architecture trained
on spectrograms of audio excerpts in end-to-end fashion. Third,
we use a calibration transformation to improve the performance of
our binaural i-vector system. Finally, we propose a late-fusion of
our binaural i-vector and the DCNN. We report the performance
of our proposed methods on the provided cross-validation setup for
the DCASE-2016 challenge. Using the late-fusion approach, we
improve the performance of the baseline by 17 percentage point in
accuracy.
Our submissions achieved ranks
first
and
second
among 49
submissions in the audio scene classification task of DCASE-2016
challenge.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the Detection and Classification of Acoustic Scenes and Events (DECASE 2016) |
| Number of pages | 5 |
| Publication status | Published - Sept 2016 |
Fields of science
- 202002 Audiovisual media
- 102 Computer Sciences
- 102001 Artificial intelligence
- 102003 Image processing
- 102015 Information systems
JKU Focus areas
- Computation in Informatics and Mathematics
- Engineering and Natural Sciences (in general)
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver