Abstract
Building an instrument detector usually requires temporally accurate ground truth that is expensive to create.
However, song-wise information on the presence of instruments is often easily available. In this work, we investigate how well we can train a singing voice detection system merely from song-wise annotations of vocal
presence. Using convolutional neural networks, multipleinstance learning and saliency maps, we can not only detect singing voice in a test signal with a temporal accuracy
close to the state-of-the-art, but also localize the spectral
bins with precision and recall close to a recent source separation method. Our recipe may provide a basis for other
sequence labeling tasks, for improving source separation
or for inspecting neural networks trained on auditory spectrograms.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR |
| Editors | Michael I. Mandel, Johanna Devaney, Douglas Turnbull, George Tzanetakis |
| Pages | 44-50 |
| Number of pages | 7 |
| ISBN (Electronic) | 9780692755068 |
| Publication status | Published - Aug 2016 |
Fields of science
- 202002 Audiovisual media
- 102 Computer Sciences
- 102001 Artificial intelligence
- 102003 Image processing
- 102015 Information systems
JKU Focus areas
- Computation in Informatics and Mathematics
- Engineering and Natural Sciences (in general)
Prizes
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver