Skip to main navigation Skip to search Skip to main content

Learning to Pinpoint Singing Voice from Weakly Labeled Examples

Research output: Chapter in Book/Report/Conference proceedingConference proceedingspeer-review

Abstract

Building an instrument detector usually requires temporally accurate ground truth that is expensive to create. However, song-wise information on the presence of instruments is often easily available. In this work, we investigate how well we can train a singing voice detection system merely from song-wise annotations of vocal presence. Using convolutional neural networks, multipleinstance learning and saliency maps, we can not only detect singing voice in a test signal with a temporal accuracy close to the state-of-the-art, but also localize the spectral bins with precision and recall close to a recent source separation method. Our recipe may provide a basis for other sequence labeling tasks, for improving source separation or for inspecting neural networks trained on auditory spectrograms.
Original languageEnglish
Title of host publicationProceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR
EditorsMichael I. Mandel, Johanna Devaney, Douglas Turnbull, George Tzanetakis
Pages44-50
Number of pages7
ISBN (Electronic)9780692755068
Publication statusPublished - Aug 2016

Fields of science

  • 202002 Audiovisual media
  • 102 Computer Sciences
  • 102001 Artificial intelligence
  • 102003 Image processing
  • 102015 Information systems

JKU Focus areas

  • Computation in Informatics and Mathematics
  • Engineering and Natural Sciences (in general)

Cite this