Abstract
Intent-aware music recommender systems are relatively new and advanced approaches in personalized music recommendation. By the integration of the user’s current intent - their main aim for the consumption of the content - these recommender systems can provide more meaningful suggestions that improve user satisfaction. At the same time, Large Language Models (LLMs) appear in many applications nowadays, so it is no surprise that they can have a significant potential and can unlock new opportunities for recommender systems too. Thus, combining these two fields - namely LLMs and intent-aware music recommendation – offers opportunities to improve personalized music recommendations. This thesis investigates the potential of LLMs for intent-aware music recommendation by studying how listening intents can be included in LLM prompts to improve the relevance and intent-alignment of recommendations. 2 main research questions guide this work: (i) How does the inclusion of listening intent and user preferences in LLM prompts affect the relevance and intent-alignment of music recommendations? and (ii) What is the impact of different track–intent assignment strategies on recommendation quality of intent-aware music recommenders using LLMs?
To address these questions, an LLM-based intent-aware music recommender framework is developed that combines the 2020 subset of the LFM-2b dataset, which includes user–track interactions, with the Spotify Million Playlist Dataset enriched with listening intent annotations for each playlist, based on which 5 distinct track-intent matching approaches are implemented to define listening intents on the track level too. 2 LLMs (Google’s Gemini 1.5 Flash and Mistral 7B Instruct v0.3) are evaluated against a Factorization Machine baseline capable of integrating contextual features such as the listening intent of the user. Recommendations are offline evaluated using fuzzy string matching between recommended songs and ground-truth user history tracks in 3 ways: (i) content relevance, where the recommended track has to match one of the tracks the user previously listened to, regardless of the listening intent, (ii) intent-aware relevance, where the listening intent also needs to match, and (iii) intent calibration, which measures distributional alignment between listening intents in the user’s history data and in the recommended songs. Standard accuracy-based and beyond-accuracy metrics such as precision, recall, F1 score, NDCG, MRR, coverage, artist diversity, and hit rate at 10 are also utilized to assess the recommended songs. The results show that including listening intents into LLM prompts improves recommendation quality and intent alignment relative to intent-agnostic prompting strategies, while the Factorization Machine model still provides a competitive baseline. However, several limitations emerge due to the restricted offline evaluation setting, the reliance on fuzzy string matching, and due to the vulnerability of track-intent matching approaches to the robustness of playlist-intent mappings. Additionally, the LLM-based intent-aware music recommender framework faces some scalability challenges, as the need to query an LLM for every user–intent pair introduces computational bottlenecks that can limit feasibility for large-scale deployment.
Overall, this thesis provides insights into the integration of LLMs to the field of intent-aware music recommendation, and highlights both their potential and their current limitations for large-scale, real-world deployment.
To address these questions, an LLM-based intent-aware music recommender framework is developed that combines the 2020 subset of the LFM-2b dataset, which includes user–track interactions, with the Spotify Million Playlist Dataset enriched with listening intent annotations for each playlist, based on which 5 distinct track-intent matching approaches are implemented to define listening intents on the track level too. 2 LLMs (Google’s Gemini 1.5 Flash and Mistral 7B Instruct v0.3) are evaluated against a Factorization Machine baseline capable of integrating contextual features such as the listening intent of the user. Recommendations are offline evaluated using fuzzy string matching between recommended songs and ground-truth user history tracks in 3 ways: (i) content relevance, where the recommended track has to match one of the tracks the user previously listened to, regardless of the listening intent, (ii) intent-aware relevance, where the listening intent also needs to match, and (iii) intent calibration, which measures distributional alignment between listening intents in the user’s history data and in the recommended songs. Standard accuracy-based and beyond-accuracy metrics such as precision, recall, F1 score, NDCG, MRR, coverage, artist diversity, and hit rate at 10 are also utilized to assess the recommended songs. The results show that including listening intents into LLM prompts improves recommendation quality and intent alignment relative to intent-agnostic prompting strategies, while the Factorization Machine model still provides a competitive baseline. However, several limitations emerge due to the restricted offline evaluation setting, the reliance on fuzzy string matching, and due to the vulnerability of track-intent matching approaches to the robustness of playlist-intent mappings. Additionally, the LLM-based intent-aware music recommender framework faces some scalability challenges, as the need to query an LLM for every user–intent pair introduces computational bottlenecks that can limit feasibility for large-scale deployment.
Overall, this thesis provides insights into the integration of LLMs to the field of intent-aware music recommendation, and highlights both their potential and their current limitations for large-scale, real-world deployment.
| Original language | English |
|---|---|
| Supervisors/Reviewers |
|
| Publication status | Published - 2025 |
Fields of science
- 102001 Artificial intelligence
- 102003 Image processing
- 202002 Audiovisual media
- 102015 Information systems
- 102 Computer Sciences
- 101019 Stochastics
- 103029 Statistical physics
- 101018 Statistics
- 101017 Game theory
- 202017 Embedded systems
- 101016 Optimisation
- 101015 Operations research
- 101014 Numerical mathematics
- 101029 Mathematical statistics
- 101028 Mathematical modelling
- 101026 Time series analysis
- 101024 Probability theory
- 102032 Computational intelligence
- 102004 Bioinformatics
- 102013 Human-computer interaction
- 101027 Dynamical systems
- 305907 Medical statistics
- 101004 Biomathematics
- 305905 Medical informatics
- 101031 Approximation theory
- 102033 Data mining
- 305901 Computer-aided diagnosis and therapy
- 102019 Machine learning
- 106007 Biostatistics
- 102018 Artificial neural networks
- 106005 Bioinformatics
- 202037 Signal processing
- 202036 Sensor systems
- 202035 Robotics
JKU Focus areas
- Digital Transformation
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver