Multilingual Music Genre Embeddings for Effective Cross-Lingual Music
Item Annotation
- URL: http://arxiv.org/abs/2009.07755v1
- Date: Wed, 16 Sep 2020 15:39:04 GMT
- Title: Multilingual Music Genre Embeddings for Effective Cross-Lingual Music
Item Annotation
- Authors: Elena V. Epure and Guillaume Salha and Romain Hennequin
- Abstract summary: Cross-lingual music genre translation is possible without relying on a parallel corpus.
By learning multilingual music genre embeddings, we enable cross-lingual music genre translation without relying on a parallel corpus.
Our method is effective in translating music genres across tag systems in multiple languages.
- Score: 9.709229853995987
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Annotating music items with music genres is crucial for music recommendation
and information retrieval, yet challenging given that music genres are
subjective concepts. Recently, in order to explicitly consider this
subjectivity, the annotation of music items was modeled as a translation task:
predict for a music item its music genres within a target vocabulary or
taxonomy (tag system) from a set of music genre tags originating from other tag
systems. However, without a parallel corpus, previous solutions could not
handle tag systems in other languages, being limited to the English-language
only. Here, by learning multilingual music genre embeddings, we enable
cross-lingual music genre translation without relying on a parallel corpus.
First, we apply compositionality functions on pre-trained word embeddings to
represent multi-word tags.Second, we adapt the tag representations to the music
domain by leveraging multilingual music genres graphs with a modified
retrofitting algorithm. Experiments show that our method: 1) is effective in
translating music genres across tag systems in multiple languages (English,
French and Spanish); 2) outperforms the previous baseline in an
English-language multi-source translation task. We publicly release the new
multilingual data and code.
Related papers
- CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models [51.03510073676228]
CLaMP 2 is a system compatible with 101 languages for music information retrieval.
By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale.
CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities.
arXiv Detail & Related papers (2024-10-17T06:43:54Z) - Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval [7.7464988473650935]
Text-to-Music Retrieval plays a pivotal role in content discovery within extensive music databases.
This paper proposes an improved Text-to-Music Retrieval model, denoted as TTMR++.
arXiv Detail & Related papers (2024-10-04T09:33:34Z) - MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models [11.834712543531756]
MuChoMusic is a benchmark for evaluating music understanding in multimodal language models focused on audio.
It comprises 1,187 multiple-choice questions, all validated by human annotators, on 644 music tracks sourced from two publicly available music datasets.
We evaluate five open-source models and identify several pitfalls, including an over-reliance on the language modality.
arXiv Detail & Related papers (2024-08-02T15:34:05Z) - ChatMusician: Understanding and Generating Music Intrinsically with LLM [81.48629006702409]
ChatMusician is an open-source Large Language Models (LLMs) that integrates intrinsic musical abilities.
It can understand and generate music with a pure text tokenizer without any external multi-modal neural structures or tokenizers.
Our model is capable of composing well-structured, full-length music, conditioned on texts, chords, melodies, motifs, musical forms, etc.
arXiv Detail & Related papers (2024-02-25T17:19:41Z) - LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT [48.28624219567131]
We introduce LyricWhiz, a robust, multilingual, and zero-shot automatic lyrics transcription method.
We use Whisper, a weakly supervised robust speech recognition model, and GPT-4, today's most performant chat-based large language model.
Our experiments show that LyricWhiz significantly reduces Word Error Rate compared to existing methods in English.
arXiv Detail & Related papers (2023-06-29T17:01:51Z) - A Dataset for Greek Traditional and Folk Music: Lyra [69.07390994897443]
This paper presents a dataset for Greek Traditional and Folk music that includes 1570 pieces, summing in around 80 hours of data.
The dataset incorporates YouTube timestamped links for retrieving audio and video, along with rich metadata information with regards to instrumentation, geography and genre.
arXiv Detail & Related papers (2022-11-21T14:15:43Z) - Contrastive Audio-Language Learning for Music [13.699088044513562]
MusCALL is a framework for Music Contrastive Audio-Language Learning.
Our approach consists of a dual-encoder architecture that learns the alignment between pairs of music audio and descriptive sentences.
arXiv Detail & Related papers (2022-08-25T16:55:15Z) - Re-creation of Creations: A New Paradigm for Lyric-to-Melody Generation [158.54649047794794]
Re-creation of Creations (ROC) is a new paradigm for lyric-to-melody generation.
ROC achieves good lyric-melody feature alignment in lyric-to-melody generation.
arXiv Detail & Related papers (2022-08-11T08:44:47Z) - Genre-conditioned Acoustic Models for Automatic Lyrics Transcription of
Polyphonic Music [73.73045854068384]
We propose to transcribe the lyrics of polyphonic music using a novel genre-conditioned network.
The proposed network adopts pre-trained model parameters, and incorporates the genre adapters between layers to capture different genre peculiarities for lyrics-genre pairs.
Our experiments show that the proposed genre-conditioned network outperforms the existing lyrics transcription systems.
arXiv Detail & Related papers (2022-04-07T09:15:46Z) - Modeling the Music Genre Perception across Language-Bound Cultures [10.223656553455003]
We study the feasibility of obtaining relevant cross-lingual, culture-specific music genre annotations.
We show that unsupervised cross-lingual music genre annotation is feasible with high accuracy.
We introduce a new, domain-dependent cross-lingual corpus to benchmark state of the art multilingual pre-trained embedding models.
arXiv Detail & Related papers (2020-10-13T12:20:32Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.