Indian Art Music Melodic Similarity Dataset

Open data API in a single place

Provided by Zenodo

Get early access to Indian Art Music Melodic Similarity Dataset API!

Let us know and we will figure it out for you.

Dataset information

Country of origin
Updated
2025.07.31 00:00
Created
2015.04.19
Available languages
English
Keywords
Quality scoring

Dataset description

This dataset comprises audio excerpts and manually done annotations of the melodic phrases in Carnatic and Hindustani music. This dataset can be used to develop and evaluate approaches for computing melodic similarity between short-time melodic patterns in Indian art music. This dataset is divided into two parts, one for Carnatic music (CMD), and the other for Hindustani music (HMD). There are two versions of the dataset available: Original version These two datasets, CMD and HMD are compiled originally by the authors of iswar2013 and ross2012, respectively. Though, they have evolved over time and have been recompiled along with the extracted audio features. Please cite  if you use the material shared here in your research work. Gulati, S., Serrà, J., & Serra, X. (2015). An evaluation of methodologies for melodic similarity in audio recordings of Indian art music. In Proceedings of the 40th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 678–682. Brisbane, Australia. [Postprint PDF@MTG] Improved version It was found that several instances of melodic phrases were not marked in the annotations. The missing phrases have been added in the improved version of the dataset. Please cite the following publication if you use the material shared here in your research work. Gulati, S., Serrà, J., & Serra, X. (2015). Improving melodic similarity in Indian art music using culture-specific melodic characteristics. In Proceedings of the 16th International Society for Music Information Retrieval Conference (ISMIR), pp. 680–686. Málaga, Spain. [Postprint PDF@MTG] Dataset structure The dataset is divided into two parts, Carnatic and Hindustani.  Carnatic has 23 folders for each song. In each folder, there are the following files named: <song identifier>.mp3: Performance audio. <song identifier>.anot: Contains the original annotations. <song identifier>.anotEdit1: Contains the improved annotations. <song identifier>.flatSegNyas: Contains nyas annotations. <song identifier>.pitch: Contains original pitch annotations. <song identifier>.pitchSilIntrpPP: Contains improved pitch annotations. <song identifier>.tonic: Tonic of the performance. <song identifier>.tonicFine: Finetuned tonic of the performance. Hindustani has 9 folders for each song. In each folder, there are the following files named: <song identifier>.wav: Performance audio. <song identifier>.anot: Contains the original annotations. <song identifier>.anotEdit4: Contains the improved annotations. <song identifier>.flatSegNyas: Contains nyas annotations. <song identifier>.tpe: Contains original pitch annotations. <song identifier>.tpe5msSilIntrpPP: Contains improved pitch annotations. <song identifier>.tonic: Tonic of the performance. <song identifier>.tonicFine: Finetuned tonic of the performance. Annotation file contains tab separated values with format as: <start_time><tab><end_time><tab><id of the melodic phrase> Mirdata This dataset is included in mirdata. Use the following code snippet to access the dataset in mirdata. # Import midata import mirdata # Initialize dataset dataset_name = 'iam_melodic_similarity' data_home = 'mirdata/dataset' dataset = mirdata.initialize(dataset_name, data_home=data_home) # Download dataset dataset.download() # Validate dataset dataset.validate() # Load dataset as a dictionary with track ids as keys and track objects as values data = dataset.load_tracks() Contact If you have any questions or comments about the dataset, please feel free to email: mtg-info@upf.edu
European data infrastructure with broad catalog discovery, free evaluation access and production-grade API options.
190K+
indexed dataset pages
32
countries and EU institutions
2019
API-first since
Free API quota
for evaluation and prototypes
SLA
history and push on production APIs
FAQ

Questions before production use

Practical answers on evaluation, licensing, freshness, versioning and support.

api.store is built and operated by Apitalks s.r.o. Company details and a direct contact path are linked in the footer for vendor checks and procurement review.
Yes. Selected APIs include a free API quota, so your team can validate coverage, freshness, response shape and workflow fit before asking for a production plan.
Often yes, but usage rights depend on the source license and dataset. We surface source, license and update metadata where available, and can help review terms before a production integration.
Maintained APIs include update metadata where available. For production integrations, we can add history, monitoring and push updates so changes are easier to detect and act on.
Production APIs can add SLA, stable identifiers, versioning support, history, push updates and direct support around the data your product or AI workflow depends on.

Didn't find the API you need?

Let us know and we will figure it out for you.

European data discovery with free evaluation access and production-grade API options.

Copyright © 2026. Made by Apitalks