Tabla Solo dataset

Open data API in a single place

Provided by Zenodo

Get early access to Tabla Solo dataset API!

Let us know and we will figure it out for you.

Dataset information

Country of origin
Updated
2020.01.24 00:00
Created
2015.01.01
Available languages
English
Keywords
Quality scoring

Dataset description

The Tabla Solo dataset is a parallel corpus comprising time-aligned syllabic scores and audio-recordings of 38 solo tabla compositions. The audio and scores for these recordings is from the instructional video DVD titled Shades of Tabla by Pt. Arvind Mulgaonkar. A companion page to the paper is here: http://compmusic.upf.edu/ismir-2015-tabla Introduction In Hindustani music, tabla is the main rhythm accompaniment (Examples of individual strokes of tabla can be obtained from here). To showcase the nuances of the tāl (the rhythmic framework of Hindustani music) as well as the skill of the percussionist with the tabla, Hindustani music concerts feature a tabla solo. A tabla solo is intricate and elaborate, with a variety of pre-composed forms used for developing further elaborations. There are specific principles that govern these elaborations. Musical forms of tabla such as the thēkā, kāyadā, palatā, ̣ rēlā, pēśkār and gat ̣are a part of the solo performance and have different functional and aesthetic roles in a solo performance. Harmonium or sarangi usually plays the role of a time-keeper in tabla solo performances.   Percussion in Hindustani music is organized and orally transmitted with the use of onomatopoeic mnemonic syllables (called the bōl) representative of the different strokes of tabla. Further, tabla has different stylistic schools called gharānās. The repertoires of major gharānās or schools of tabla differ in aspects such as the use of specific bōls, the dynamics of strokes, ornamentation and rhythmic phrases. But there are also many similarities due to the fact that same forms and same standard phrases reappear across these repertoires. The Dataset The syllabic representation for tabla solos provide a meaningful representation for analysis. This dataset uses a such a representation. The dataset comprises audio recordings, scores and time aligned syllabic transcriptions for 38 tabla solo compositions of different forms in tīntāl (a metrical cycle of 16 time units). The compositions are from the instructional video DVD Shades Of Tabla by Pandit Arvind Mulgaonkar, who is among the most renowned contemporary tabla maestros. Out of the 120 compositions in the DVD, we chose 38 representative compositions spanning all the gharānās of tabla (Ajrada, Benaras, Dilli, Lucknow, Punjab, Farukhabad). The dataset contains about 17 minutes of audio with over 8200 syllables. Audio The audio is extracted from the DVD video and segmented at the level of compositions from the full audio recording. The audio files are mono wav files, sampled at 44.1 kHz with a bit depth of 16 bits. All audios have a soft harmonium accompaniment. Annotations The booklet accompanying the DVD provides a syllabic transcription for each composition. We used Tesseract, an open source Optical Character Recognizer (OCR) engine to convert printed scores to a machine readable format. The scores obtained from OCR were manually verified and corrected for errors, adding the the vibhāgs (sections) of the tāl to the syllabic transcription. A time aligned syllabic transcription for each score and audio file pair was obtained using a spectral flux based onset detector followed by manual correction. The score for each composition has additional metadata describing gharānā, composer and its musical form. The scores in the booklet consists of 41 different mnemonic syllables that are reduced and mapped to 18 syllables based on the timbral similarity between the syllables. The list of syllables along with their mapping can be found here: Syllable Mappings Dataset Organization The dataset consists of set of four files for each composition: WAV audio file (*.wav) The syllable scores as retrieved from the booklet with the metadata (*.txt) Time-aligned non-mapped syllabic score with stroke onset times (*.csv) Time-aligned mapped syllabic score with stroke onset times (*.csv) Possible Uses of the Dataset The dataset can be used for variety of of MIR tasks such as onset detection, percussion transcription, rhythm and percussion pattern analysis, and tabla stroke modeling. Using this dataset If you use the dataset in your work, please cite the following publication: S. Gupta, A. Srinivasamurthy, M. Kumar, H. A. Murthy, X. Serra. Discovery of Syllabic Percussion Patterns in Tabla Solo Recordings. In Proc. of the 16th International Society for Music Information Retrieval Conference (ISMIR), 2015. http://hdl.handle.net/10230/25697 We are interested in knowing if you find our datasets useful! If you use our dataset please email us at mtg-info@upf.edu and tell us about your research. Contact Ajay Srinivasamurthy PhD Student, Music Technology Group Universitat Pompeu Fabra, Barcelona, Spain ajays.murthy@upf.edu Swapnil Gupta Masters Student, Sound and Music Computing Universitat Pompeu Fabra, Barcelona, Spain suapnilgupta.iiith@gmail.com Xavier Serra Head, Music Technology Group Universitat Pompeu Fabra, Barcelona, Spain xavier.serra@upf.edu   http://compmusic.upf.edu/tabla-solo-dataset
European data infrastructure with broad catalog discovery, free evaluation access and production-grade API options.
190K+
indexed dataset pages
32
countries and EU institutions
2019
API-first since
Free API quota
for evaluation and prototypes
SLA
history and push on production APIs
FAQ

Questions before production use

Practical answers on evaluation, licensing, freshness, versioning and support.

api.store is built and operated by Apitalks s.r.o. Company details and a direct contact path are linked in the footer for vendor checks and procurement review.
Yes. Selected APIs include a free API quota, so your team can validate coverage, freshness, response shape and workflow fit before asking for a production plan.
Often yes, but usage rights depend on the source license and dataset. We surface source, license and update metadata where available, and can help review terms before a production integration.
Maintained APIs include update metadata where available. For production integrations, we can add history, monitoring and push updates so changes are easier to detect and act on.
Production APIs can add SLA, stable identifiers, versioning support, history, push updates and direct support around the data your product or AI workflow depends on.

Didn't find the API you need?

Let us know and we will figure it out for you.

European data discovery with free evaluation access and production-grade API options.

Copyright © 2026. Made by Apitalks