CoRal - Danish Conversational and Read-aloud Dataset - version 2

Open data API in a single place

Provided by Alexandra Instituttet

Get early access to CoRal - Danish Conversational and Read-aloud Dataset - version 2 API!

Let us know and we will figure it out for you.

Dataset information

Country of origin
Updated
2025.06.26 00:00
Created
2025.03.10
Available languages
English
Keywords
Kultur, Sprog og retskrivning, Kulturarv, Uddannelse, kultur og sport, kultur
Quality scoring

Dataset description

CoRal v2 is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations. Key Features: Dialect and Accent Diversity: The dataset includes speech samples from all major Danish dialects as well as multiple accents, ensuring broad geographical coverage and the inclusion of regional linguistic features. Gender Representation: Both male and female speakers are well-represented, offering balanced gender diversity. Age Range: The dataset includes speakers from a wide range of age groups, providing a comprehensive resource for age-agnostic ASR model development. High-Quality Audio: All recordings are of high quality, ensuring that the dataset can be used for both training and evaluation of high-performance ASR models. Forbidden Use Cases Speech Synthesis and Biometric Identification are not allowed using the CoRal dataset. For more information, see addition 4 in our license (https://huggingface.co/datasets/alexandrainst/coral/blob/main/LICENSE). Access information: This dataset has gated access, meaning access must be requested. Access is available to everyone upon application. A research paper will be submitted soon, but until then, if you use the CoRal dataset in your research or development, please cite it as follows: @dataset{coral2024, author = {Dan Saattrup Nielsen, Sif Bernstorff Lehmann, Simon Leminen Madsen, Anders Jess Pedersen, Anna Katrine van Zee and Torben Blach}, title = {CoRal: A Diverse Danish ASR Dataset Covering Dialects, Accents, Genders, and Age Groups}, year = {2024}, url = {https://hf.co/datasets/alexandrainst/coral}, }
European data infrastructure with broad catalog discovery, free evaluation access and production-grade API options.
190K+
indexed dataset pages
32
countries and EU institutions
2019
API-first since
Free API quota
for evaluation and prototypes
SLA
history and push on production APIs
FAQ

Questions before production use

Practical answers on evaluation, licensing, freshness, versioning and support.

api.store is built and operated by Apitalks s.r.o. Company details and a direct contact path are linked in the footer for vendor checks and procurement review.
Yes. Selected APIs include a free API quota, so your team can validate coverage, freshness, response shape and workflow fit before asking for a production plan.
Often yes, but usage rights depend on the source license and dataset. We surface source, license and update metadata where available, and can help review terms before a production integration.
Maintained APIs include update metadata where available. For production integrations, we can add history, monitoring and push updates so changes are easier to detect and act on.
Production APIs can add SLA, stable identifiers, versioning support, history, push updates and direct support around the data your product or AI workflow depends on.

Didn't find the API you need?

Let us know and we will figure it out for you.

European data discovery with free evaluation access and production-grade API options.

Copyright © 2026. Made by Apitalks