Phonetic Corpus of Estonian Spontaneous Speech v.1.0.3

Open data API in a single place

Provided by Eesti Keele Instituut

Get early access to Phonetic Corpus of Estonian Spontaneous Speech v.1.0.3 API!

Let us know and we will figure it out for you.

Dataset information

Country of origin
Updated
2024.05.16 00:00
Created
2023.06.08
Available languages
English
Keywords
[object Object]
Quality scoring

Dataset description

The aim of the corpus is to compile a large amount of quality recordings of spontaneous Estonian and segment it phonetically on different levels. The project started in autumn 2006. The total size of the corpus is approximately 80 hours of speech from 120 speakers with different dialectological and social background. Speakers are from different age groups. They are asked to participate with face-to-face invitation and they are aware of the purpose of the recordings. Most of the recordings are made in a recording studio, some also on fieldwork. The signal of each speaker is recorded in a separate channel. The distance between the speakers is about 3 meters to minimize the effect of overlaps. For the field-work recordings head-set microphones are used. Recordings are saved in PCM wav-format and are not compressed. Background information about the recordings is collected in a text-file. Segmentation and annotation files are saved as Praat TextGrid files and get same filenames as recordings segmented. Segmentation and annotation Segmentation and annotation is done with the Praat program (www.praat.org). Recordings are segmented manually on different levels (automatic segmentation program is also elaborated and tested). Following tiers are used: -Words (in orthographic spelling), -Phonemes (SAMPA adjusted for Estonian is used for transcription), -Syllables (short – long, open – closed), -Prosodic feet, -Intonation phrases or inter-pausal units; -Changes in voice quality (e.g. creaky voice);
European data infrastructure with broad catalog discovery, free evaluation access and production-grade API options.
190K+
indexed dataset pages
32
countries and EU institutions
2019
API-first since
Free API quota
for evaluation and prototypes
SLA
history and push on production APIs
FAQ

Questions before production use

Practical answers on evaluation, licensing, freshness, versioning and support.

api.store is built and operated by Apitalks s.r.o. Company details and a direct contact path are linked in the footer for vendor checks and procurement review.
Yes. Selected APIs include a free API quota, so your team can validate coverage, freshness, response shape and workflow fit before asking for a production plan.
Often yes, but usage rights depend on the source license and dataset. We surface source, license and update metadata where available, and can help review terms before a production integration.
Maintained APIs include update metadata where available. For production integrations, we can add history, monitoring and push updates so changes are easier to detect and act on.
Production APIs can add SLA, stable identifiers, versioning support, history, push updates and direct support around the data your product or AI workflow depends on.

Didn't find the API you need?

Let us know and we will figure it out for you.

European data discovery with free evaluation access and production-grade API options.

Copyright © 2026. Made by Apitalks