Norwegian Parliamentary Speech Corpus 1.1

Open data API in a single place

Provided by Nasjonalbiblioteket

Get early access to Norwegian Parliamentary Speech Corpus 1.1 API!

Let us know and we will figure it out for you.

Dataset information

Country of origin
Updated
2021.11.30 00:00
Created
2019.08.01
Available languages
English
Keywords
språkteknologi, tale, språkbanken, korpus, språkforskning
Quality scoring

Dataset description

This is version 1.1 of The Norwegian Parliamentary Speech Corpus (NPSC). The following changes have been made in the update from version 1.0 to 1.1: - The data has been split into official training, evaluation and test sets. - Manual dialect annotations were added for each speaker. - The end time of one sentence in 20171208 (sentence_id 45886), was changed, as a 30 minute break was included in the sentence time span in version 1.0. The corresponding audio file (20171208-085509_6122400_6124160.wav was shortened accordingly. - Some of the metadata in the transcriptions of 20171213 were lacking in the json transcription files. These are added in version 1.1. - The documentation has been updated to reflect these changes. The corpus is developed by the Norwegian Language Bank at the National Library of Norway from 2019-2021. The NPSC consists of audio recordings of meetings in Stortinget (the Norwegian parliament), with corresponding orthographic transcriptions in either Norwegian Bokmål or Norwegian Nynorsk, as well as various metadata about the speakers. The official proceedings from the meetings are also included in the corpus for reference. The recordings add up to 140 hours of running speech (including pauses) from 267 unique speakers, and contain 65,000 sentences and 1.2 million words in total. Transcription was first done automatically; subsequently, the output of the automatic process was manually checked and corrected by trained linguists and philologists. Finally, all transcriptions were proofread to ensure consistency and accuracy. NPSC is primarily intended as an open-source dataset for ASR development. The individual audio files in the corpus contain the speech of entire days of plenary meetings from 2017 and 2018 (or, if a meeting lasts more than six hours, the first six hours of the meeting). Since the audio files are quite large, individual audio files for each sentence are also included. Beta releases of the NPSC were published in 2020 and 2021. Note that we have run postprocessing scripts since the last release (0.2) which affect all transcriptions, and the formatting of the transcriptions is different from previous releases. Users should therefore replace old transcription files with the files in this release. We greatly appreciate any feedback and suggestions for improvement. Please use our e-mail address, sprakbanken@nb.no.
European data infrastructure with broad catalog discovery, free evaluation access and production-grade API options.
190K+
indexed dataset pages
32
countries and EU institutions
2019
API-first since
Free API quota
for evaluation and prototypes
SLA
history and push on production APIs
FAQ

Questions before production use

Practical answers on evaluation, licensing, freshness, versioning and support.

api.store is built and operated by Apitalks s.r.o. Company details and a direct contact path are linked in the footer for vendor checks and procurement review.
Yes. Selected APIs include a free API quota, so your team can validate coverage, freshness, response shape and workflow fit before asking for a production plan.
Often yes, but usage rights depend on the source license and dataset. We surface source, license and update metadata where available, and can help review terms before a production integration.
Maintained APIs include update metadata where available. For production integrations, we can add history, monitoring and push updates so changes are easier to detect and act on.
Production APIs can add SLA, stable identifiers, versioning support, history, push updates and direct support around the data your product or AI workflow depends on.

Didn't find the API you need?

Let us know and we will figure it out for you.

European data discovery with free evaluation access and production-grade API options.

Copyright © 2026. Made by Apitalks