This resource contains n-grams - i.e. unigrams, bigrams and trigrams - from all books and newspapers that had been digitized at the National Library of Norway up to September 2013. The n-grams have been extracted from a material consisting of approximately 220.000 books and 540.000 newspapers.
The n-grams are available in two formats, CSV and SQlite: CSV is probably the most interesting format for most developers, because it is very easy to import these files into standard applications. The SQLite files contain indexed databases, which are used in the service NB N-gram. Users who want to contribute to the development of NB N-gram can download the source code on GitHub and the SQLite databases from this page.
A word count by source (books/newspapers) and language variety (Bokmål/Nynorsk) is presented in the json file.
Build on reliable and scalable technology
FAQ
Frequently Asked Questions
Some basic informations about API Store ®.
Operation and development of APIs are currently fully funded by company Apitalks and its usage is for free.
Yes, you can.
All important information such as time of last update, license and other information are in response of each API call.
In case of major update that would not be compatible with previous version of API, we keep for 30 days both versions so you will have enough time to transfer to new version. We will inform you about the changes in advance by e-mail.