INCEpTION is a certified open-source web annotation service that has been developed by the Faculty of Computer Science of Technische Universität Darmstadt and is available to all registered users of the CLARIN:EL Research Infrastructure.

INCEpTION offers a generic multi-user annotation environment aiming

  1. to cover three essential aspects of text annotation in a single tool: corpus building, knowledge modelling and annotation and
  2. to combine them with machine-learning-based assistive mechanisms (so-called recommenders) to improve the annotation efficiency and quality.

INCEpTION service is hosted at Kielipankki’s CLARIN partners at CLARIN:EL in Greece. (Click here to view their Privacy Policy.)

To start using the INCEpTION service Click ”Use Service” > ”Log in to access” > ”CLARIN Service Provider Federation login” and select your home organization.

For more information  see the INCEpTION User Documentation.

This resource group page has a Persistent Identifier:

Nordic Tweet Stream (NTS) haku- ja visualisointikäyttöliittymä

In English

NTS on monikielinen monitorikorpus, joka sisältää maantieteellisesti paikannettuja twiittejä ja niihin liittyviä metatietoja Pohjoismaista. Kaikkiaan se sisältää lähes 74 miljoonaa viestiä sadoilta tuhansilta käyttäjätileiltä Tanskasta, Suomesta, Islannista, Norjasta ja Ruotsista. NTS-tiedot kattavat ajanjakson tammikuun 2013 ja toukokuun 2023 välillä, ja ne kerättiin Twitter Academic API:n avulla, joka on nyt suljettu.

NTS:n tarkoituksena on helpottaa SSH:n perustutkimusta. NTS:ssä on helppokäyttöinen graafinen käyttöliittymä, joka tukee nopeaa tiedonsaantia, jotta tutkijat voivat keskittyä tietojen analysointiin. Tietoaineisto mahdollistaa erityyppiset tutkimukset. Esimerkiksi on mahdollista tutkia julkista keskustelua ja tunteita lähihistorian tapahtumista (esim. COVID-19-pandemia, Nato-jäsenyysprosessi jne.). Tietokokonaisuus on myös resurssi sosiolingvistiselle tutkimukselle ja monikielisyyden tutkijoille.

Tutustu verkkosivustoon.

Lisää tietoa NTS:stä

Jos käytät NTS-käyttöliittymää ja hyödynnät tuloksia julkaisuissasi, mainitse hiljattain julkaistu artikkeli, joka on saatavilla verkossa:
[1] Laitinen, Mikko, Jonas Lundberg, Magnus Levin & Rafael Martins. 2018. The Nordic Tweet Stream: A Dynamic Real-Time Monitor Corpus of Big and Rich Language Data, Proc. of Digital Humanities in the Nordic Countries 3rd Conference, Helsinki, Finland, March 7-9, 2018,, online

Tämän sivun pysyvä tunniste:

Nordic Tweet Stream (NTS) search & visualization interface


The NTS is a multilingual monitor corpus of geolocated tweets and associated metadata from the Nordic region. Altogether, it contains nearly 74 million messages from hundreds of thousands of user accounts from Denmark, Finland, Iceland, Norway, and Sweden. The NTS data cover the period between January 2013 and May 2023 and were collected using the Twitter Academic API, which is now closed.

The purpose of the NTS is to facilitate fundamental research in SSH. The NTS comes with an easy-to-use graphic interface that supports quick data access so that researchers can focus on data analysis. The dataset enables various types of research. For instance, it is possible to study public discourses and sentiment concerning events in recent history (e.g., the COVID-19 pandemic, the NATO membership process, etc.). The dataset is also a resource for sociolinguistic research and for scholars of multilingualism.

Please visit the website.

About NTS

If you use the NTS interface and use the findings in your publications, please cite the recent paper, which is available online:
[1] Laitinen, Mikko, Jonas Lundberg, Magnus Levin & Rafael Martins. 2018. The Nordic Tweet Stream: A Dynamic Real-Time Monitor Corpus of Big and Rich Language Data, Proc. of Digital Humanities in the Nordic Countries 3rd Conference, Helsinki, Finland, March 7-9, 2018,, online

This resource group page has a Persistent Identifier:


UDPipe is a trainable pipeline for tokenization, tagging, lemmatization and dependency parsing of CoNLL-U files. UDPipe is language-agnostic and can be trained given annotated data in CoNLL-U format. Trained models are provided for nearly all UD treebanks. UDPipe is available as a binary for Linux/Windows/OS X, as a library for C++, Python, Perl, Java, C#, and as a web service. Third-party R CRAN package also exists.

UDPipe is a free software distributed under the Mozilla Public License 2.0 and the linguistic models are free for non-commercial use and distributed under the CC BY-NC-SA license, although for some models the original data used to create the model may impose additional licensing conditions. UDPipe is versioned using Semantic Versioning.

Copyright 2017 by the Institute of Formal and Applied Linguistics, Faculty of Mathematics and Physics, Charles University, Czech Republic.

Kielipankki version:  
UDPipe Kielipankki version
icon-info-circle Metadata and license
Access to Puhti
Source version:  
icon-info-circle Metadata and license
Access to GitHub
Look for all versions of this tool in META-SHARE  

For more information on this tool have a look at the UDPipe User’s manual


More information on the Kielipankki version:

Using UDPipe on CSC’s servers requires a CSC user account:

UDPipe is installed in CSC’s computing environment (invoke with: module load udpipe) in the following configuration:
Software: UDPipe 1.2.0
Models: 2.3-181115

UDPipe was compiled and installed from Source without local modifications. Please refer to the user’s manual.

The tool was installed using Ansible scripts that can be found here:

This resource group page has a Persistent Identifier:

Finnish Dependency Parsing Pipeline

Kielipankki version:  
Turku Dependency Parser Pipeline, Kielipankki version (TDPP-LBF)
icon-info-circle Metadata and license
Access to GitHub
TurkuNLP Finnish Dependency Parser:  
Finnish dependency parser developed by TurkuNLP (TDPP)
icon-info-circle Metadata and license
Access to GitHub
Look for all versions of this tool in META-SHARE  

The Turku Dependency Parser Pipeline, Kielipankki version (TDPP-LBF) is a version of the open source dependency parsing pipeline developed by the University of Turku NLP group for analyzing Finnish text, adapted by Kielipankki – the Language Bank of Finland.

For further information on the source version please visit the project’s website.


On Kielipankki’s GitHub repository you can find VRT tools adapted from the original pipeline (vrt-tdp-…):

  • vrt-tdp-alpha-fillup
  • vrt-tdp-alpha-lookup
  • vrt-tdp-alpha-marmot
  • vrt-tdp-alpha-parse


This resource group page has a Persistent Identifier:


Transkribus is a comprehensive platform for the digitisation, AI-powered text recognition, transcription and searching of historical documents.

Open the website

User instructions

This resource group page has a Persistent Identifier:

TurkuNLP word embedding demo (word2vec)

A tool developed for analyzing the semantic similarity of words.

The demo is based on word embeddings induced using the word2vec method, trained on 4.5B words of Finnish from the Finnish Internet Parsebank project and over 2B words of Finnish from Suomi24. On the Parsebank project page you can also download the vectors in binary form. The software behind the demo is open-source, available on GitHub. The demo is maintained by the Turku NLP group.

TurkuNLP word embedding demo
icon-info-circle Metadata and license
Try out the demo
icon-info-circle Metadata and license
Tool project page
Search for all versions of this resource in META-SHARE

For word embeddings trained with word2vec and available in Kielipankki – The Language Bank of Finland please visit the wordvec resource group page.


This resource group page has a Persistent Identifier:


The Language Bank’s Webanno instance was shutdown 15.8.2024

Existing and new users are encouraged to start using the much newer INCEpTION service hosted at our CLARIN partners at CLARIN:EL in Greece. (Click here to view their Privacy Policy.)

To start using the INCEpTION service Click ”Use Service” > ”Log in to access” > ”CLARIN Service Provider Federation login” and select your home organization.

For more information  see the INCEpTION User Documentation. If you have questions contact us at kielipankki (ät)

For reference: Historical documentation about WebAnno is available in Github.

This resource group page has a Persistent Identifier:


Mylly service has been discontinued

Due to very low usage, the Mylly service was shut down. If you still have data in Mylly or in case you wish to utilise the Mylly tool scripts on other services, read the instructions here.

Mylly is a versatile data analysis platform with interactive visualizations and workflows. It can be used to build workflows with a variety of tools, including morphosyntactic parsing, character set conversion and speech recognition.

About Mylly

Mylly User Guide

This resource group page has a Persistent Identifier:

Sparv Pipeline

Sparv, Språkbanken’s text analysis tool, is a multilingual toolkit provided by the Swedish Språkbanken for parsing and annotating text in various languages.

Latest version:  
icon-info-circle Metadata and license
Look for all versions of this tool in META-SHARE  

User manual

Latest Sparv release on GitHub

Sparv GUI


This resource group page has a Persistent Identifier:

Turku Neural Parser Pipeline

The Turku Neural Parser Pipeline is a neural parsing pipeline for segmentation, morphological tagging, dependency parsing and lemmatization with pre-trained models for more than 50 languages.

The pipeline is installed in CSC’s computing environment as a Singularity container for the languages Finnish, Swedish and English.

Kielipankki version:  
Turku Neural Parser Pipeline, Kielipankki version (TNPP-LBF)
icon-info-circle Metadata and license
Access to Puhti
TurkuNLP Finnish Neural Parser:  
Turku Neural Parser Pipeline (TNPP)
icon-info-circle Metadata and license
Access to GitHub
Look for all versions of this tool in META-SHARE  

Kielipankki – the Language Bank of Finland has adapted the parser for its VRT format ( CWB-VRT):
Source for the Kielipankki version on GitHub

On Puhti you can see a list of all installed versions and languages using:
module use /appl/soft/ai/singularity/modulefiles/
module spider turku-neural-parser

For more information on this tool have a look at the following links:
Parser Demo
Turku-neural-parser-pipeline manual TNPP no longer maintained by TurkuNLP, see the note from May 2024!
TurkuNLP DockerHub

This resource group page has a Persistent Identifier:

Finnish Tagtools

This software package provides finnish-postag, a part-of-speech and morphology tagger for Finnish, and finnish-nertag, a named entity recogniser for Finnish.
This software is also installed in CSC’s computing environment (module load finnish-tagtools).

Both tools take running text from standard input and produce tabular output (one token per line) to standard output. See –help messages for more details.

An installer is provided in the form of a Makefile. More information can be found in the README file in the download folder.

Latest version:
Finnish Tagtools 1.6
icon-info-circle Metadata and license
Download the resource
Look for all versions of this tool in META-SHARE

This resource group page has a Persistent Identifier:

Search the Language Bank Portal:
Tuukka Törö
Researcher of the Month: Tuukka Törö


Upcoming events


The Language Bank's technical support:
kielipankki (at)
tel. +358 9 4572001

Requests related to language resources:
fin-clarin (at)
tel. +358 29 4129317

More contact information