
Kesän 2026 aikana Puhti korvataan CSC:n uudella Roihu-supertietokoneella. Puhti poistuu käytöstä alustavan aikataulun mukaan seuraavasti:
Tarkistathan ajantasaisen aikataulun CSC:n verkkosivuilta.
Yleiseen käyttöön asennettujen työkalujen lisäksi CSC:n laskentaympäristössä on pieni joukko työkaluja, joita Kielipankki ylläpitää. Kielipankki siirtää useimmat näistä työkaluista uuteen ympäristöön ja osa työkaluista kuuluu puolestaan CSC:n yleiseen ylläpitoon. Joitakin työkaluja ei kuitenkaan enää pystytä asentamaan Roihuun.
Seuraavat työkalut säilyvät myös Roihulla osana Kielipankin valikoimaa:
Seuraavat työkalut tarjotaan Roihussa keskitettynä palveluna:
Seuraavia moduuleita ei enää siirretä Roihuun:
Tähän saakka joidenkin ladattavien aineistojen kopiot ovat olleet käytön helpottamiseksi saatavilla Puhdissa. Samojen aineistojen kopiot on tarkoitus siirtää myös Roihuun. Uudessa järjestelmässä korpuksia ei kuitenkaan käytännön syistä enää pureta valmiiksi, eli ne on tarkoitus tarjota zip-paketteina samassa muodossa, jossa ne löytyvät myös latauspalvelusta.
During summer 2026, Puhti will be replaced by Roihu, the new supercomputer at CSC. Puhti will be gradually retired according to the following tentative schedule:
In the high-performance computing (HPC) environment at CSC, there is a small number of tools that are maintained by the Language Bank of Finland, in addition to the centrally installed tools. Most of these tools will be migrated by the Language Bank, some tools will be supported as part of the general services at CSC, and a few tools can no longer be supported in Roihu.
The following tools will continue to be offered by the Language Bank on Roihu:
The following tools will be offered centrally on HPC:
The following modules will no longer be provided on Roihu:
Until now, the copies of some of the downloadable resources in the Language Bank have been available on Puhti, to make them more conveniently accessible without downloading and transferring large datasets. We aim to keep these resource copies available in Roihu as well. In the new system, however, the resource copies will be provided as zipped archives.
Kaikkien RES-kategoriaan kuuluvien CLARIN-lisenssien PLAN-ehtoon on tehty tarkentava lisäys, joka näkyy seuraavassa korostettuna:
Edellä mainittu päivitys on 12.6.2026 alkaen mukana Kielipankin yleisissä käyttöehdoissa ja se koskee automaattisesti kaikkia kyseisen päivän jälkeen Kielipankin kautta myönnettyjä RES-tyyppisiä lisenssejä.
Muutoksesta ilmoitetaan vielä erikseen sähköpostitse kaikille niille käyttäjille, joilla on vähintään yhden aineiston voimassa oleva käyttöoikeus Kielipankin oikeudet -järjestelmässä. Uudet ehdot astuvat voimaan myös aikaisemmin myönnettyjen lisenssien osalta 60 päivän kuluttua muutosviestin lähettämisestä.
Pääsy Kielipankin kautta saatavilla oleviin luvanvaraisiin aineistoihin on jo tähänkin saakka voitu myöntää vain hakemuksessa annetun asianmukaisen selvityksen perusteella, joten lisenssin tarkentuminen ei aiheuta käytännön muutoksia Kielipankin lupaprosessiin.
Jos sinulla on jonkin luvanvaraisen aineiston käyttöoikeus Kielipankissa, varmista, että kyseisen aineiston käyttäminen on edelleen tarpeen tutkimuksesi kannalta. Mikäli et enää tarvitse aineistoa alkuperäisen hakemuksesi mukaiseen tarkoitukseen, sinun tulee luopua käyttöoikeudestasi joko sulkemalla hakemus Kielipankin oikeudet -palvelussa tai ilmoittamalla asiasta Kielipankille (ks. yhteystiedot). Tarvittaessa voit hakea pääsyä myöhemmin uudelleen, jos vaikkapa aloitat uuden tutkimushankkeen muutaman vuoden kuluttua.
Uutena käytäntönä CLARIN PUB-, CLARIN ACA- tai CLARIN RES-luokkiin kuuluvien lisenssien vakioehtojen lyhenteet merkitään lisenssin otsikon perään. Tarkoitus on, että jo pelkän otsikon perusteella voi saada alustavan käsityksen siitä, minkä tyyppisiä ehtoja lisenssiin sisältyy.
Huomaa kuitenkin, että CLARIN-lisenssien otsikot ovat edelleen viitteellisiä. Siirtymävaiheen aikana lisenssien otsikoissa voi myös vielä esiintyä vaihtelua. Kunkin lisenssin tarkat ja sitovat ehdot löytyvät lisenssin tekstistä.
Erilaisiin CLARIN-lisensseihin ja lisäehtojen yhdistelmiin voi tutustua tukisivulla. Yksittäisten Kielipankin kautta välitettävien aineistojen lisenssitiedot löytyvät Kielipankin verkkosivuilla olevasta aineistoluettelosta (Lisenssi-sarake).
A small clarification has been made to the PLAN condition of all CLARIN licenses in the RES category. The added text is highlighted below:
The aforementioned update is included in the Terms of Use of the Language Bank of Finland, effective June 12, 2026, and the updated condition will automatically apply to all RES-type licenses granted through Kielipankki on or after that date.
All users who have a valid license for at least one resource in the Language Bank Rights system will be notified of this change separately via email. The new terms and conditions will take effect for previously granted licenses 60 days after the change notification is sent.
Even until now, access to restricted resources available through the Language Bank has only been granted on the basis of an appropriate plan included in the application. Thus, no changes will be needed in the licensing process in practice.
If you have access to a restricted resource via the Language Bank, please check whether the use of the material is still required for your research purpose. In case you no longer need the data for the purpose specified in your original application, you must give up your access rights either by closing the application in the Language Bank Rights service or by notifying the Language Bank (see contact information). You can reapply for access later, e.g., if you start working on a new research project in a few years.
As a new practice, the abbreviations for the standard terms and conditions of licenses falling under the CLARIN PUB, CLARIN ACA, or CLARIN RES categories are now listed after the license title. The title can thus provide a preliminary overview of the types of terms included in the license.
Please note, however, that the titles of the CLARIN licenses are still provided for reference purposes only. Moreover, there may still be some variation in the titles of the licenses during the transition period. The detailed and binding terms and conditions are specified in the license text.
You can find information about the various CLARIN licenses and combinations of the additional terms on the support page. The license information for individual resources distributed through the Language Bank of Finland can be found on the list of corpora (”License” column) on the website of the Language Bank.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 367751
Start date: 01-01-2026
Duration: 24 months
WP 2.3: Report on developing licensing and protection schemes for sharing sign language data
Date of reporting: 11-06-2026
Report author: Mietta Lennes (University of Helsinki)
Deliverable location: https://www.kielipankki.fi/corpora/resource-families-fin-clarin/sign-language-resources/
Keywords for the deliverable page: sign language; personal data; video processing; sensitive data; SD Desktop
Several resource groups containing sign language material are available via the Language Bank of Finland. The very first sign language resource containing sign language was The Kipo Corpus (2010 The Language Policy Programme for the National Sign Languages in Finland), published openly in 2015. In the years 2016, 2019 and 2024-2025, large numbers of annotated sign language recordings have been published in the CFINSL and CFSTS resource groups of Finnish and Finland-Swedish Sign Language. The sign language corpora can be found on the website of the Language Bank, under the Sign Language Resource Family.
Most sign language resources tend to contain personal data, as the signers are identifiable on the video recordings on the basis of their face, physical appearance and movements. In free signing and signed conversation, the signers may also refer to other people. The data usually cannot be anonymized for research purposes. Due to the personal data, the decisions on the appropriate end-user licenses and the data protection schemes must be made on the basis of the information given to the data subjects (the participating signers), and on the evaluation of the potential risks vs. benefits regarding the processing of the types of data in question.
For sign language communities, it is often desirable to make some language data publicly available. By informing the participating signers in an appropriate way, it is possible to publish the content openly, given that the publication is not considered harmful to the people involved. Some of the above-mentioned resources were made publicly available via the Language Bank of Finland, whereas others are only available for research purposes upon application.
The depositing organization is generally responsible for setting the terms and conditions on how the personal data can be processed and redistributed. If protection is needed, the Language Bank offers options for managing and restricting access to the data via federated academic login (CLARIN ACA type licenses) or individual access granted upon application (CLARIN RES type licenses).
For additional protection, it is even possible to share the data in packages that are separately encrypted for individual users, or the data can be made accessible via SD Desktop provided by CSC. However, the latter two options are currently not used for sign language data. The encryption of large amounts of video files is time-consuming and would often not be in proportion with the protection requirements, since encrypted data would still need to be decrypted for the actual research use. SD Desktop offers a secure environment for analyzing and processing data. The current tools and technical properties of SD Desktop may not yet be sufficient for the convenient playback, annotation and analysis of sign language videos. However, we are collaborating with CSC to investigate the possibilities for adding tools on SD Desktop that would enable users to run useful analyses in manual or batch mode, to produce data that can be safely exported from the secure environment. For further details on sensitive data, see Deliverable 2.1.1 and the support page regarding sensitive data in the Language Bank.
The FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 367751.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 367751
Start date: 01-01-2026
Duration: 24 months
WP 2.3: Report on developing policies for processing and sharing translation memories
Date of reporting: 11-06-2026
Report authors: Mietta Lennes (University of Helsinki)
Deliverable location: https://www.kielipankki.fi/support/data-management/dela/
Keywords for the deliverable page: translation memory, machine translation
A translation memory is a bilingual or multilingual database of previous translations of text segments of varying size. In some cases, a translation memory (or a translation memory manager/system) can also refer to a language-technological tool that helps translators by suggesting translations on the basis of similar, previously translated segments of text. (For the definitions in Finnish, see https://tieteentermipankki.fi/wiki/Language_Technology:translation-memory.)
Professional translators use translation memories as part of their workflow. Translation memories help translators in maintaining consistent terminology across documents and make their work significantly faster. Translation memories are thus a valuable source for research in language technology, linguistics, terminology and translation studies. For example, by using machine learning techniques, translation memories could be used for analyzing specific types of translation solutions, for enriching and extending the existing data, or for extending the translation solutions to other languages. However, due to, e.g., copyright reasons, translation memories can often only be used and shared within a company or an organization. It can be difficult to share translation memories with a larger research community.
In case a translation memory is publicly available, or in case a deposition agreement can be reached with the rightholders about the appropriate restrictions of use regarding research purposes, it is possible to provide access to the data via the Language Bank of Finland. The translation memory database can be made available via the Language Bank of Finland as a downloadable package. It is also possible to consider different platforms for accessing and querying the content, e.g., as a parallel corpus via the Korp concordancer, or as a lexical resource via the Karp platform. Karp is currently not yet available in the Language Bank of Finland, but the Language Bank will discuss the possibilities of installing it (see the original Karp in Språkbanken, the Language Bank of Sweden). In case the translation memory includes very sensitive content that requires additional protection, it is also possible to use the secure SD Desktop environment for providing access to the data (cf., Deliverable 2.1.1 and the support page regarding sensitive data in the Language Bank).
The FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 367751.
This page outlines the project deliverables for 2022-2023 (completed).
(completed)
| D1.1.1 | Updating LBF resource selection | 2022-09 |
| D1.1.2 | Ingesting new unstructured resources | 2023-12 |
| D1.2.1 | Forced-Alignment Service | 2022-09 |
| D1.2.2 | Transcription Service for Finnish Interviews | 2023-09 |
| D1.3.1 | Corpora of non-standard language | 2022-09 |
| D1.3.2 | System for detecting toxic language | 2023-06 |
| D1.3.3 | Models for retrieving QA pairs from the web | 2023-09 |
| D1.3.4 | QA pair corpora | 2023-12 |
| D2.1.1 | Licensing agreements for personal data | 2022-09 |
| D2.1.2 | Licensing agreements for special categories | 2023-06 |
| D2.2.1 | Speech recognition for L2 | 2022-12 |
| D2.2.2 | Speech recognition for L2 update | 2023-12 |
| D2.3.1 | Licensing interpretation sessions | 2022-12 |
| D2.3.2 | Aligning and retrieving | 2023-12 |
| D2.4.1 | Term discovery procedures | 2022-09 |
| D2.4.2 | Terminology application | 2023-06 |
| D2.4.3.1 | Initializing terminology collections | 2022-09 |
| D2.4.3.2 | Initializing terminology collections | 2023-06 |
| D2.4.3.3 | Initializing terminology collections | 2023-12 |
| D2.5.1 | Test performances storage | 2022-12 |
| D2.5.2 | Analysis and annotation tools for learner performances | 2023-12 |
| D3.1.1 | Initial NLF data | 2022-09 |
| D3.1.2 | Ingestion framework | 2022-12 |
| D3.1.3 | Versioning support | 2023-06 |
| D3.1.4 | Incremental update process | 2023-12 |
| D3.2.1 | Pipeline for transferring archival data | |
| D3.2.2 | Annotation & analysis tools for NARC data | 2023-12 |
| D3.3.1 | Qualitative survey data concept network | 2022-09 |
| D3.3.2 | R package for data concept network |
| D3.4.1 | Livestream data collector | 2022-12 |
| D3.5.1 | Text network analysis of political texts | |
| D3.5.2 | Text network analysis of political texts |
| D4.1.1 | Harmonized FNB | 2022-09 |
| D4.1.2 | Harmonization code | 2022-12 |
| D4.1.3 | Visualisation workflow | 2023-06 |
| D4.1.4 | R/Python module | 2023-12 |
| D4.2.1 | LDF knowledge extraction tools | 2022-12 |
| D4.2.2 | Parliament of Finland Ontology | 2023-12 |
| D4.3.1 | Subsetting tool | 2022-09 |
| D4.3.2 | Statistical overviews and bias detection | 2023-06 |
| D4.3.3 | Representative Twitter dataset | 2023-12 |
| D5.1.1 | User experience questionnaire | 2022-09 |
| D5.1.2 | Log data collection and analysis | 2023-06 |
| D5.1.3 | Protocol for collecting workshop data | 2023-12 |
| D5.2.1 | Actor network | 2022-12 |
| D5.2.2 | Educational material | 2023-12 |
This page outlines the project deliverables for 2024-2025 (completed).
| D1.1.1 | Named-entity annotation | 2024-09 |
| D1.1.2 | Ingesting new unstructured resources | 2025-11 |
| D1.2.1 | Data collection for minority languages | 2024-09 |
| D1.2.2 | Transcription service for minority languages |
| D1.3.1 | Tools and guidelines for video processing | 2025-06 |
| D2.1.1 | Integrate environment for personal data | 2024-09 |
| D2.1.2 | Framework for processing copyrighted data for verification of research |
| D2.2.1 | Transformer training for specialised data | |
| D2.2.2 | Transformer adaptation for specialised data | 2025-12 |
| D2.3.1 | Remote access to text data repositories | |
| D2.3.2 | Remote access to video data repositories | 2025-12 |
| D2.4.1 | Term definition discovery procedures | 2024-09 |
| D2.4.2 | Initializing terminology collections | 2025-12 |
| D3.1.1 | Comprehensive data versioning | 2024-09 |
| D3.1.2 | Workflow automation and version syncing | 2025-09 |
| D3.2.1 | Ingestion of structured data from Finna (NLF) | |
| D3.2.2 | Ingestion of heritage and societal data from Sampo | 2025-06 |
| D3.2.3 | Ingestion of multimodal societal data from the Web | 2025-12 |
| D3.3.1 | Automated metadata of archival data from NAF | |
| D3.3.2 | Automated harmonisation and enrichment of metadata | |
| D3.3.3 | Machine-learning -based enrichment of social media | |
| D3.3.4 | Machine-learning -based enrichment of textual and audio-visual social media contents | 2025-11 |
| D3.3.5 | Forensic linguistics corpus and search interface C.R.I.M.E | 2025-09 |
| D3.3.6 | Reliable image labelling with computer vision | 2025-09 |
| D4.1.1 | Analysis of video stream interactions with AI solutions | |
| D4.1.2 | Analysis Tools for Multimodal Born-digital Social Media | 2024-12 |
| D4.1.3 | Advanced analytic social media tools and data | 2025-12 |
| D4.1.4 | Analysis of multimodal properties of naturalistic speech | 2025-12 |
| D4.1.5 | Analysis of multimodal cultural heritage | 2025-12 |
| D4.1.6 | Enrich survey data with register data and unstructured text | 2025-06 |
| D5.1.1 | Community engagement: multim. societal data researchers | 2024-09 |
| D5.1.2 | Community engagement: multim. heritage researchers | 2025-06 |
| D5.1.3 | Evidence-based infrastructure development | 2024-12 |
| D5.1.4 | Educational resource development | 2025-12 |
The European Summer School in Logic, Language and Information (ESSLLI) on vuodesta 1989 lähtien ollut yksi johtavista monitieteisistä kesäkouluista, joka yhdistää logiikan, kielitieteen, tietojenkäsittelytieteen ja tekoälyn.
ESSLLI kokoaa vuosittain noin 400 osallistujaa eri puolilta Eurooppaa sekä Pohjois- ja Latinalaisesta Amerikasta ja Aasiasta. Tapahtumasta on muodostunut logiikan, kielitieteen ja tietojenkäsittelytieteen nuorille tutkijoille ja opiskelijoille tärkeä kohtaamispaikka, jossa keskustellaan ajankohtaisesta tutkimuksesta ja jaetaan osaamista.
ESSLLI 2026
Date: 3.–14. elokuuta 2026
Location: Praha, Tšekki
ESSLLI 2026 tarjoaa:
Tekoälyn parissa työskenteleville opiskelijoille ja tutkijoille ESSLLI on erityisen hyödyllinen sen vahvan päättelyä, kieliteknologiaa, formaaleja menetelmiä, laskentaa ja kognitiota painottavan sisällön ansiosta.
Mukaan ehtii ilmoittautua edullisemmalla early bird -hinnalla 31.5. asti. Tämän jälkeen ilmoittautuminen jatkuu normaalihinnalla 18.7. asti.
Tutustu ESSLLI 2026 -kesäkouluun alla olevien linkkien kautta (englanniksi):
The European Summer School in Logic, Language and Information (ESSLLI) is one of the leading interdisciplinary summer schools connecting logic, linguistics, computer science, and AI since 1989.
ESSLLI attracts every year around 400 participants from all parts of Europe, as well as from North and Latin America, and Asia. The ESSLLI has become the main meeting place for young researchers and students in logic, linguistics and computer science to discuss current research and to share knowledge.
ESSLLI 2026
Date: August 3–14, 2026
Location: Prague, Czech Republic
ESSLLI 2026 offers:
For students and researchers working in Artificial Intelligence, ESSLLI is especially valuable given the strong focus on reasoning, language technologies, formal methods, computation, and cognition.
If you are considering attending, now is the time to register before the early bird deadline.
Tervetuloa työpajaan Digital Language Sovereignty – Euskadi–Finland AI & Language, joka kokoaa yhteen tutkijoita ja asiantuntijoita Suomesta ja Baskimaasta (Euskadi) keskustelemaan tekoälyn ja kieliteknologian ajankohtaisista kehityssuunnista sekä tulevaisuuden mahdollisuuksista. Tilaisuus on maksuton ja se järjestetään Helsingin yliopiston keskustakampuksella 22.4.2026 klo 9.00-16.00.
Keskustelemme päivän aikana mm. seuraavista teemoista:
Tilaisuus on suunnattu tutkijoille ja jatko-opiskelijoille, kieliasiantuntijoille, julkisen hallinnon edustajille sekä tekoälyn ja kielimallien parissa työskenteleville asiantuntijoille.
Lue lisää täältä: https://www.kielipankki.fi/tapahtumat/digital-language-sovereignty-euskadi-finland-ai-language-workshop/
Ilmoittaudu mukaan tapahtumasivulla: https://euskorpora.eus/en/evento/workshop-digital-language-sovereignty-euskadi-finland-ai-language-workshop/
A Finnish-language volume Sanovat syntaksiksi – Aineistopohjaisia tutkimuksia murteiden lauseopista (”They call it syntax. Data-based approaches to Finnish dialects”) has been published in the series Suomalaisen Kirjallisuuden Seuran Toimituksia. The volume presents recent research on the syntax of Finnish dialects and offers useful information about the history of the field as well as the spoken-language corpora available for research. The work also refers to several resources familiar to users of Kielipankki – the Language Bank of Finland, such as:
The volume also provides interesting background information on how these corpora have been compiled and the various stages they have gone through over the course of their existence.
Sanovat syntaksiksi – Aineistopohjaisia tutkimuksia murteiden lauseopista is openly available in digital form on the SKS (Finnish Literature Society) website: https://doi.org/10.21435/skst.1505
Kielipankin tutkimusjohtaja Krister Lindén oli Svenska Ylen haastateltavana marraskuussa 2025.
Lue uutinen Svenska Ylen sivuilta
Project: FIN-CLARIAH
Grant agreement: Academy of Finland no. 358720
Start date: 2024-01-01
Duration: 24 months
Report author: Jussi Piitulainen (UHEL)
WP 1.1: Report on Ingesting new unstructured resources
Date of reporting: 2024-11-28
Contributors: Jussi Piitulainen, Jyrki Niemi, Jack Rueter, Erik Axelson, Ute Dieckmann, Mietta Lennes, Tommi Jauhiainen (UHEL), Sam Hardwick, Martin Matthiesen (CSC)
Deliverable location: https://www.kielipankki.fi/corpora/
Keywords for the deliverable page: conversion; annotation; interoperability; VRT; UralicUD; Korp; Mink
The Language Bank of Finland receives and obtains text resources in different formats ranging from plain text documents to text enriched with complex annotations and document-level metadata. We aim to ensure that the material is made available to researchers in formats that are usable and interoperable. For text corpora, the Language Bank particularly supports and promotes VRT (VeRticalized Text) as an interchange format by developing, maintaining and utilizing the set of open-source VRT Tools for converting, enriching and ingesting resources containing text. All currently supported formats can be found via the Standards Information System of CLARIN.
The Suomi24 resource group was extended with the discussions from the years 2021–2023 (The Suomi24 Corpus 2021-2023, VRT version, and The Suomi24 Sentences Corpus 2021-2023, Korp version). Moreover, the entire The Suomi24 Sentences Corpus 2001-2023, Korp version and The Suomi 24 Corpus 2001-2023, VRT version now include named-entity and identified-language annotations. The Ylenews resource group was also extended with material from the years 2022-2024, which was made available for download (Yle Finnish News Archive 2022-2024, source). The Korp version of this extension will be published soon.
The Language Bank contributes to the Universal Dependencies (UD) project in order to maintain validity and coverage of the treebanks not only for Finnish but also more generally for Finnic, Finno-Ugric and Uralic languages (Uralic UD). Samples of languages in these groups will also be included in the text resources licensed by the Institute for Bible Translation and in other multilingual text collections that are currently being processed for publication.
In addition to other corpora, the Language Bank participated in publishing several resources prepared by the Ancient Near Eastern Empires (ANEE) research group, including Oracc, Achemenet and Babylonian Administrative and Legal Texts (BALT), available via Korp with linkage from their corresponding lexical networks.
The Trankit toolbox (see Nguyen et al. 2021), a recommended replacement for the old dependency parsers by the Turku NLP group, was installed in the CSC Puhti environment. Trankit was tested to be robust for the kind of morpho-syntactic annotation of pre-segmented Finnish that we need for the existing KLK and Suomi24 corpora. Once adapted for the CWB-VRT format, Trankit would be used to re-annotate the existing corpora with the Universal Dependencies (UD2) features and dependency syntax. Trankit could also be adapted for the segmentation of paragraphs into sentences and tokens, and it adds support for many other languages apart from Finnish.
The Mink platform, developed by Språkbanken Text in Sweden, was test-installed by the Language Bank. Mink allows users to process their own text corpora and to access the result via a private Korp instance. After the new version of the Korp platform is officially published at the Language Bank, it will be possible to make Mink available for wider use by the community. Support for user authentication in Mink is to be added in the year 2026.
The Language Bank participates in the recently launched CLARIN PressMint project that aims to compile a multilingual, comparable, annotated, translated and interoperable set of corpora of European historical newspapers by using a common TEI format. For PressMint, we will transform the out-of-copyright data from our existing KLK corpora (newspapers and magazines from the National Library) from the CWB-VRT format to the appropriate TEI format.
FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Academy of Finland under grant number 358720.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 358720
Start date: 01-01-2024
Duration: 24 months
WP 2.2: Report on Transformer adaptation for specialised data
Date of reporting: 25-11-2025
Report author: Erik Axelson, Jack Rueter (University of Helsinki)
Contributors: Jack Rueter (University of Helsinki), Sam Hardwick, Martin Matthiesen (CSC)
Deliverable location: N/A
In this work package, we aim to provide an MCP server for facilitation of fst-tool and LLM linking for less technically oriented people.
MCP (Model Context Protocol) provides a powerful new opportunity to bring large language model (LLM) capabilities into the research and learning of low-resource languages by creating a bridge between rule-based, finite-state linguistic tools and LLM-based modern chatbots. By hosting HFST [1] analyzers and open-source dictionaries designed and authored by individual humans and teams at GiellaLT [2] and Apertium [3] through UralicNLP [4] libraries on an MCP server, even users with no technical background — and working from a laptop or cellphone — can access lemmatizers, morphological analyzers, and translation dictionaries for dozens of minority languages. This approach opens the door to more inclusive language technology, making advanced tools available to communities that have historically lacked computer-aided support.
We have familiarized ourselves with the use of a local MCP server from a laptop, and have run into memory issues. A so-called free server with a larger memory set at CSC would provide an ideal solution for individual users, as the server would host the model. Some language communities might want to have their specific language data housed as private, i.e., there would have to be different access to this material. The Language Bank of Finland is making plans for the installation of MCP service to allow extensive testing.
[1] HFST – Helsinki Finite-State Technology
[2] GiellaLT – an infrastructure for rule-based language technology aimed at minority and indigenous languages
[3] Apertium – a free/open-source machine translation platform
[4] UralicNLP – an NLP library for Uralic languages
FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 358720.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 358720
Start date: 01-01-2024
Duration: 24 months
WP 2.1: Report on Framework for processing copyrighted data for verification of research
Date of reporting: 28-11-2025
Report authors: Mietta Lennes (UH)
Contributors: Sirpa Kovanen (UH), Krister Lindén(UH), Martin Matthiesen (CSC)
Deliverable location: https://www.kielipankki.fi/support/data-management/dela/
Keywords for the deliverable page: copyrighted data, personal data, social media data, data protection, safeguards
Researchers in Social Sciences and Humanities often need to use data collected from social media platforms. Currently, the reuse of social media data for research purposes is legally challenging. Some part of the content originating from social media is usually protected by copyright or related rights. Social media postings (often including images and videos) may also contain personal data. The terms of use of social media platforms tend to be volatile and non-transparent, and individual permissions cannot be requested due to the large numbers of potential rightholders and data subjects.
Since neither the related EU regulations nor the Finnish legislation are well established in current legal practice, the possibilities for depositing research data from social media must be considered on a case by case basis. It may be possible to archive data obtained from social media and make it available for restricted purposes under certain conditions, according to Section 13 b of the Finnish Copyright Act (i.e., Tekijänoikeuslaki 13 b §), concerning data mining.
Two social media datasets, Finnish presidential elections 2024 in social media (somepressa24), collected by researchers at UHEL, and Nordic Tweet Stream 2013-2023 (nts) collected by a team at UEF, both teams participating in the FIN-CLARIAH project, have been suggested for deposition to the Language Bank of Finland. Using the potential redistribution of these two resources as an example, a review of the current legal risks and restrictions was performed by the legal advisors at UHEL. The negotiations for depositing the first dataset are nearly complete, and the dataset is to be delivered to the Language Bank in December 2025 and to be made available under a RES category license in early 2026. After the first experiences with somepressa24 at UHEL, we aim for a similar deposition agreement with UEF regarding the nts dataset.
The Language Bank of Finland offers frameworks, instructions and technical solutions for deposition agreements and end-user licenses, for access management (the Language Bank Rights system at CSC), and for data encryption or secure processing in a restricted environment if necessary (SD services at CSC). Step-by-step instructions to using the Sensitive Data services (cf. Deliverable 2.1.1.), including the secure SD Desktop environment, are now available both in Finnish and in English for researchers in Social Sciences and Humanities. The Language Bank also collects and shares the links to the privacy notices published by the users of the Language Bank.
FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 358720.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 358720
Start date: 01-01-2024
Duration: 24 months
WP 4.1: Report on Advanced analytic social media tools and data
Date of reporting: 26-11-2025
Report author: Mikko Laitinen (UEF)
Contributors: Masoud Fatemi (UEF), Mehrdad Salimi (UEF)
Deliverable location:
Keywords: social media corpora; social network tools; ego networks; gender
Our work has resulted in building four massive social media corpora from one social media application. The purpose is to enable research access to large-scale and curated social media data, which is often a bottle neck in SSH (Laitinen & Rautionaho 2025). The four datasets are named Digital Social Network Corpora (DSN), as they not only consist of user-generated texts but also of detailed information of people’s social networks. They cover four geographic areas: Australia (DSN Ozzie), the Nordic countries (DSN Nordic), the United Kingdom (DSN British), and the United States (DSN America).
In total, they include 19,345 ego networks, consisting of a central node (ego), its directly connected neighbors (alters), and the connections between the alters. These networks were filtered using a semi-automated method to target what we call genuine human accounts, meaning that we aimed to exclude accounts with unusual network qualities, such as bots, celebrities, politicians, organizations, and businesses. Recreating a comparable dataset to the DSN corpora under the current paid data access policies of the social media application (X) would cost over 3 million euros and take around 58 years, given the current limitations of data access policies.
The resulting datasets are extremely large but contain carefully curated social networks with user-generated textual material. The network datasets contain material from 829,608 users, and the data range from 2006 to 2023. Altogether, they contain more than 700 million messages and nearly 10 billion words keyed in by users.
With their detailed structure, massive size, and coverage over 17 years, the DSN corpora support new research and enable re-examining old questions in the humanities. A case in point is the role of weak ties in the spread of innovations, where prior empirical evidence in sociolinguistics comes from ethnographic observations based on very small networks. One clear limitation of ethnographic network investigations is that participant observation methods are limited to networks of 30–50 individuals. The networks in the DSN corpora are substantially larger and close to average human networks in general, making it possible to investigate a variety of networks of different sizes and structures.
Publications:
Laitinen, Mikko & Paula Rautionaho. 2025. Reuse of social media data in corpus linguistics. International Journal of Corpus Linguistics. doi: 10.1075/ijcl.24136.lai
Masoud Fatemi & Mikko Laitinen. 2025. From tweets to networks: Introducing four large network-based social media corpora. CLARIN Annual Conference Proceedings, 2025. Ed by Cristina Crisot and Thalassia Kontino. Vienna, Austria, 2025. pp. 100–104. (https://www.clarin.eu/sites/default/files/CLARIN2025_ConferenceProceedings.pdf)
Events:
CLARIN 2025 conference Vienna 30 Sept – 2 October 2025 (https://www.clarin.eu/event/2025/clarin-annual-conference-2025)
FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 358720.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 358720
Start date: 01-01-2024
Duration: 24 months
WP 3.3: Report on Machine-learning-based enrichment of textual and audio-visual social media contents
Date of reporting: 20-11-2025
Report authors: Jari Lindroos (JYU), Raine Koskimaa (JYU)
Contributors: Jari Lindroos (University of Jyväskylä), Raine Koskimaa (JYU), Ida Toivanen (JYU), Tanja Välisalo (NAF), Jaakko Peltonen (TAU)
Deliverable locations:
Keywords: video clip analysis; multimodal; MLLM; video summarization; data enrichment; Twitch
The proliferation of short-form video on livestreaming platforms like Twitch presents a significant challenge for multimodal content analysis. Each clip contains a vast amount of diverse information: the visual action, the auditory context from caster commentary, and the text-based reactions from the live chat, all representing dense and valuable data for understanding online communities. However, the sheer volume and complexity of this data creates a need for efficient analysis tools. Our previous tools have focused on chat-analysis or chat content detection [1, 2].
This deliverable presents a continuation of the deliverable D4.1.1 tool for the automated understanding and enrichment of such clips. The tool is powered by state-of-the-art Multimodal Large Language Models (MLLMs) from the Google Gemini family, guided by a multi-step Chain-of-Thought prompt. This prompt instructs the MLLM to focus on data enrichment, systematically analyzing the clip’s metadata, audio-visual content, and chat log, producing a JSON file.
This structured JSON data is organized into three parts. The analysis begins with the audiovisual analysis of the content in the video. It identifies all key entities involved, logs chronological actions in the video, transcribes the on-screen text, and breaks down caster commentary into key quotes and emotional tones. Next, the “chat reaction” section shows how the audience reacted to the jargon used by the community while also providing a glossary to explain the cultural meaning behind this. Finally, the “causal synthesis” connects these two modalities. It provides a narrative summary explaining why the clip matters and establishes direct causal links between the audiovisual triggers to the exact chat reactions they caused.
All generated analyses are automatically saved and accessible within the video_descriptions category of the data viewer section.
Publications
[1] Jari Lindroos, Jaakko Peltonen, Tanja Välisalo, Raine Koskimaa, and Ida Toivanen. ”From PogChamps to Insights: Detecting Original Content in Twitch Chat.” In Hawaii International Conference on System Sciences, pp. 2542-2551. Hawaii International Conference on System Sciences, 2025. https://doi.org/10.24251/hicss.2025.308
[2] Jari Lindroos, Ida Toivanen, Jaakko Peltonen, Tanja Välisalo, Raine Koskimaa, and Sami Äyrämö. ”Participant profiling on Twitch based on chat activity and message content.” In International GamiFIN Conference, pp. 18-29. CEUR Workshop Proceedings, 2025. https://ceur-ws.org/Vol-4012/paper18.pdf
FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 358720.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 358720
Start date: 01-01-2024
Duration: 24 months
WP 3.2: Report on Ingestion of multimodal societal data from the Web
Date of reporting: 20-11-2025
Report authors: Matti Nelimarkka (University of Helsinki), Jari Lindroos (JYU), Raine Koskimaa (JYU)
Contributors: Matti Nelimarkka (University of Helsinki), Denis Davydov (University of Helsinki), Anita Braida (University of Helsinki), Jari Lindroos (University of Jyväskylä), Raine Koskimaa (JYU), Ida Toivanen (JYU), Tanja Välisalo (NAF), Jaakko Peltonen (TAU)
Deliverable locations:
Keywords for the deliverable page: Twitch, YouTube, chat data, video data
This deliverable focuses on infrastructures for acquisition of multimodal and societal data harvested from the web. The task includes the implementation and maintenance of data collection tools for most popular Finnish discussion forums, YouTube, and Twitch. This deliverable contains two parts ⎯ part A conducted by the Centre for Social Data Science, University of Helsinki and part B by the University of Jyväskylä.
PART A: FINNISH DISCUSSION FORUMS
To ensure that researchers have access beyond global platforms (where data collection is a shared global concern) University of Helsinki build and maintain forum scrapers which extract the content to user-generated content including vauva.fi, kaksplus.fi and comments on yle.fi and hs.fi. These can be used through a command line interface which produces the content as a CSV file for further analysis. We also provided modifications to the 4CAT platform (https://4cat.nl/) to ensure it correctly operates with Finnish language.
PART B: YOUTUBE CHAT COLLECTOR & TWITCH VIDEO COLLECTOR
The team from the University of Jyväskylä presents a continuation of the deliverable for the Twitcher data collector tool. We present new added features such as the option to collect chat data from YouTube from either live or past broadcasts. The collected YouTube chat data can also be viewed in the data viewer section and are also automatically saved in CSC Allas. We also implemented the option to collect videos from Twitch past broadcasts in regard to the video clip analysis tool presented in D3.3.4 and D4.1.1.
FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 358720.
Project: FIN-CLARIAH
Grant agreement: Research Council of Finland no. 358720
Start date: 01-01-2024
Duration: 24 months
WP 4.1: Report on Analysis of multimodal cultural heritage
Date of reporting: 20-11-2025
Report author: Ilkka Lähteenmäki (University of Oulu)
Contributor: Ilkka Lähteenmäki (University of Oulu)
Deliverable location: 10.5281/zenodo.17700648
This paper examines if historians and cultural heritage researchers can justifiably depend on multimodal AI systems for accessing large visual collections from social epistemology point of view. Building on Inkeri Koskinen’s “necessary trust view” and Jakob Ortmann’s account of task-specific epistemic reliance, it argues that digital history and cultural heritage form a non-typical setting for current social epistemology of AI. In contrast to the physical sciences, where AI tools such as AlphaFold are embedded in long-standing evaluation regimes and well-defined tasks, historical research involves open-ended, exploratory questions, fuzzy and historically shifting concepts, and interpretive practices centred on individual researchers and small teams.
The paper uses examples from recent proposals for using multimodal AI for text-to-image, image-to-text and image-to-image retrieval, and for AI-assisted metadata generation and “distant viewing” of images. It shows how hopes for a multimodal turn in digital humanities confronts the essential epistemic opacity of deep neural networks and the difficulty of evaluating reliability for complex open ended retrieval tasks. Three suggested mitigation strategies are discussed: critical analyses of models and training data; historically informed reflection on bias and concept change; and fine-tuning or post-processing of models for specific purposes. From a social epistemology perspective, each strategy encounters limits when generalised to research infrastructure meant to support many corpora, tasks and user communities.
The paper then turns to approaches that argue for using multimodality theory to design metadata schemas and guide AI-based annotation. It shows how this is a attempt to shift epistemic trust from AI systems back to scholars (at least partially) in effort to make use of the developing technology. However, this brings into discussion old debates of between theories of meaning. Especially with image data the theoretical discussion of how images meanings should be established and if these theories are implementable to computational models need to be explored. Couple examples from contemporary photography and medieval manuscript research illustrate both the potential of AI-supported exploration and the need for additional contextual and theoretical work to render outputs historically interpretable.
The central claim is that, given the essential epistemic opacity of AI, it currently looks like justified epistemic dependence in history and cultural heritage research needs be organised around situated, task-specific, and accountable uses of multimodal models rather than general-purpose models. The options for research infrastructures for establishing trust are therefore focus on building mechanisms for task-specific reliability assessment, or embedding trusted identifiable human agents or institutions between users and models.
FIN-CLARIAH project has received funding from the European Union – NextGenerationEU instrument and is funded by the Research Council of Finland under grant number 358720.
