North Saami Corpus (Literature) (UHLCS)

The North Saami Corpus contains Kerttu Vuolab’s novel Cheppari cháráhus written in Northern Sami. The corpus is a part of the UHLCS corpus collection.

UHLCS has many different IPR holders. Should you have any questions regarding the collection, please contact Pirkko Suihkonen (suihkonen.pirkko@gmail.com).

Latest versions/subcorpora:
North Saami Corpus (Literature) (UHLCS)
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
The data is available upon request via CSC’s computing environment
North Saami Corpus (Literature) (UHLCS), Helsinki Korp Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
The resource will be available soon
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021061605

Finnish Corpus (Literature) (UHLCS)

The Finnish Corpus is a part of the UHLCS corpus collection.

UHLCS has many different IPR holders. Should you have any questions regarding the collection, please contact Pirkko Suihkonen (suihkonen.pirkko@gmail.com).

Latest versions/subcorpora:
Finnish Corpus (Literature) (UHLCS)
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
The data is available upon request via CSC’s computing environment
Finnish Corpus (Literature) (UHLCS), Helsinki Korp Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
The resource will be available soon
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021061604

Testipiste Corpus

Testipiste is a language assessment centre for adult migrants. This corpus contains texts written by 2397 different persons, 3 texts from each person. It also contains assignments and other related texts. The essays contain i.a. information on the starting level of their authors, as defined by Testipiste.

Latest versions/subcorpora:
Testipiste Corpus, source
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
The resource will be available soon
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021061603

DIALUKI – Diagnosing reading and writing in a second or foreign language

The project studies the diagnosis of reading and writing abilities in a second or foreign language. It seeks to identify the cognitive features which predict a learner’s strengths and weaknesses in those areas. The project brings together scholars from applied linguistics, psychology and assessment to engage in multidisciplinary work and to develop innovative ways of diagnosing the development of second and foreign language abilities.

More information on the corpus: https://www.jyu.fi/dialuki

Latest versions/subcorpora:
DIALUKI – Diagnosing reading and writing in a second or foreign language
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
The resource will be made available in Korp
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021061602

CEFLING Project Corpus

Finnish as a second language and English as a foreign language writing performances collected from comprehensive school students (grades 7 – 9) in the project CEFLING – Linguistic Basis of the Common European Framework for L2 English and L2 Finnish. Data from several hundred learners; 4-5 writing tasks from each learner; background information, self-assessments of proficiency.

More information:
https://www.jyu.fi/hytk/fi/laitokset/kivi/tutkimus/hankkeet/paattyneet-tutkimushankkeet/cefling/en/cefling

Latest versions/subcorpora:
CEFLING Project Corpus
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
The resource will be available soon
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021061601

Corpus of Historical American English

The Corpus of Historical American English (COHA) contains about 385 million words and 115 000 texts from the years 1810-2009. Each decade has roughly the same balance of fiction, popular magazine, newspaper, and non-fiction books.

For general terms and conditions for this and other corpora from BYU please see https://www.corpusdata.org/restrictions.asp

More information on the BYU corpora at Kielipankki

Latest versions/subcorpora:
Corpus of Historical American English – Kielipankki Korp version 2017H1
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Corpus of Historical American English – Kielipankki download version 2017H1
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2017061924

Corpus of Contemporary American English

The Corpus of Contemporary American English (COCA) contains about 440 million words and 190 000 texts from the years 1990-2012. The corpus is evenly divided into spoken, fiction, magazine, newspaper, academic genres (~88 million words each).

For general terms and conditions for this and other corpora from BYU please see https://www.corpusdata.org/restrictions.asp

More information on the BYU corpora at Kielipankki

Latest versions/subcorpora:
Corpus of Contemporary American English – Kielipankki Korp version 2017H1
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Corpus of Contemporary American English – Kielipankki download version 2017H1
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2017061921

Corpus of Global Web-Based English

The Corpus of Global Web-Based English (GloWbE) contains about 1.8 billion words and 1 800 000 texts from web pages in the United States, Great Britain, Australia, India, and 16 other countries. About 60 % of the texts come from blogs.

For general terms and conditions for this and other corpora from BYU please see https://www.corpusdata.org/restrictions.asp

More information on the BYU corpora at Kielipankki

Latest versions/subcorpora:
Corpus of Global Web-Based English – Kielipankki Korp version 2017H1
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Corpus of Global Web-Based English – Kielipankki download version 2017H1
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2017061927

SFNET Corpus

The corpus contains written discussion in the SFNET Internet discussion forum in Finnish from 2002-2003.

Latest versions/subcorpora:
SFNET Corpus
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
SFNET Corpus, Helsinki Korp Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Resource will be available soon
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021052501

The Corpus of Beserman Udmurt, Kielipankki Version

The Corpus of Beserman Udmurt comprises 65 000 tokens. The Beserman dialect of Udmurt is used in daily communication approximately by 2 000 speakers (according to the 2010 census). The Beserman live in the basin of the Cheptsa river in the Republic of Udmurtia and in the Kirov Oblast of the Russian Federation. In the scientific literature Beserman is considered to be a dialect of the Udmurt language which is characterized by an unusual combination of specifically Beserman phenomena (concentrated in vocabulary and phonetics) with certain traits of Northern and Southern Udmurt dialects, mostly morphological and phonological. The dialect remains the main means of everyday communication in Beserman villages, at least for the older generation.

The texts contained in the corpus have been collected in the villages of Shamardan (109 texts of 117), Vortsa (4 of 117), Malaya Yunda (1 of 117) and Zhuvam (3 of 117) in the Republic of Udmurtia in the years 2003-2015. There are 33 informants in total. The texts have been recorded, transcribed and grammatically annotated in the SIL FieldWorks software. The corpus contains narratives, life stories, dialogues, recipes, and recordings of psycholinguistic experiments. Each sentence is provided with interlinear glossing (according to the Leipzig Glossing Rules) and translation. Both the full text version with audio files and the corpus version are available at http://beserman.ru/corpus/search/?interface_language=en

Latest versions/subcorpora:  
The Corpus of Beserman Udmurt, Kielipankki Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Search for all versions in META-SHARE  

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021052406

Corpus of Age-related Voice Disguise

This corpus includes normal and age-related disguised speech uttered by 60 native Finnish speakers (31 females and 29 males). The speakers were asked to read the same text fragments several times, in their modal voice and in two disguised voices, first pretending to be an elderly speaker and then pretending to be a child. The texts consisted of the Finnish translations of The Rainbow Passage and The North Wind and the Sun, and two selected English sentences from the TIMIT[1] corpus (SA1, SA2). The corpus includes samples of 78 different sentences per speaker (66 Finnish, 12 English). The speech was recorded simultaneously with a portable recorder with close-talking microphone, and two smartphones applications, yielding a total of 14040 audio files (3 * 4680). The material was recorded in summer 2015 in order to study the effect of voice disguise on automatic speaker recognition.

Data protection policy for this corpus: http://urn.fi/urn:nbn:fi:lb-2018121021

Guidelines for processing corpora containing personal data in the Language Bank of Finland: http://urn.fi/urn:nbn:fi:lb-2020081522

Latest versions/subcorpora:
Corpus of Age-related Voice Disguise
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021052405

ArkiSyn Database of Finnish Conversational Discourse

The Arkisyn corpus contains Finnish everyday conversations which have been morphologically and syntactically annotated. The data comes from the Conversation Analysis Archive at the University of Helsinki and the Finnish language Recording Archive at the University of Turku.

Latest versions/subcorpora:
ArkiSyn Database of Finnish Conversational Discourse, Helsinki Korp Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2014073026

The Helsinki Korp JRC-Acquis Bilingual Parallel Corpora

The Helsinki Korp JRC-Acquis Bilingual Parallel Corpora are:

The Helsinki Korp JRC-Acquis Finnish-English Corpus
The Helsinki Korp JRC-Acquis Finnish-Swedish Corpus
The Helsinki Korp JRC-Acquis Finnish-German Corpus
The Helsinki Korp JRC-Acquis Finnish-French Corpus
The Helsinki Korp JRC-Acquis Finnish-Spanish Corpus
The Helsinki Korp JRC-Acquis Finnish-Italian Corpus
The Helsinki Korp JRC-Acquis Finnish-Estonian Corpus
The Helsinki Korp JRC-Acquis Finnish-Hungarian Corpus
The Helsinki Korp JRC-Acquis Finnish-Polish Corpus

The corpora contain texts of the JRC-Acquis Multilingual Parallel Corpus. The Acquis Communautaire (AC) is the total body of European Union (EU) law applicable in the the EU Member States.

For more information on the JRC-Acquis Multilingual Parallel Corpus see http://urn.fi/urn:nbn:fi:lb-20140730162 or https://ec.europa.eu/jrc/en/language-technologies/jrc-acquis

Latest versions/subcorpora:
The Helsinki Korp JRC-Acquis Bilingual Parallel Corpora
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
The Finnish Sub-corpus of the JRC-Acquis Multilingual Parallel Corpus
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
The Finnish Sub-corpus of the JRC-Acquis Multilingual Parallel Corpus, Downloadable Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021052404

The Helsinki Korp Europarl Bilingual Corpora

The Helsinki Korp Europarl Bilingual Corpora are:

The Helsinki Korp Europarl Finnish-English Corpus
The Helsinki Korp Europarl Finnish-Swedish Corpus
The Helsinki Korp Europarl Finnish-German Corpus
The Helsinki Korp Europarl Finnish-French Corpus
The Helsinki Korp Europarl Finnish-Spanish Corpus
The Helsinki Korp Europarl Finnish-Estonian Corpus

The corpora contain texts of the Europarl Parallel Corpus v7.

The Europarl parallel corpus is extracted from the proceedings of the European Parliament. The goal of the extraction and processing was to generate sentence aligned text for statistical machine translation systems. For this purpose matching items were extracted and labeled with corresponding document IDs. By using a preprocessor, sentence boundaries were identified. The data was sentence aligned by using a tool based on the Church and Gale algorithm.

For more information on the Europarl Parallel Corpus see http://urn.fi/urn:nbn:fi:lb-20140730195 and http://www.statmt.org/europarl/

Latest versions/subcorpora:
The Helsinki Korp Europarl Bilingual Corpora
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021052403

Opus, Helsinki Korp Version

The Helsinki Korp version of the Opus open parallel corpus (http://opus.lingfil.uu.se/), containing scrambled sentences, has been published in Kielipankki.

Latest versions/subcorpora:
Opus, Helsinki Korp Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Search for all versions in META-SHARE

The subcorpora of Opus, Helsinki Korp Version are:

OPUS Finnish–Czech
OPUS Finnish–Danish
OPUS Finnish–Dutch
OPUS Finnish–English
OPUS Finnish–Estonian
OPUS Finnish–French
OPUS Finnish–German
OPUS Finnish–Greek
OPUS Finnish–Hungarian
OPUS Finnish–Italian
OPUS Finnish–Polish
OPUS Finnish–Portuguese
OPUS Finnish–Russian
OPUS Finnish–Swedish
OPUS Finnish–Spanish
OPUS Finnish–Turkish

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021052402

The Karelian Finnish Newspaper Corpus

The corpus contains issues of the ’Karjalan Sanomat’ newspaper published in 2012-2014.

Latest versions/subcorpora:
The Karelian Finnish Newspaper Corpus
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021052401

The HS.fi News and Comments Corpus

The HS.fi News and Comments Corpus contains the domestic news of the Helsingin Sanomat website and their comments from 5.9.2011 to 4.9.2012. The corpus starts with the first news of 5.9.2011 and ends with a news published in the morning on 3.9.2012 and the comments published on the website by 5.9.2012.

Latest versions/subcorpora:
The HS.fi News and Comments Corpus
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021051910

AI2D-RST: A multimodal corpus of 1000 primary school science diagrams

AI2D-RST is a multimodal corpus of 1000 English-language diagrams that represent topics in primary school natural science, such as food webs, life cycles, moon phases and human physiology. The corpus is based on the Allen Institute for Artificial Intelligence Diagrams (AI2D) dataset, a collection of diagrams with crowd-sourced descriptions.

Building on the layout segmentation in AI2D, the AI2D-RST corpus presents a multi-layer annotation schema that provides a rich, graph-based description of diagram structure. The annotation was performed by trained experts.

Latest versions/subcorpora:
AI2D-RST: A multimodal corpus of 1000 primary school science diagrams version 1.1
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021051909

A Multimodal Corpus of Tourist Brochures Produced by the City of Helsinki, Finland (1967-2008)

This multimodal corpus, which consists of the tourist brochures produced by the city of Helsinki, Finland, is fully annotated using XML schema provided for the Genre and Multimodality (GeM) model. The GeM model is used to describe the content, layout, graphic and typographic appearance, and rhetorical structure of 58 double-pages published between 1967 and 2008.

Latest versions/subcorpora:
A Multimodal Corpus of Tourist Brochures Produced by the City of Helsinki, Finland (1967-2008)
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are (or might be in the future) published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021051908

The Advanced Finnish Learners’ Corpus

The Advanced Finnish Learners’ Corpus (in Finnish Edistyneiden suomenoppijoiden korpus) consists mainly of texts written by non-native MA students of Finnish language. At the end of 2009 it consisted of the following:

– digitalized exam essays,
– digitalized theses,
– other academic writings digitalized.

The subcorpora containing digitalized exam essays (esseet) and course papers (tentit) have been made available at http://korp.csc.fi/

Important: Due to the nature of the material, the resource should be handled with care in order to respect the privacy of the personal data. If samples of the data are published, they must be anonymized according to best practices.

More information on the corpus: https://www.utu.fi/fi/yliopisto/humanistinen-tiedekunta/suomen-kieli-ja-suomalais-ugrilainen-kielentutkimus/lauseopin-arkisto

Latest versions/subcorpora:
The Advanced Finnish Learners’ Corpus
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Select the corpus in Korp
The Advanced Finnish Learners’ Corpus, Downloadable Version
icon-info-circle Metadata and license
icon-quote-right Attribution instructions
Download the resource
Search for all versions in META-SHARE

Of this language corpus different versions/subcorpora are published in the Language Bank of Finland. The versions are available through the Language Bank Download Service and/or through the Korp concordance tool. The links to the different versions can be found from the list above.

Detailed information on the content of each version, user rights and licenses can be found from it’s specific metadata record in META-SHARE.

This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021051907