FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements (with L1, L2 and Imitated L2 Accents) (frost-ema)
Suomeksi
Currently available versions of this resource
| Shortname | Name and metadata | License | Location | Cite | Resource group and help | Apply | Publication year | Support level | Group Name | Clarin Resource Family En |
|---|
| Shortname | Name and metadata | License | Location | Cite | Resource group and help | Apply | Publication year | Support level | Group Name | Clarin Resource Family En |
|---|
Upcoming versions of this resource
These resource versions are not yet available in the Language Bank of Finland.
| Shortname | Name and metadata | License | Formats | Support level | Contact Person | Resource group and help | Location | Other information | Group Name |
|---|
| Shortname | Name and metadata | License | Formats | Support level | Contact Person | Resource group and help | Location | Other information | Group Name |
|---|
Resource information
The dataset contains speech recordings from 18 bilingual speakers of Finnish and Russian, as well as articulation movements, which were recorded using EMA equipment in the Logolab facility at the University of Eastern Finland.
The data includes recordings of the speakers’ speech in both their first language (L1) and second language (L2), as well as speech delivered with an imitated foreign accent. The speakers’ task was to read a rainbow story, word lists (‘say [x] again’) and to describe, in their own words, the picture they had been given.
In addition to the audio data, the dataset includes articulation measurements produced using the AG5012 equipment from Carstens Medizinelektronik GmbH. In practice, these are raw text data files (in .pos format) containing X, Y and Z coordinates (as well as other data provided by the equipment). The speech data is also to be segmented automatically (forced alignment).
The data was collected between 2023 and 2025.
References
Hopponen, S., Kinnunen, T., Nikolaev, A., González Hautamäki, R., Tavi, L., Meister, E. (2025) FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents. Proc. Interspeech 2025, 364-368, doi: 10.21437/Interspeech.2025-1853
License and access
- Some versions of this resource are available publicly (PUB), whereas others require you to log in as an academic user (ACA) or to apply for individual access rights (RES).
- Click on the license image to see the resource-specific license text.
- (Some/all versions of this resource may contain personal data (license condition +PRIV). The license may then include additional data protection terms and conditions that you must follow. If processing personal data, maintain a public Privacy Notice regarding your project and provide the link to the Language Bank of Finland, see instructions.)
- (Some versions of this resource are available in the computing environment (see column ’Location’). icon-question-circle)
This page has a persistent identifier: http://urn.fi/urn:nbn:fi:lb-2026050603
Last modified on 2026-09-02