FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements (with L1, L2 and Imitated L2 Accents) (frost-ema)

Suomeksi


Currently available versions of this resource

ShortnameName and metadataLicenseLocationCiteResource group and helpApplyPublication yearSupport levelGroup NameClarin Resource Family En
ShortnameName and metadataLicenseLocationCiteResource group and helpApplyPublication yearSupport levelGroup NameClarin Resource Family En

Upcoming versions of this resource

These resource versions are not yet available in the Language Bank of Finland.

ShortnameName and metadataLicenseFormatsSupport levelContact PersonResource group and helpLocationOther informationGroup Name
ShortnameName and metadataLicenseFormatsSupport levelContact PersonResource group and helpLocationOther informationGroup Name

Resource information

The dataset contains speech recordings from 18 bilingual speakers of Finnish and Russian, as well as articulation movements, which were recorded using EMA equipment in the Logolab facility at the University of Eastern Finland.
The data includes recordings of the speakers’ speech in both their first language (L1) and second language (L2), as well as speech delivered with an imitated foreign accent. The speakers’ task was to read a rainbow story, word lists (‘say [x] again’) and to describe, in their own words, the picture they had been given.
In addition to the audio data, the dataset includes articulation measurements produced using the AG5012 equipment from Carstens Medizinelektronik GmbH. In practice, these are raw text data files (in .pos format) containing X, Y and Z coordinates (as well as other data provided by the equipment). The speech data is also to be segmented automatically (forced alignment).
The data was collected between 2023 and 2025.

References

Hopponen, S., Kinnunen, T., Nikolaev, A., González Hautamäki, R., Tavi, L., Meister, E. (2025) FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents. Proc. Interspeech 2025, 364-368, doi: 10.21437/Interspeech.2025-1853

License and access

  • Some versions of this resource are available publicly (PUB), whereas others require you to log in as an academic user (ACA) or to apply for individual access rights (RES).
  • Click on the license image to see the resource-specific license text.
  • (Some/all versions of this resource may contain personal data (license condition +PRIV). The license may then include additional data protection terms and conditions that you must follow. If processing personal data, maintain a public Privacy Notice regarding your project and provide the link to the Language Bank of Finland, see instructions.)
  • (Some versions of this resource are available in the computing environment (see column ’Location’). icon-question-circle)

 

 


This page has a persistent identifier: http://urn.fi/urn:nbn:fi:lb-2026050603

Last modified on 2026-09-02