University of Oulu

Rantanen, T., Tolvanen, H., Roose, M., Ylikoski, J., & Vesakoski, O. (2022). Best practices for spatial language data harmonization, sharing and map creation—A case study of Uralic. PLOS ONE, 17(6), e0269648.

Best practices for spatial language data harmonization, sharing and map creation : a case study of Uralic

Saved in:
Author: Rantanen, Timo1; Tolvanen, Harri1; Roose, Meeli1;
Organizations: 1Department of Geography and Geology, University of Turku, Turku, Finland
2Giellagas Institute for Saami Studies, University of Oulu, Oulu, Finland
3Department of Biology, University of Turku, Turku, Finland
4Department of Finnish language and Finno-Ugric Linguistics, University of Turku, Turku, Finland
Format: article
Version: published version
Access: open
Online Access: PDF Full Text (PDF, 3.6 MB)
Persistent link:
Language: English
Published: Public Library of Science, 2022
Publish Date: 2023-02-09


Despite remarkable progress in digital linguistics, extensive databases of geographical language distributions are missing. This hampers both studies on language spatiality and public outreach of language diversity. We present best practices for creating and sharing digital spatial language data by collecting and harmonizing Uralic language distributions as case study. Language distribution studies have utilized various methodologies, and the results are often available as printed maps or written descriptions. In order to analyze language spatiality, the information must be digitized into geospatial data, which contains location, time and other parameters. When compiled and harmonized, this data can be used to study changes in languages’ distribution, and combined with, for example, population and environmental data. We also utilized the knowledge of language experts to adjust previous and new information of language distributions into state-of-the-art maps. The extensive database, including the distribution datasets and detailed map visualizations of the Uralic languages are introduced alongside this article, and they are freely available.

see all

Series: PLoS one
ISSN: 1932-6203
ISSN-E: 1932-6203
ISSN-L: 1932-6203
Volume: 17
Issue: 6
Article number: e0269648
DOI: 10.1371/journal.pone.0269648
Type of Publication: A1 Journal article – refereed
Field of Science: 6121 Languages
Funding: TR was financially supported by the University of Turku Graduate School (UTU-BGG);, Kone Foundation (UraLex and personal grant);, Finno-Ugrian Society;, and UIT – The Arctic University of Norway; OV was funded by Kone Foundation (SumuraSyyni, AikaSyyni);, and the Academy of Finland, grant number 329257; MR was funded by the Academy of Finland, grant number 329257; The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Dataset Reference: Geographical database of the Uralic Languages (S2 Appendix) is available from All other relevant data are within the paper and its Supporting Information files.
Copyright information: © 2022 Rantanen et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.