Research Article

Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View

Volume: 10 Number: 1 June 30, 2026
EN

Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View

Abstract

Datasets play a crucial role in advancing NLP tasks by providing resources for training and evaluating models. These datasets serve as the foundation for developing models that can understand and generate human language. However, little research has been performed on creating datasets in low-resource languages, such as Turkish, compared to English. This research focuses on paraphrase datasets, which is a limited area of study in NLP. To the best of our knowledge, nine Turkish paraphrase datasets are available. Accordingly, we provide a comprehensive review of Turkish paraphrase datasets by examining their scope, design, and linguistic properties. The pros and cons of the datasets are discussed, including their size, variety, and domain coverage limitations. Suggestions are presented for a more diverse, large-scale, and expandable Turkish paraphrase corpora that could better represent the language’s linguistic richness.

Keywords

References

  1. Alkurdi, B., Sarıoğlu, H. Y., & Amasyalı, M. F. (2022). Semantic similarity based filtering for Turkish paraphrase dataset creation. In M. Abbas & A. A. Freihat (Eds.), Proceedings of the 5th International Conference on Natural Language and Speech Processing (ICNLSP 2022) (pp. 119-127). Association for Computational Linguistics. google scholar
  2. Alkurdi, B., Sarıoğlu, H. Y., & Amasyalı, M. F. (2024). Scaling Up Paraphrase Generation Datasets with Machine Translation and Semantic Similarity Filtering. In M. Abbas (Ed.), Practical Solutions for Diverse Real-World NLP Applications (pp. 21-36). Springer International Publishing. 10.1007/978-3-031-44260-5_2 google scholar
  3. Bağcı, A., & Amasyalı, M. F. (2021). Comparison of Turkish Paraphrase Generation Models. In 2021 International Conference on INnovations in Intelligent SysTems and Applications (INISTA) (pp. 1-6). IEEE. https://doi.org/10.1109/INISTA52262.2021.954833510.1109/INISTA52262.2021.9548335 google scholar
  4. Bhagat, R., & Hovy, E. (2013). What is a paraphrase? Computational Linguistics, 39(3), 463-472. https://doi.org/10.1162/COLI_a_00166 google scholar
  5. Biber, D. (1993). Representativeness in Corpus Design. Literary and Linguistic Computing, 8(4), 243-257. https://doi.org/10.1093/llc/8.4.243https://doi.org/10.1093/llc/8.4.243 google scholar
  6. Biber, D., Conrad, S., & Reppen, R. (1998). Corpus Linguistics: Investigating Language Structure and Use. https://doi.org/10.1017/CBO9780511804489 google scholar
  7. Binbir, M., & Aksoy, S. (2021). Paraphrase generation for Turkish language. Retrieved Aralık 22, 2024, from https://github.com/sercaksoy/MT5-Turkish-Paraphrase-Generation/blob/main/paper.pdf google scholar
  8. Cambridge University Press. (n.d.). Corpus. Retrieved Aralık 5, 2024, from https://dictionary.cambridge.org/dictionary/english/corpus google scholar

Details

Primary Language

English

Subjects

Natural Language Processing

Journal Section

Research Article

Publication Date

June 30, 2026

Submission Date

May 10, 2025

Acceptance Date

February 12, 2026

Published in Issue

Year 2026 Volume: 10 Number: 1

APA
Teker, G., & Koşaner, Ö. (2026). Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View. Acta Infologica, 10(1), 80-96. https://doi.org/10.26650/acin.1696617
AMA
1.Teker G, Koşaner Ö. Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View. ACIN. 2026;10(1):80-96. doi:10.26650/acin.1696617
Chicago
Teker, Görkem, and Özgün Koşaner. 2026. “Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View”. Acta Infologica 10 (1): 80-96. https://doi.org/10.26650/acin.1696617.
EndNote
Teker G, Koşaner Ö (June 1, 2026) Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View. Acta Infologica 10 1 80–96.
IEEE
[1]G. Teker and Ö. Koşaner, “Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View”, ACIN, vol. 10, no. 1, pp. 80–96, June 2026, doi: 10.26650/acin.1696617.
ISNAD
Teker, Görkem - Koşaner, Özgün. “Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View”. Acta Infologica 10/1 (June 1, 2026): 80-96. https://doi.org/10.26650/acin.1696617.
JAMA
1.Teker G, Koşaner Ö. Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View. ACIN. 2026;10:80–96.
MLA
Teker, Görkem, and Özgün Koşaner. “Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View”. Acta Infologica, vol. 10, no. 1, June 2026, pp. 80-96, doi:10.26650/acin.1696617.
Vancouver
1.Görkem Teker, Özgün Koşaner. Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View. ACIN. 2026 Jun. 1;10(1):80-96. doi:10.26650/acin.1696617