Comprehensive Review of Turkish Paraphrase Datasets: A Linguistics View
Abstract
Datasets play a crucial role in advancing NLP tasks by providing resources for training and evaluating models. These datasets serve as the foundation for developing models that can understand and generate human language. However, little research has been performed on creating datasets in low-resource languages, such as Turkish, compared to English. This research focuses on paraphrase datasets, which is a limited area of study in NLP. To the best of our knowledge, nine Turkish paraphrase datasets are available. Accordingly, we provide a comprehensive review of Turkish paraphrase datasets by examining their scope, design, and linguistic properties. The pros and cons of the datasets are discussed, including their size, variety, and domain coverage limitations. Suggestions are presented for a more diverse, large-scale, and expandable Turkish paraphrase corpora that could better represent the language’s linguistic richness.
Keywords
References
- Alkurdi, B., Sarıoğlu, H. Y., & Amasyalı, M. F. (2022). Semantic similarity based filtering for Turkish paraphrase dataset creation. In M. Abbas & A. A. Freihat (Eds.), Proceedings of the 5th International Conference on Natural Language and Speech Processing (ICNLSP 2022) (pp. 119-127). Association for Computational Linguistics. google scholar
- Alkurdi, B., Sarıoğlu, H. Y., & Amasyalı, M. F. (2024). Scaling Up Paraphrase Generation Datasets with Machine Translation and Semantic Similarity Filtering. In M. Abbas (Ed.), Practical Solutions for Diverse Real-World NLP Applications (pp. 21-36). Springer International Publishing. 10.1007/978-3-031-44260-5_2 google scholar
- Bağcı, A., & Amasyalı, M. F. (2021). Comparison of Turkish Paraphrase Generation Models. In 2021 International Conference on INnovations in Intelligent SysTems and Applications (INISTA) (pp. 1-6). IEEE. https://doi.org/10.1109/INISTA52262.2021.954833510.1109/INISTA52262.2021.9548335 google scholar
- Bhagat, R., & Hovy, E. (2013). What is a paraphrase? Computational Linguistics, 39(3), 463-472. https://doi.org/10.1162/COLI_a_00166 google scholar
- Biber, D. (1993). Representativeness in Corpus Design. Literary and Linguistic Computing, 8(4), 243-257. https://doi.org/10.1093/llc/8.4.243https://doi.org/10.1093/llc/8.4.243 google scholar
- Biber, D., Conrad, S., & Reppen, R. (1998). Corpus Linguistics: Investigating Language Structure and Use. https://doi.org/10.1017/CBO9780511804489 google scholar
- Binbir, M., & Aksoy, S. (2021). Paraphrase generation for Turkish language. Retrieved Aralık 22, 2024, from https://github.com/sercaksoy/MT5-Turkish-Paraphrase-Generation/blob/main/paper.pdf google scholar
- Cambridge University Press. (n.d.). Corpus. Retrieved Aralık 5, 2024, from https://dictionary.cambridge.org/dictionary/english/corpus google scholar
Details
Primary Language
English
Subjects
Natural Language Processing
Journal Section
Research Article
Publication Date
June 30, 2026
Submission Date
May 10, 2025
Acceptance Date
February 12, 2026
Published in Issue
Year 2026 Volume: 10 Number: 1