Research Article

Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts

Volume: 9 Number: 1 September 17, 2026

Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts

Abstract

AI text detectors are used in academic integrity settings. Their validity may change across time, providers, and writing conditions. This study examined AI-text detection in Turkish academic writing using a leakage-audited historical corpus and a prespecified provider-known benchmark. The historical audit included 50,913 human-written documents, including 40,381 dated from 2000 to 2019. The benchmark contained 300 project-lineage-held-out human sources and 2,398 valid variants. Variants were produced by OpenAI, Gemini, DeepSeek, and Claude under four writing conditions: full generation, AI polishing, deeply mixed human–AI writing, and humanization. Six supervised detectors, one lineage-unresolved deployed system, and a frozen zero-shot baseline were evaluated at locked thresholds. Human false-positive rates and AI recall were estimated with exact binomial intervals. Score uncertainty was estimated with a source-seed clustered bootstrap. XLM-R achieved the highest held-out AUROC at 0.9154. Its human false-positive rate was 0.0400, with AI recall of 0.6952. BERTurk had a human false-positive rate of 0.0267 and AI recall of 0.6197. The deployed system reached 0.8603 recall and falsely flagged 35.0% of human texts. OpenAI outputs were the most difficult provider condition for all project-lineage-held-out supervised detectors. Deeply mixed text produced the lowest recall for the transformer models. AI-polished text was the most difficult condition for the TF-IDF baselines. The zero-shot baseline remained near chance at an AUROC of 0.5069. Its recall was zero at the locked threshold. Historical threshold transfer varied by architecture. The findings support local validation and explicit reporting of human false-positive rates in consequential academic use.

Keywords

References

  1. S. Messick, "Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning," American Psychologist, vol. 50, no. 9, pp. 741–749, 1995, doi: 10.1037/0003-066X.50.9.741.
  2. R. J. Vandenberg and C. E. Lance, "A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research," Organizational Research Methods, vol. 3, no. 1, pp. 4–70, 2000, doi: 10.1177/109442810031002.
  3. S. Gehrmann, H. Strobelt, and A. M. Rush, "GLTR: Statistical detection and visualization of generated text," in Proc. 57th Annu. Meeting Assoc. Comput. Linguist.: Syst. Demonstrations, 2019, pp. 111–116, doi: 10.18653/v1/P19-3019.
  4. E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, "DetectGPT: Zero-shot machine-generated text detection using probability curvature," in Proc. 40th Int. Conf. Mach. Learn., ser. PMLR, vol. 202, 2023, pp. 24950–24962, doi: 10.48550/arXiv.2301.11305.
  5. G. Bao, Y. Zhao, Z. Teng, L. Yang, and Y. Zhang, "Fast-DetectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature," in Proc. Int. Conf. Learn. Representations (ICLR), 2024, doi: 10.48550/arXiv.2310.05130.
  6. A. Hans, A. Schwarzschild, V. Cherepanova, H. Kazemi, A. Saha, M. Goldblum, J. Geiping, and T. Goldstein, "Spotting LLMs with Binoculars: Zero-shot detection of machine-generated text," in Proc. 41st Int. Conf. Mach. Learn., ser. PMLR, vol. 235, 2024, pp. 17519–17537, doi: 10.48550/arXiv.2401.12070.
  7. W. Hao, R. Li, W. Zhao, J. Yang, and C. Mao, "Learning to rewrite: Generalized LLM-generated text detection," in Proc. 63rd Annu. Meeting Assoc. Comput. Linguist. (Volume 1: Long Papers), 2025, pp. 6421–6434, doi: 10.18653/v1/2025.acl-long.322.
  8. X. Yang, W. Cheng, Y. Wu, L. Petzold, W. Y. Wang, and H. Chen, "DNA-GPT: Divergent N-gram analysis for training-free detection of GPT-generated text," in Proc. Int. Conf. Learn. Representations (ICLR), 2024, doi: 10.48550/arXiv.2305.17359.

Details

Primary Language

English

Subjects

Natural Language Processing

Journal Section

Research Article

Publication Date

September 17, 2026

Submission Date

September 4, 2026

Acceptance Date

September 15, 2026

Published in Issue

Year 2026 Volume: 9 Number: 1

APA
Tüker, M., Sinecen, M., & Eser, M. T. (2026). Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts. Scientific Journal of Mehmet Akif Ersoy University, 9(1), 53-67. https://doi.org/10.70030/sjmakeu.2033004
AMA
1.Tüker M, Sinecen M, Eser MT. Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts. Techno-Science. 2026;9(1):53-67. doi:10.70030/sjmakeu.2033004
Chicago
Tüker, Mustafa, Mahmut Sinecen, and Mehmet Taha Eser. 2026. “Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts”. Scientific Journal of Mehmet Akif Ersoy University 9 (1): 53-67. https://doi.org/10.70030/sjmakeu.2033004.
EndNote
Tüker M, Sinecen M, Eser MT (September 1, 2026) Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts. Scientific Journal of Mehmet Akif Ersoy University 9 1 53–67.
IEEE
[1]M. Tüker, M. Sinecen, and M. T. Eser, “Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts”, Techno-Science, vol. 9, no. 1, pp. 53–67, Sept. 2026, doi: 10.70030/sjmakeu.2033004.
ISNAD
Tüker, Mustafa - Sinecen, Mahmut - Eser, Mehmet Taha. “Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts”. Scientific Journal of Mehmet Akif Ersoy University 9/1 (September 1, 2026): 53-67. https://doi.org/10.70030/sjmakeu.2033004.
JAMA
1.Tüker M, Sinecen M, Eser MT. Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts. Techno-Science. 2026;9:53–67.
MLA
Tüker, Mustafa, et al. “Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts”. Scientific Journal of Mehmet Akif Ersoy University, vol. 9, no. 1, Sept. 2026, pp. 53-67, doi:10.70030/sjmakeu.2033004.
Vancouver
1.Mustafa Tüker, Mahmut Sinecen, Mehmet Taha Eser. Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts. Techno-Science. 2026 Sep. 1;9(1):53-67. doi:10.70030/sjmakeu.2033004