Text classification by machine learning algorithms using a new text feature extraction method based on image processing
Abstract
Accurate text and character identification on documents using smart technologies is a very important method of obtaining data. The complex and irregular text and characters on the images, as well as the use of different writing styles, affect the text recognition success of both Artificial Intelligence (AI) and Machine Learning (ML) technologies. Manually transferring texts and characters from paper format documents to digital media creates a great waste of time and labor. In addition, when documents containing direct text are scanned and transferred in a computer environment, the texts cannot be edited. OCR (Optical Character Recognition) methods, which are proposed as a solution to this situation, are one of the Natural Language Processing (NLP) tasks. In particular, it has been observed that even in current artificial intelligence-based OCR software, the characters 0 and O are confused with each other. In this study, it is suggested that image pre-processing should be done on images containing characters in order to increase the success of character recognition. In the study, a new model was designed to increase the success of correctly recognizing 0 and O characters that are very similar to each other. In the study, image pre-processing was applied to the images of 408 characters. Classification successes were measured by using kNN, SVM and Logistic Regression algorithms on the data set. Additionally, the classification performance of 0 and O characters was measured on the artificial intelligence-based Google Documents tool. According to the results obtained, the success of recognizing 0 and O characters with the LR machine learning algorithm was realized at the rate of 1.00 according to the performance metrics.
Keywords
References
- Manwatkar, P.M. & Singh, K.R. (2015). A technical review on text recognition from images. Proceedings of the IEEE 9th International Conference on Intelligent Systems and Control (ISCO), Coimbatore, India, 1-5. https://doi.org/10.1109/ISCO.2015.7282362.
- Prabu & Sundar, K.J.A. (2023) Enhanced attention-based encoder-decoder framework for text recognition. Intelligent Automation & Soft Computing, 35(2), 2071-2086.
- Guan, T., Shen, W., Yang, X., Feng, Q., Jiang, Z., & Yang, X. (2023). Self-supervised character-to-character distillation for text recognition. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 19473-19484. https://doi.org/10.48550/arXiv.2211.00288.
- Zhou, L., Wang, L., Ge X., Shi, Q. (2010). A clustering-Based KNN improved algorithm CLKNN for text classification. Proceedings of the 2nd International Asia Conference on Informatics in Control, Automation and Robotics (CAR 2010), Wuhan, China, 212-215. https://doi.org/10.1109/CAR.2010.5456668.
- Alzoubi, Y., Topcu, A., & Erkaya, A. (2023). Machine learning-based text classification comparison: Turkish language context. Applied Sciences, 13(16), 9428. https://doi.org/10.3390/app13169428.
- Jing, Y., Gou, H., & Zhu, Y. (2013). An improved density-based method for reducing training data in KNN. Proceedings of the International Conference on Computational and Information Sciences, Shiyang, China, 972-975. https://doi.org/10.1109/ICCIS.2013.261.
- Gowda, D.K., & Kanchana, V. (2022). Kannada handwritten character recognition and classification through OCR using hybrid machine learning techniques. Proceedings of the IEEE International Conference on Data Science and Information System (ICDSIS), Hassan, India, 1-6. https://doi.org/10.1109/ICDSIS55133.2022.9915906.
- Nahar, K.M.O., Alsmadi, I., Al Mamlook, R.E., Nasayreh, A., Gharaibeh, H., Almuflih A.S., & Alasim, F. (2023). Recognition of Arabic air-written letters: machine learning, Convolutional Neural Networks, and Optical Character Recognition (OCR) Techniques, Sensors, 23(23), 9475. https://doi.org/10.3390/s23239475.
Details
Primary Language
English
Subjects
Information Systems (Other)
Journal Section
Research Article
Publication Date
October 8, 2025
Submission Date
June 12, 2025
Acceptance Date
September 17, 2025
Published in Issue
Year 2025 Volume: 9 Number: 4
APA
Çelik, A., & Kaptan, D. (2025). Text classification by machine learning algorithms using a new text feature extraction method based on image processing. Turkish Journal of Engineering, 9(4), 712-724. https://doi.org/10.31127/tuje.1718023
AMA
1.Çelik A, Kaptan D. Text classification by machine learning algorithms using a new text feature extraction method based on image processing. TUJE. 2025;9(4):712-724. doi:10.31127/tuje.1718023
Chicago
Çelik, Ahmet, and Deniz Kaptan. 2025. “Text Classification by Machine Learning Algorithms Using a New Text Feature Extraction Method Based on Image Processing”. Turkish Journal of Engineering 9 (4): 712-24. https://doi.org/10.31127/tuje.1718023.
EndNote
Çelik A, Kaptan D (October 1, 2025) Text classification by machine learning algorithms using a new text feature extraction method based on image processing. Turkish Journal of Engineering 9 4 712–724.
IEEE
[1]A. Çelik and D. Kaptan, “Text classification by machine learning algorithms using a new text feature extraction method based on image processing”, TUJE, vol. 9, no. 4, pp. 712–724, Oct. 2025, doi: 10.31127/tuje.1718023.
ISNAD
Çelik, Ahmet - Kaptan, Deniz. “Text Classification by Machine Learning Algorithms Using a New Text Feature Extraction Method Based on Image Processing”. Turkish Journal of Engineering 9/4 (October 1, 2025): 712-724. https://doi.org/10.31127/tuje.1718023.
JAMA
1.Çelik A, Kaptan D. Text classification by machine learning algorithms using a new text feature extraction method based on image processing. TUJE. 2025;9:712–724.
MLA
Çelik, Ahmet, and Deniz Kaptan. “Text Classification by Machine Learning Algorithms Using a New Text Feature Extraction Method Based on Image Processing”. Turkish Journal of Engineering, vol. 9, no. 4, Oct. 2025, pp. 712-24, doi:10.31127/tuje.1718023.
Vancouver
1.Ahmet Çelik, Deniz Kaptan. Text classification by machine learning algorithms using a new text feature extraction method based on image processing. TUJE. 2025 Oct. 1;9(4):712-24. doi:10.31127/tuje.1718023
Cited By
AI-Based Assistant for Generating and Analyzing Legal Documents
Turkish Journal of Engineering
https://doi.org/10.31127/tuje.1843894