Optimalisasi Teknologi OCR dalam Digitalisasi Manuskrip Hadis: Studi Akurasi dan Tantangan Linguistik Arab Klasik
DOI:
https://doi.org/10.64691/6vn20070Keywords:
OCR; Hadith Digitization; Arabic Linguistics; Text Accuracy; Hadith ManuscriptsAbstract
The digitization of hadith manuscripts is a strategic step in preserving and disseminating Islamic intellectual treasures. Still, this process faces significant challenges related to the accuracy of classical Arabic text recognition using Optical Character Recognition (OCR) technology. The complexity of Arabic orthography, the diversity of calligraphic styles, and the presence of diacritics and ligatures typical of classical manuscripts often lead to errors in character recognition, which impact the reliability of the digitization results. This study aims to analyze the accuracy of various OCR systems in reading hadith manuscripts, identify the main linguistic constraints that cause distortions in letter and word recognition, and develop an optimization model based on linguistic and computational integration to improve system performance. Using a mixed methods approach, this study combines quantitative analysis of text processing results from several popular OCR platforms with qualitative analysis of linguistic error patterns that appear in classical Arabic manuscripts. The results show that the accuracy of OCR for hadith manuscripts varies widely, ranging from 68% to 91%, depending on image quality, calligraphy type, and the system’s ability to recognize morphological forms and punctuation typical of classical Arabic. The most common errors occur in letters with graphemic similarities, such as “س” and “ش,” as well as in words with full vowels or complex Iʻrāb markings. An optimization model developed through a combination of machine learning and Arabic morphological analysis has been shown to improve accuracy by up to 94%, while accelerating the post-correction process for digital texts. In conclusion, the digitization of hadith manuscripts requires an integrative approach between OCR technology and an understanding of classical Arabic linguistics to ensure the validity, efficiency, and sustainability of digital religious manuscript preservation.
References
Ahmed, R., Gogate, M., Tahir, A., Dashtipour, K., Al-Tamimi, B., Hawalah, A., … Hussain, A. (2021). Deep neural network-based contextual recognition of arabic handwritten scripts. Entropy, 23(3), 4–6. https://doi.org/10.3390/e23030340
Alghamdi, M., & Teahan, W. (2017). Experimental evaluation of Arabic OCR systems. PSU Research Review, 1(3), 229–241. https://doi.org/10.1108/PRR-05-2017-0026
Alghyaline, S. (2023). Arabic Optical Character Recognition: A Review. CMES - Computer Modeling in Engineering and Sciences, 135(3), 1825–1861. https://doi.org/10.32604/cmes.2022.024555
Ayadi, M., Masmoudi, N., Almuqren, L., Alshahrani, H. S., & Aljohani, R. O. (2025). Designing a Novel CNN–LSTM-based Model for Arabic Handwritten Character Recognition for the Visually Impaired Person. Journal of Disability Research, 4(1), 1–13. https://doi.org/10.57197/JDR-2024-0080
Balaha, H. M., Ali, H. A., Youssef, E. K., Elsayed, A. E., Samak, R. A., Abdelhaleem, M. S., … Mohammed, M. M. (2021). Recognizing arabic handwritten characters using deep learning and genetic algorithms. Multimedia Tools and Applications, 80(21), 32473–32509. https://doi.org/10.1007/s11042-021-11185-4
Chaudhuri, A., Mandaviya, K., Badelia, P., & Ghosh, S. K. (2017). Optical Character Recognition Systems. In A. Chaudhuri, K. Mandaviya, P. Badelia, & S. K Ghosh (Eds.), Optical Character Recognition Systems for Different Languages with Soft Computing (pp. 9–41). Cham: Springer International Publishing. https://doi.org/10.1007/978-3-319-50252-6_2
Daffa, M. (2022). Analysis Of Hadith Understanding Of Social Media Phenomena As A Communication Tool In The Digital Era. Riwayah : Jurnal Studi Hadis, 8(1), 69–86. https://doi.org/10.21043/riwayah.v8i1.11209
Daud, Z., & Junus, R. A. (2023). The Use of Maktabah Syamilah Software in Increasing Students’ Interest of Learning Hadith: A Survey. Journal Of Hadith Studies, 8(2), 94–104. https://doi.org/10.33102/johs.v8i2.249
Hládek, D., Staš, J., Ondáš, S., Juhár, J., & Kovács, L. (2017). Learning string distance with smoothing for OCR spelling correction. Multimedia Tools and Applications, 76(22), 24549–24567. https://doi.org/10.1007/s11042-016-4185-5
Mosbah, L., Moalla, I., Hamdani, T. M., Neji, B., Beyrouthy, T., & Alimi, A. M. (2024). ADOCRNet: A Deep Learning OCR for Arabic Documents Recognition. IEEE Access, 12, 55620–55631. https://doi.org/10.1109/ACCESS.2024.3379530
Park, J., Lee, E., Kim, Y., Kang, I., Koo, H. I., & Cho, N. I. (2020). Multi-Lingual Optical Character Recognition System Using the Reinforcement Learning of Character Segmenter. IEEE Access, 8, 174437–174448. https://doi.org/10.1109/ACCESS.2020.3025769
Siregar, U. N. (2025). Penggunaan AI dan Big Data dalam Pembelajaran Nilai-Nilai Al-Qur’an dan Hadis di Madrasah. Arba: Jurnal Studi Keislaman, 1(3), 160–175. https://doi.org/10.64691/arba.v1i3.13
Yeh, C., Chen, Y., Wu, A., Chen, C., Viégas, F., & Wattenberg, M. (2024). AttentionViz: A Global View of Transformer Attention. IEEE Transactions on Visualization and Computer Graphics, 30(1), 262–272. https://doi.org/10.1109/TVCG.2023.3327163
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Zeindri Abu Umar & Mahmud Yunus

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.













