A Corpus-based Model of Methodological Transformation in Arabic Studies within the Digital Humanities Era

Authors

  • Harmanto Raharjo Islamic University of Sunan Gunung Djati Bandung, Jl. A.H. Nasution No. 105, Cipadung, Cibiru, Bandung, West Java, Indonesia
  • Muhammad Arwani Islamic University of Sunan Gunung Djati Bandung, Jl. A.H. Nasution No. 105, Cipadung, Cibiru, Bandung, West Java, Indonesia
  • Ahmad Syahid Islamic University of Sunan Gunung Djati Bandung, Jl. A.H. Nasution No. 105, Cipadung, Cibiru, Bandung, West Java, Indonesia
  • Mahrus As’ad Islamic University of Sunan Gunung Djati Bandung, Jl. A.H. Nasution No. 105, Cipadung, Cibiru, Bandung, West Java, Indonesia
  • Ateng Ruhendi Islamic University of Sunan Gunung Djati Bandung, Jl. A.H. Nasution No. 105, Cipadung, Cibiru, Bandung, West Java, Indonesia
  • Aulia Rahmi Islamic University of Sunan Gunung Djati Bandung, Jl. A.H. Nasution No. 105, Cipadung, Cibiru, Bandung, West Java, Indonesia
  • Henny Septia Utami Queen's University Belfast, United Kingdom

DOI:

https://doi.org/10.24252/jad.v26i2a3

Keywords:

Arabic digital corpus, Corpus linguistics, Digital Humanities, Methodological transformation model, Arabic studies

Abstract

The rise of the Digital Humanities has reshaped the methodological landscape of the humanities, including Arabic studies, through the use of digital corpora as a basis for linguistic analysis. Yet studies that specifically link corpus linguistics to the methodological transformation of Arabic studies and that offer an explicit conceptual framework remain limited. This research analyzes the position of corpus linguistics within the Digital Humanities, models the methodological transformation of Arabic studies that it engenders, and maps the challenges and prospects of its development. Adopting a qualitative, library-research design of a meta-methodological character, it combines conceptual and analytical approaches and examines reputable literature published between 2019 and 2026 through content analysis. As its principal contribution, the research formulates a Three-Layer Model of Corpus-Based Methodological Transformation in Arabic Studies linking a data-infrastructure layer, an analytical-procedure layer, and an epistemological layer together with a typology of corpus applications across four domains of Arabic studies (the Qur’an, turāth, modern media, and pedagogy). The model is illustrated concretely through frequency, collocation, and concordance analysis of the root ع-ل-م in the Quranic Arabic Corpus. The findings show that corpus linguistics drives a shift from text-centered toward data-driven studies that are more empirical and reproducible, although it still faces the limited availability of standardized corpora, morphological complexity and diglossia, academic digital literacy, and technological infrastructure. The research recommends developing representative, open-access Arabic digital corpora, strengthening scholars’ digital literacy, and pursuing a balanced integration of classical philological methods and data-based approaches.

ملخص

أحدث التحوّل نحو الإنسانيّات الرقميّة (Digital Humanities) نقلةً في المشهد المنهجيّ للعلوم الإنسانيّة، ومنها الدراسات العربيّة، عبر توظيف المدوّنات الرقميّة أساسًا لتحليل اللغة. غير أنّ الدراسات التي تربط لسانيّات المدوّنة (corpus linguistics) تحديدًا بالتحوّل المنهجيّ في الدراسات العربيّة، وتقدّم إطارًا مفاهيميًّا صريحًا، لا تزال محدودة. يهدف هذا البحث إلى تحليل موقع لسانيّات المدوّنة ضمن الإنسانيّات الرقميّة، ونمذجة التحوّل المنهجيّ الذي تُحدثه في الدراسات العربيّة، ورسم تحدّياته وآفاقه. اعتمد البحث المنهج الكيفيّ عبر الدراسة المكتبيّة (library research) ذات الطابع الميتا-منهجيّ، جامعًا بين المقاربتين المفاهيميّة والتحليليّة، ومحلِّلًا أدبيّات محكَّمة صادرة بين عامَي 2019 و2026 عبر تحليل المضمون (content analysis). وكإسهام رئيس، صاغ البحث «نموذج الطبقات الثلاث للتحوّل المنهجيّ المستند إلى المدوّنة في الدراسات العربيّة»، الذي يربط بين طبقة البنية التحتيّة للبيانات، وطبقة الإجراء التحليليّ، والطبقة المعرفيّة (الإبستمولوجيّة)، إلى جانب تصنيف (typology) لتطبيقات المدوّنة في أربعة مجالات للدراسات العربيّة: القرآن الكريم، والتراث، والإعلام الحديث، والتعليم. وقد جرى توضيح النموذج تطبيقيًّا من خلال تحليل التواتر، والمصاحبة اللفظيّة، والسياقات الكشفيّة (concordance) لجذر (ع-ل-م) في المدوّنة القرآنيّة العربيّة (Quranic Arabic Corpus). وأظهرت النتائج أنّ لسانيّات المدوّنة تدفع نحو الانتقال من الدراسات المتمركزة حول النصّ (text-centered) إلى الدراسات المقودة بالبيانات (data-driven) الأكثر تجريبيّةً وقابليّةً للتكرار، على الرغم مما تواجهه من محدوديّة المدوّنات العربيّة المعياريّة، وتعقيد البنية الصرفيّة والازدواجيّة اللغويّة (diglossia)، وضعف المعرفة الرقميّة الأكاديميّة، وقصور البنية التحتيّة التقنيّة. ويوصي البحث بتطوير مدوّنات عربيّة رقميّة تمثيليّة ومفتوحة الوصول (open access)، وتعزيز المعرفة الرقميّة لدى الباحثين، وتحقيق تكامل متوازن بين المناهج الفيلولوجيّة الكلاسيكيّة والمقاربات المستندة إلى البيانات.

 Abstrak

Transformasi Digital Humanities telah mengubah lanskap metodologis ilmu humaniora, termasuk studi Arab, melalui pemanfaatan korpus digital sebagai basis analisis kebahasaan. Namun, kajian yang secara khusus menghubungkan linguistik korpus dengan transformasi metodologi studi Arab dan menawarkan kerangka konseptual yang eksplisit masih terbatas. Penelitian ini bertujuan menganalisis posisi linguistik korpus dalam Digital Humanities, memodelkan transformasi metodologi studi Arab yang ditimbulkannya, serta memetakan tantangan dan prospek pengembangannya. Penelitian menggunakan pendekatan kualitatif dengan metode kepustakaan (library research) yang bersifat meta-metodologis, memadukan pendekatan konseptual dan analitis, serta menganalisis literatur bereputasi terbitan 2019–2026 melalui analisis isi (content analysis). Sebagai kontribusi utama, penelitian merumuskan Model Tiga Lapis Transformasi Metodologi Studi Arab Berbasis Korpus yang menautkan lapis infrastruktur data, lapis prosedur analitis, dan lapis epistemologis serta tipologi penerapan korpus pada empat ranah studi Arab (Al-Qur’an, turats, media modern, dan pedagogi). Model tersebut diilustrasikan secara konkret melalui analisis frekuensi, kolokasi, dan konkordansi atas akar ع-ل-م dalam Quranic Arabic Corpus. Hasil menunjukkan bahwa linguistik korpus mendorong pergeseran dari pendekatan text-centered menuju data-driven studies yang lebih empiris dan reproduktif, meski masih menghadapi keterbatasan korpus terstandardisasi, kompleksitas morfologi dan diglosia, literasi digital akademik, serta infrastruktur teknologi. Penelitian merekomendasikan pengembangan Arabic digital corpus yang representatif dan open-access, peningkatan literasi digital akademisi, serta integrasi seimbang antara metode filologis klasik dan pendekatan berbasis data.

Downloads

Download data is not yet available.

References

Abazoglu, Muhammet, and Mohammad Issa Alhourani. “The Use of Language Corpora in Teaching Arabic to Turkish Speakers Within the Framework of Computational Linguistics.” Social Sciences & Humanities Open 12 (2025): 101947. doi:10.1016/j.ssaho.2025.101947.

Alayba, Abdulaziz M. “Arabic Natural Language Processing (NLP): A Comprehensive Review of Challenges, Techniques, and Emerging Trends.” Computers 14, no. 11 (November 15, 2025): 497. doi:10.3390/computers14110497.

Alkaabi, Hussein Ala’a, Ali Kadhim Jasim, and Ali Darroudi. “Arabic NLP: A Survey of Pre-Processing and Representation Techniques.” JCoSITTE: Journal of Computer Science, Information Technology and Telecommunication Engineering 6, no. 2 (2025): 876–90.

Aziz, Abd, and Saihu Saihu. “Interpretasi Humanistik Kebahasaan: Upaya Kontekstualisasi Kaidah Bahasa Arab.” Arabiyatuna : Jurnal Bahasa Arab 3, no. 2 (November 13, 2019): 299. doi:10.29240/jba.v3i2.1000.

Bashir, Muhammad Huzaifa, Aqil M. Azmi, Haq Nawaz, Wajdi Zaghouani, Mona Diab, Ala Al-Fuqaha, and Junaid Qadir. “Arabic Natural Language Processing for Qur’anic Research: A Systematic Review.” Artificial Intelligence Review 56, no. 7 (July 2023): 6801–54. doi:10.1007/s10462-022-10313-2.

Belinkov, Yonatan, Alexander Magidow, Alberto Barrón-Cedeño, Avi Shmidman, and Maxim Romanov. “Studying the History of the Arabic Language: Language Technology and a Large-Scale Historical Corpus.” Language Resources and Evaluation 53, no. 4 (December 2019): 771–805. doi:10.1007/s10579-019-09460-w.

Besdouri, Fatma Zahra, Inès Zribi, and Lamia Hadrich Belguith. “Arabic Automatic Speech Recognition: Challenges and Progress.” Speech Communication 163 (September 2024): 103110. doi:10.1016/j.specom.2024.103110.

Bonde Thylstrup, Nanna, Mikkel Flyverbom, and Rasmus Helles. “Datafied Knowledge Production: Introduction to the Special Theme.” Big Data & Society 6, no. 2 (July 2019): 205395171987598. doi:10.1177/2053951719875985.

Brookes, Gavin, and Tony McEnery. “Correlation, Collocation and Cohesion: A Corpus-Based Critical Analysis of Violent Jihadist Discourse.” Discourse & Society 31, no. 4 (July 2020): 351–73. doi:10.1177/0957926520903528.

Charfi, Anis, Wajdi Zaghouani, Syed Hassan Mehdi, and Esraa Mohamed. “A Fine-Grained Annotated Multi-Dialectal Arabic Corpus.” In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019), 198–204. Varna: INCOMA Ltd., 2019. doi:10.26615/978-954-452-056-4_023.

Dukes, Kais, and Nizar Habash. “Morphological Annotation of Quranic Arabic.” In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC 2010), 2530–36. Valletta: European Language Resources Association (ELRA), 2010.

Egbert, Jesse, Douglas Biber, and Bethany Gray. Designing and Evaluating Language Corpora: A Practical Framework for Corpus Representativeness. 1st ed. Cambridge: Cambridge University Press, 2022. doi:10.1017/9781316584880.

Elewa, Abdelhamid. “Not Ready for Democracy: A Corpus-Assisted Discourse Analysis of the Representation of Majorities in Arabic Media Narratives.” Discourse & Society 36, no. 6 (November 2025): 863–82. doi:10.1177/09579265251323906.

El-Farahaty, Hanem, Nouran Khallaf, and Amani Alonayzan. “Building the Leeds Monolingual and Parallel Legal Corpora of Arabic and English Countries’ Constitutions: Methods, Challenges and Solutions.” Corpus Pragmatics 7, no. 2 (June 2023): 103–19. doi:10.1007/s41701-023-00138-x.

Fauziyah, Nur Nabilah, and Ni Gusti Ayu Roselani. “A Corpus-Based Critical Discourse Analysis of English Discourse in The Jakarta Post.” K@ta 27, no. 1 (June 30, 2025): 80–94. doi:10.9744/kata.27.1.80-94.

Flanagan, Joseph. “Reproducibility, Replicability, Robustness, and Generalizability in Corpus Linguistics.” International Journal of Corpus Linguistics 30, no. 2 (October 2, 2025): 130–49. doi:10.1075/ijcl.24113.fla.

Gamal, Donia, Marco Alfonse, El-Sayed M. El-Horbaty, and Abdel-Badeeh M. Salem. “Implementation of Machine Learning Algorithms in Arabic Sentiment Analysis Using N-Gram Features.” Procedia Computer Science 154 (2019): 332–40. doi:10.1016/j.procs.2019.06.048.

Gries, Stefan Th. “A New Approach to (Key) Keywords Analysis: Using Frequency, and Now Also Dispersion.” Research in Corpus Linguistics 9, no. 2 (2021): 1–33. doi:10.32714/ricl.09.02.02.

Hallberg, Andreas. “An 81-Million-Word Multi-Genre Corpus of Arabic Books.” Data in Brief 60 (June 2025): 111456. doi:10.1016/j.dib.2025.111456.

Hamamah, Hamamah. “To Data Driven Research and Beyond: An Overview of Corpus Linguistics Research in Indonesia.” In Proceedings of the 2nd International Conference on Language, Literature, Education, and Culture, ICOLLEC 2022, 11–12 November 2022, Malang, Indonesia. Malang, Indonesia: EAI, 2023. doi:10.4108/eai.11-11-2022.2329401.

Hashimoto, Brett, and Kyra Nelson. “Recent Trends in Corpus Design and Reporting: A Methodological Synthesis.” Research in Corpus Linguistics 12, no. 1 (2024): 59–88. doi:10.32714/ricl.12.01.03.

Hizbullah, Nur, Zakiyah Arifa, Yoke Suryadarma, Ferry Hidayat, Luthfi Muhyiddin, and Eka Kurnia Firmansyah. “Source-Based Arabic Language Learning: A Corpus Linguistic Approach.” Humanities & Social Sciences Reviews 8, no. 3 (June 17, 2020): 940–54. doi:10.18510/hssr.2020.8398.

Hizbullah, Nur, Fazlur Rachman, and Fuzi Fauziah. “Penyusunan Model Korpus Al-Qur’an Digital.” Jurnal Al-Azhar Indonesia Seri Humaniora 3, no. 3 (December 20, 2017): 215. doi:10.36722/sh.v3i3.209.

Hizbullah, Nur, Iin Suryaningsih, and Zaqiatul Mardiah. “Manuskrip Arab Di Nusantara Dalam Tinjauan Linguistik Korpus.” Arabi : Journal of Arabic Studies 4, no. 1 (July 1, 2019): 65. doi:10.24865/ajas.v4i1.145.

Joyeux-Prunel, Béatrice. “Digital Humanities in the Era of Digital Reproducibility: Towards a Fairest and Post-Computational Framework,.” International Journal of Digital Humanities 6, no. 1 (January 3, 2024): 23–43. doi:10.1007/s42803-023-00079-6.

Mak, Matthew H.C. “Corpus Linguistics Will Benefit from Greater Adoption of Pre-Registration: A Novice-Friendly Split-Corpus Approach to Pre-Registration.” Applied Corpus Linguistics 4, no. 3 (December 2024): 100111. doi:10.1016/j.acorp.2024.100111.

Mohamed, Eid. “The Potential and Limits of Arabic Digital Humanities.” Journal of Cultural Analytics 9, no. 3 (June 17, 2024): 889. doi:10.22148/001c.116818.

Mohd Yousof, Noor Mohamed, Ali Selamat, Zatul Alwani Shaffiei, Siti Nur Khadijah Aishah Ibrahim, Liyana ‘Adilla Burhanuddin, and Hamido Fujita. “An Intelligent NLP-Based Framework for Digital Scripts Concordance and Semantic Exploration.” In Frontiers in Artificial Intelligence and Applications, edited by Hamido Fujita, Andres Hernandez-Matamoros, and Yutaka Watanobe. Amsterdam: IOS Press, 2025. doi:10.3233/FAIA250545.

Pérez-Paredes, Pascual, and Niall Curry. “Epistemologies of Corpus Linguistics across Disciplines.” Research Methods in Applied Linguistics 3, no. 3 (December 2024): 100141. doi:10.1016/j.rmal.2024.100141.

Priem, Karin, and Lynn Fendler. “Shifting Epistemologies for Discipline and Rigor in Educational Research: Challenges and Opportunities from Digital Humanities.” European Educational Research Journal 18, no. 5 (September 2019): 610–21. doi:10.1177/1474904118820433.

Ramadhan, Adrianto, Weningtyas Parama Iswari, and Aridah. “Corpus-Based Study On The Use Of Reporting Verbs In Applied Linguistics Journal Articles Published From 2020-2024.” E3L: Journal of English Teaching, Linguistic, and Literature 8, no. 1 (June 25, 2025): 9–26. doi:10.30872/e3l.v8i1.5242.

Sawalha, Majdi, Faisal Al-Shargi, Sane Yagi, Abdallah T. AlShdaifat, Bassam Hammo, Mariam Belajeed, and Lubna R. Al-Ogaili. “Morphologically-Analyzed and Syntactically-Annotated Quran Dataset.” Data in Brief 58 (February 2025): 111211. doi:10.1016/j.dib.2024.111211.

Shormani, Mohammed Q., and Mohammad A. N. Alenezi. “Arabic Nominalization as Form of Pragmatic Actions in Digital Discourse: A Corpus-Based Study.” Corpus Pragmatics 10, no. 1 (December 2026): 35. doi:10.1007/s41701-026-00237-5.

Smadi, Alaa’ Mohammad. “A Corpus-Based Analysis of the Lexical and Grammatical Properties of Arabic Abstracts in Social Sciences Articles.” International Journal of Linguistics 12, no. 6 (November 24, 2020). doi:10.5296/ijl.v12i6.17849.

Syahrullah, Syahrullah, Wildan Imaduddin Muhammad, Eva Nugraha, and Aulia Raudhatul Jannah. “Corpus Coranicum and Digital Philology: A Methodological Model for Advancing Qur’anic Manuscript Studies in Indonesia.” Khazanah Theologia 6, no. 3 (December 29, 2024): 201–22. doi:10.15575/kt.v6i3.45610.

Published

2026-09-26

How to Cite

Raharjo, H., Arwani, M., Syahid, A., As’ad, M., Ruhendi, A., Rahmi, A., & Utami, H. S. (2026). A Corpus-based Model of Methodological Transformation in Arabic Studies within the Digital Humanities Era. Jurnal Adabiyah, 26(2). https://doi.org/10.24252/jad.v26i2a3