Use of AI for Generating and Grading Exam Questions in an Undergraduate English Linguistics Course
DOI:
https://doi.org/10.5294/edu.2026.29.1.6Keywords:
AI-generated exam grading, ChatGPT, DeepSeek, educational assessment, exams, artificial intelligenceAbstract
This study explores the application of artificial intelligence (AI) in the generation and grading of exam questions for undergraduate English linguistics courses. Specifically, it compares the performance of two AI models, ChatGPT and DeepSeek, in creating fair, coherent, and varied questions, as well as their effectiveness in grading subjective responses. Using a mixed-methods approach that combines quantitative analysis of AI-generated content with comparative evaluation against human criteria, the study identifies key advantages in efficiency, objectivity, and adaptability, along with challenges such as limited contextual understanding and difficulties in assessing creative responses. The findings suggest that while AI can significantly support teaching, its successful implementation requires careful calibration, ongoing supervision, and ethical considerations. The study concludes with practical recommendations for educators to optimize the use of AI-based assessment tools in order to maintain academic standards and improve student performance.
Downloads
References
Alers, H., Malinowska, A., Meghoe, G. y Apfel, E. (2024). Using ChatGPT-4 to grade open question exams. En Lecture Notes in Networks and Systems (vol. 919, pp. 1-9). Springer.
Bewersdorff, A., Seßler, K., Baur, A., Kasneci, E. y Nerdel, C. (2023). Assessing student errors in experimentation using artificial intelligence and large language models: A comparative study with human raters. Computers and Education: Artificial Intelligence, 5. https://doi.org/10.1016/j.caeai.2023.100177
Choi, J. H. (2024). How to use large language models for empirical legal research. Journal of Institutional and Theoretical Economics, 180(2), 214-233. https://doi.org/10.1628/jite-2024-0006
Chu, C.-H. y Liu, Y.-L. (2023). Augmented reality user interface design and experimental evaluation for human-robot collaborative assembly. Journal of Manufacturing Systems, 68, 313-324. https://doi.org/10.1016/j.jmsy.2023.04.007
Chu, H. y Liu, Y. (2024). Research on reforming the training model of master’s talents in design in ethnic regions of universities based on artificial intelligence. En 2024 The 9th International Conference on Information and Education Innovations (pp. 57-62). ACM. https://doi.org/10.1145/3664934.3664949
Cooper, G. (2023). Examining Science education in ChatGPT: An exploratory study of generative artificial intelligence. Journal of Science Education and Technology, 32(3), 444-452. https://doi.org/10.1007/s10956-023-10039-y
Daun, M. y Brings, J. (2023). How ChatGPT will change software engineering education. En ITiCSE 2023: Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education (vol. 1, pp. 110-116). ACM. https://doi.org/10.1145/3587102.3588815
Dai, W., Lin, J., Jin, H., Li, T., Tsai, Y.‑S., Gašević, D., y Chen, G. (2023). Can large language models provide feedback to students? A case study on ChatGPT. 2023 IEEE International Conference on Advanced Learning Technologies (ICALT), 323-325. https://doi.org/10.1109/ICALT58122.2023.00100
Deng, R., Jiang, M., Yu, X., Lu, Y. y Liu, S. (2025). Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Computers and Education, 227. https://doi.org/10.1016/j.compedu.2024.105224
Falchikov, N. y Goldfinch, J. (2000). Student peer assessment in higher education: A meta-analysis comparing peer and teacher marks. Review of Educational Research, 70(3), 287-322. https://doi.org/10.3102/00346543070003287
Farazouli, A., Cerratto-Pargman, T., Bolander-Laksov, K. y McGrath, C. (2024). Hello GPT! Goodbye home examination? An exploratory study of AI chatbots impact on university teachers’ assessment practices. Assessment and Evaluation in Higher Education, 49(3), 363-375. https://doi.org/10.1080/02602938.2023.2241676
Gencer, A. y Aydin, S. (2023). Can ChatGPT pass the thoracic surgery exam? American Journal of the Medical Sciences, 366(4), 291-295. https://doi.org/10.1016/j.amjms.2023.08.001
Kayaalp, M. E., Prill, R., Sezgin, E. A., Cong, T., Królikowska, A. y Hirschmann, M. T. (2025). DeepSeek versus ChatGPT: Multimodal artificial intelligence revolutionizing scientific discovery. From language editing to autonomous content generation—Redefining innovation in research and practice. Knee Surgery, Sports Traumatology, Arthroscopy, 33(5), 1553-1556. https://doi.org/10.1002/ksa.12628
Kung, J. E., Marshall, C., Gauthier, C., Gonzalez, T. A. y Jackson, J. B. (2023). Evaluating ChatGPT performance on the orthopaedic in-training examination. The Journal of Bone and Joint Surgery, 8(3), e23.00056. https://doi.org/10.2106/JBJS.23.0005
Leddo, J. (2021). Comparing the effectiveness of AI-powered educational software to human teachers. International Journal of Social Science and Economic Research, 6(3). https://doi.org/10.46609/IJSSER.2021.v06i03.015
Lee, G.-G. y Zhai, X. (2024). Using ChatGPT for science learning: A study on pre-service teachers’ lesson planning. IEEE Transactions on Learning Technologies, 17, 1683-1700. https://doi.org/10.1109/TLT.2024.3401457
Leiker, D., Finnigan, S., Gyllen, A. R. y Cukurova, M. (2023). Prototyping the use of Large Language Models (LLMs) for adult learning content creation at scale. arXiv:2306.01815, 3-7. https://doi.org/10.48550/arXiv.2306.01815
Mabrito, M. (2025). Collaborating with generative AI in the English classroom. International Journal of Technology, Knowledge and Society, 21(2), 1-23. https://doi.org/10.18848/1832-3669/CGP/v21i02/1-23
Markowitz, D. M. (2024). From complexity to clarity: How AI enhances perceptions of scientists and the public’s understanding of science. PNAS Nexus, 3(9). https://doi.org/10.1093/pnasnexus/pgae387
Markowitz, D. M., Hancock, J. T. y Bailenson, J. N. (2024). Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews. Journal of Language and Social Psychology, 43(1), 63-82. https://doi.org/10.1177/0261927X231200201
Mizumoto, A. y Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2(2). https://doi.org/10.1016/j.rmal.2023.100050
Peng, Y., Malin, B. A., Rousseau, J. F., Wang, Y., Xu, Z., Xu, X., Weng, C. y Bian, J. (2025). From GPT to DeepSeek: Significant gaps remain in realizing AI in healthcare. Journal of Biomedical Informatics, 163, 104791. https://doi.org/10.1016/j.jbi.2025.104791
Pinto, G., Cardoso-Pereira, I., Monteiro, D., Lucena, D., Souza, A. y Gama, K. (2023). Large language models for education: Grading open-ended questions using ChatGPT. arXiv:2307.16696, 293-302. https://doi.org/10.1145/3613372.3614197
Rana, S., Sheshadri, T., Malhotra, N. y Mahabub Basha, S. (2024). Creating digital learning environments: Tools and technologies for success. En Transdisciplinary teaching and technological integration for improved learning: Case studies and practical approaches (pp. 1-21). IGI Global.
Sadler, P. M. y Good, E. (2006). The impact of self- and peer-grading on student learning. Educational Assessment, 11(1), 1-31. https://doi.org/10.1207/s15326977ea1101_1
Sun X., y Lin H. Investigación práctica sobre inteligencia artificial en la evaluación de exámenes - tomando como ejemplo las preguntas de geografía del examen de ingreso a la universidad de la provincia de Jiangsu de 2023. Geography Teaching, 2024(5): 21-23, 34.
Viera, A. J. y Garrett, J. M. (2008). Preliminary study of a school-based program to improve hypertension awareness in the community. Family Medicine, 40(4), 264-270.
Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P. y Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(1). https://doi.org/10.1007/s40979-023-00146-z
Xu Wenbo, Zhou Xiaoping. “Exploración de la aplicación de modelos de lenguaje grandes tipo ChatGPT en la evaluación de cursos de enfermería - basado en pruebas con ChatGPT, Wenxin Yiyan y iFlytek Spark”. China Medical Education Technology, 2024, 38(5): 567-571.
Yeadon, W., Agra, E., Inyang, O.-O., Mackay, P. y Mizouri, A. (2024). Evaluating AI and human authorship quality in academic writing through physics essays. European Journal of Physics, 45(5). https://doi.org/10.1088/1361-6404/ad669d
Zhang, Y., Fen, B. W., Zhang, C. y Pi, S. (2024). Transforming music education through artificial intelligence: A Systematic literature review on enhancing music teaching and learning. International Journal of Interactive Mobile Technologies, 18(18), 76-93. https://doi.org/10.3991/ijim.v18i18.50545
Zupanc, K. y Bosnić, Z. (2020). Improvement of automated essay grading by grouping similar graders. Fundamenta Informaticae, 172(3), 239-259. https://doi.org/10.3233/FI-2020-1904
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Tina Xu Geng, Raquel Raya-Porcel

This work is licensed under a Creative Commons Attribution 4.0 International License.
1. Proposed Policy for Journals That Offer Open Access
Authors who publish with this journal agree to the following terms:
-
This journal and its papers are published with the Creative Commons License CC BY 4.0 DEED Atribución 4.0 Internacional. You are free to share copy and redistribute the material in any medium or format if you: give appropriate credit, provide a link to the license, and indicate if changes were made; don’t use our material for commercial purposes; don’t remix, transform, or build upon the material.



