Human–AI Collaboration in Academic Writing Instruction: Calibrated Feedback Model for EAP

Authors

  • Reza Farzi UNIVERSITY OF OTTAWA

DOI:

https://doi.org/10.18192/olbij.v15i1.7773

Keywords:

Generative AI, academic writing, feedback, ChatGPT, second language writing, EAP, Ai in education

Abstract

This study adopts a mixed-methods design to compare the feedback generated by a human instructor and ChatGPT-3.5 on student essays. It analyzes scoring consistency, students’ perceptions of feedback usefulness, and the extent to which each form of feedback aligns with instructional objectives. The dataset comprises 30 first-year students enrolled in a Canadian English for Academic Purposes (EAP) course. Each student submitted two argumentative essays, which were evaluated independently by both the instructor and ChatGPT-3.5 across three calibration stages. The findings suggest that while ChatGPT-3.5 consistently identified surface-level issues and demonstrated close alignment with human scoring, it was less effective in addressing rhetorical structure and contextual nuance. Student responses indicated a clear preference for a hybrid feedback model, which combines the immediacy of AI-generated comments with the depth and pedagogical value of instructor feedback. Through calibrated prompts and pedagogical (re-)framing, GenAI can supplement human feedback effectively but will not replace it. This study puts forward a hybrid model of feedback provision in which language instructors can leverage GenAI to support formative assessment while maintaining pedagogical oversight. The findings offer guidance for integrating GenAI into academic writing instruction in ways that preserve instructional quality and student agency.

References

Barkaoui, K. (2010). Variability in ESL essay rating processes: The role of the rating scale and rater experience. Language Assessment Quarterly, 7(1), 54–74. https://doi.org/10.1080/15434300903464418

Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. https://doi.org/10.18653/v1/2020.acl-main.463

Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa

Cotos, E. (2011). Potential of automated writing evaluation feedback. CALICO Journal, 28(2), 420–459. https://doi.org/10.11139/cj.28.2.420-459 Creswell, J. W., & Plano Clark, V. L. (2017). Designing and conducting mixed methods research (3rd ed.). SAGE.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv: 1810.04805. arXiv. https://arxiv.org/abs/1810.04805

Dikli, S. (2006). Automated essay scoring. Turkish Online Journal of Distance Education, 7(1), 49–62. https://dergipark.org.tr/en/pub/tojde/article/176620

Escalante, J., Pack, A., & Barrett, A. (2023). AI-generated feedback on writing: Insights into effcacy and ENL student preference. International Journal of Educational Technology in Higher Education, 20, Article 57. https://doi.org/10.1186/s41239-023-00425-2

Evmenova, A. S., Regan, K., Mergen, R., & Hrisseh, R. (2024). Improving writing feedback for struggling writers: Generative AI to the rescue? TechTrends, 68(4), 790–802. https://doi.org/10.1007/s11528-024-00965-y

Fan, Y., Tan, S., & Lim, G. Y. W. (2024). EAP teacher feedback in the age of AI: Supporting first-year students in EFL disciplinary writing. Australian Journal of Applied Linguistics, 7(3), 250–267. https://doi.org/10.29140/ajal.v7n3.1943

Farzi, R. (2024). Calibrating generative AI for second language writing assessment: Combining statistical validation with prompt design. Assessment and Practice in Educational Sciences, 2(4), 1–12. https://journalapes.com/index.php/apes/article/view/91

Ferris, D. R. (2014). Responding to student writing: Teachers’ philosophies and practices. Assessing Writing, 19, 6–23. https://doi.org/10.1016/j.asw.2013.09.004

Fleckenstein, J., Liebenow, L. W., & Meyer, J. (2023). Automated feedback and writing: A multi-level meta-analysis of effects on students’ performance. Frontiers in Artificial Intelligence, 6, Article 1162454. https://doi.org/10.3389/frai.2023.1162454

García-López, I. M., & Trujillo-Liñán, L. (2025). Ethical and regulatory challenges of generative AI in education: A systematic review. Frontiers in Education, 10, Article 1565938. https://doi.org/10.3389/feduc.2025.1565938

Gayed, J. M., Carlon, M. K. J., Oriola, A. M., & Cross, J. S. (2022). Exploring an AI-based writing assistant’s impact on english language learners. Computers and Education: Artificial Intelligence, 3, Article 100055. https://doi.org/10.1016/j.caeai.2022.100055

Harunasari, S. Y. (2023). Examining the effectiveness of AI-integrated approach in EFL writing: A case of ChatGPT. International Journal of Progressive Sciences and Technologies, 39(2), 357–368. https://doi.org/10.52155/ijpsat.v39.2.5516

Hyland, K. (2016). Teaching and researching writing (3rd ed.). Routledge.

Hyland, K., & Hyland, F. (Eds.). (2019). Feedback in second language writing: Contexts and issues (2nd ed.). Cambridge University Press. https://doi.org/10.1017/9781108635547

Kaliisa, R., Misiejuk, K., López-Pernas, S., & Saqr, M. (2026). How does artificial intelligence compare to human feedback? a meta-analysis of performance, feedback perception, and learning dispositions. Educational Psychology, 46(1), 80–111. https://doi.org/10.1080/01443410.2025.2553639

Karagoz, I. (2025). AI-generated feedback in english writing instruction for language learners: A systematic review. The Reading Matrix: An International Online Journal, 25(1), 51–67. https://www.readingmatrix.com/files/34-9gp90gjo.pdf

Landauer, T. K. (2003). Automatic essay assessment. Assessment in Education: Principles, Policy & Practice, 10(3), 295–308.

https://doi.org/10.1080/0969594032000148154

Lee, S. S., & Moore, R. L. (2024). Harnessing generative AI (GenAI) for automated feedback in higher education: A systematic review. Online Learning, 28(3), 82–104. https://doi.org/10.24059/olj.v28i3.4593

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2021). Retrieval-augmented generation for knowledge-intensive NLP tasks. arXiv. https://doi.org/10.48550/arXiv.2005.11401

Li, S. (2025). Generative AI and second language writing. Digital Studies in Language and Literature, 2(1), 122–152. https://doi.org/10.1515/dsll-2025-0007

Liang, W., Zhang, Y., Codreanu, M., Wang, J., Cao, H., & Zou, J. (2025). The widespread adoption of large language model-assisted writing across society. arXiv: 2502.09747. arXiv. https://arxiv.org/abs/2502.09747

Lund, B. D., Lee, T. H., Mannuru, N. R., & Arutla, N. (2025). AI and academic integrity: Exploring student perceptions and implications for higher education. Journal of Academic Ethics, 23(3), 1545–1565. https://doi.org/10.1007/s10805-025-09613-3

Luo, M., Hu, X., & Zhong, C. (2025). The collaboration of AI and teacher in feedback provision and its impact on EFL learners’ argumentative writing. Education and Information Technologies, 30, 17695–17715. https://doi.org/10.1007/s10639-025-13488-7

McNamara, T. (2000). Language testing. Oxford University Press.

Nelson, A. S., Santamaría, P. V., Javens, J. S., & Ricaurte, M. (2025). Students’ perceptions of generative artificial intelligence (GenAI) use in academic writing in english as a foreign language. Education Sciences, 15(5), Article 611. https://doi.org/10.3390/educsci15050611

Nguyen, A., Hong, Y., Dang, B., & Huang, X. (2024). Human-AI collaboration patterns in AI-assisted academic writing. Studies in Higher Education, 49(5), 847–864. https://doi.org/10.1080/03075079.2024.2323593

Nguyen, K. V. (2025). The use of generative AI tools in higher education: Ethical and pedagogical principles. Journal of Academic Ethics, 23, 1435–1455. https://doi.org/10.1007/s10805-025-09607-1

OECD. (2026). OECD digital education outlook 2026: Exploring effective uses of generative AI in education. https://doi.org/10.1787/062a7394-en

Ogunleye, B., Zakariyyah, K. I., Ajao, O., Olayinka, O., & Sharma, H. (2024). A systematic review of generative AI for teaching and learning practice. Education Sciences, 14(6), Article 636. https://doi.org/10.3390/educsci14060636

OpenAI. (2023). ChatGPT (GPT-3.5). https://chatgpt.com/

OpenAI. (2025). ChatGPT (GPT-5). https://chatgpt.com/

Page, E. B. (2003). Project essay grade: PEG. In M. D. Shermis & J. Burstein (Eds.), Automated essay scoring: A cross-disciplinary perspective (pp. 43–54). Lawrence Erlbaum Associates.

Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI.

https://cdn.openai.com/better-languagemodels/language_models_are_unsupervised_multitask_learners.pdf

Rana, A., Vaidya, P., & Hu, Y. C. (2026). Transformative role of large language models in education and comparative analysis with traditional technologies. Education and Information Technologies, 31(4), 1137–1162. https://doi.org/10.1007/s10639-025-13854-5

Samala, A. D., Rawas, S., Wang, T., Reed, J. M., Kim, J., Howard, N.-J., & Ertz, M. (2024). Unveiling the landscape of generative artificial intelligence in education: A comprehensive taxonomy of applications, challenges, and future prospects. Education and Information Technologies, 30, 3239–3278. https://doi.org/10.1007/s10639-024-12936-0

Shi, H., & Aryadoust, V. (2024). A systematic review of AI-based automated written feedback research. ReCALL, 36(2), 187–209.

https://doi.org/10.1017/S0958344023000265

Steiss, J., Tate, T., Graham, S., Cruz, J., Hebert, M., Wang, J., & Olson, C. B. (2024). Comparing the quality of human and ChatGPT feedback on students’ writing. Learning and Instruction, 91, Article 101894. https://doi.org/10.1016/j.learninstruc.2024.101894

Uludag, P., & McDonough, K. (2022). Validating a rubric for assessing integrated writing in an EAP context. Assessing Writing, 52, Article 100609. https://doi.org/10.1016/j.asw.2022.100609

van Niekerk, J., Delport, P. M., & Sutherland, I. (2025). Addressing the use of generative AI in academic writing. Computers and Education: Artificial Intelligence, 8, Article 100342. https://doi.org/10.1016/j.caeai.2024.100342

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. https://proceedings.neurips.cc/paper/7181-attention-is-all-you-need

Wang, D. (2024). Teacher- versus AI-generated (Poe application) corrective feedback and language learners’ writing anxiety, complexity, fluency, and accuracy. The International Review of Research in Open and Distributed Learning, 25(3), 37–56. https://doi.org/10.19173/irrodl.v25i3.7646

Zhai, C., Wibowo, S., & Li, L. D. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: A systematic review. Smart Learning Environments, 11(1), Article 28. https://doi.org/10.1186/s40561-024-00316-7

Zhang, Z., Aubrey, S., Huang, X., & Chiu, T. K. F. (2025). The role of generative AI and hybrid feedback in improving L2 writing skills: A comparative study. Innovation in Language Learning and Teaching. Advance online publication. https://doi.org/10.1080/17501229.2025.2503890

Downloads

Published

2026-10-02

Issue

Section

Articles

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.