Using Large Language Models for Automated Grading of Student Writing about Science

被引:0
|
作者
Impey, Chris [1 ]
Wenger, Matthew [1 ]
Garuda, Nikhil [1 ]
Golchin, Shahriar [2 ]
Stamer, Sarah [1 ]
机构
[1] Univ Arizona, Dept Astron, Tucson, AZ 85721 USA
[2] Univ Arizona, Dept Comp Sci, Tucson, AZ 85721 USA
基金
美国国家科学基金会;
关键词
Student writing; Science classes; Online education; Assessment; Machine learning; Large language models; ONLINE; ASTRONOMY; RATER;
D O I
10.1007/s40593-024-00453-7
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Assessing writing in large classes for formal or informal learners presents a significant challenge. Consequently, most large classes, particularly in science, rely on objective assessment tools such as multiple-choice quizzes, which have a single correct answer. The rapid development of AI has introduced the possibility of using large language models (LLMs) to evaluate student writing. An experiment was conducted using GPT-4 to determine if machine learning methods based on LLMs can match or exceed the reliability of instructor grading in evaluating short writing assignments on topics in astronomy. The audience consisted of adult learners in three massive open online courses (MOOCs) offered through Coursera. One course was on astronomy, the second was on astrobiology, and the third was on the history and philosophy of astronomy. The results should also be applicable to non-science majors in university settings, where the content and modes of evaluation are similar. The data comprised answers from 120 students to 12 questions across the three courses. GPT-4 was provided with total grades, model answers, and rubrics from an instructor for all three courses. In addition to evaluating how reliably the LLM reproduced instructor grades, the LLM was also tasked with generating its own rubrics. Overall, the LLM was more reliable than peer grading, both in aggregate and by individual student, and approximately matched instructor grades for all three online courses. The implication is that LLMs may soon be used for automated, reliable, and scalable grading of student science writing.
引用
收藏
页数:35
相关论文
共 50 条
  • [1] Automated Grading in Coding Exercises Using Large Language Models
    Lagakis, Paraskevas
    Demetriadis, Stavros
    Psathas, Georgios
    SMART MOBILE COMMUNICATION & ARTIFICIAL INTELLIGENCE, VOL 1, IMCL 2023, 2024, 936 : 363 - 373
  • [2] Automated Grading of Students’ Short Answers Using Language Models
    Ch. B. Minnegalieva
    I. I. Kashapov
    O. D. Morozova
    Automatic Documentation and Mathematical Linguistics, 2024, 58 (Suppl 3) : S109 - S114
  • [3] Sparks: Inspiration for Science Writing using Language Models
    Gero, Katy Ilonka
    Liu, Vivian
    Chilton, Lydia B.
    PROCEEDINGS OF THE 2022 ACM DESIGNING INTERACTIVE SYSTEMS CONFERENCE, DIS 2022, 2022, : 1002 - 1019
  • [4] Sparks: Inspiration for Science Writing using Language Models
    Gero, Katy Ilonka
    Liu, Vivian
    Chilton, Lydia B.
    PROCEEDINGS OF THE FIRST WORKSHOP ON INTELLIGENT AND INTERACTIVE WRITING ASSISTANTS (IN2WRITING 2022), 2022, : 83 - 84
  • [5] The Role of First Language in Automated Essay Grading for Second Language Writing
    Hwang, Haerim
    ARTIFICIAL INTELLIGENCE IN EDUCATION, PT II, AIED 2024, 2024, 14830 : 302 - 310
  • [6] PERSUASIVE LEGAL WRITING USING LARGE LANGUAGE MODELS
    Curran, Damian
    Levy, Inbar
    Mistica, Meladel
    Hovy, Eduard
    LEGAL EDUCATION REVIEW, 2024, 34 (01):
  • [7] EvaAI: A Multi-agent Framework Leveraging Large Language Models for Enhanced Automated Grading
    Lagakis, Paraskevas
    Demetriadis, Stavros
    GENERATIVE INTELLIGENCE AND INTELLIGENT TUTORING SYSTEMS, PT I, ITS 2024, 2024, 14798 : 378 - 385
  • [8] Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability
    Pack A.
    Barrett A.
    Escalante J.
    Computers and Education: Artificial Intelligence, 2024, 6
  • [9] Large language models in science
    Kowalewski, Karl-Friedrich
    Rodler, Severin
    UROLOGIE, 2024, 63 (09): : 860 - 866
  • [10] Automated Natural Language Explanation of Deep Visual Neurons with Large Models (Student Abstract)
    Zhao, Chenxu
    Qian, Wei
    Shi, Yucheng
    Huai, Mengdi
    Liu, Ninghao
    THIRTY-EIGTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 21, 2024, : 23712 - 23713