This research expands on the use of Large Language Models (LLM), to evaluate and provide feedback on student papers. This paper uses a custom GPT created through OpenAI’s ChatGPT. Previous work by the authors found that there is significant time savings from using an LLM to assist in providing feedback to students but there was not a strong agreement between the evaluator ratings and the LLM ratings. The authors hypothesize that the LLM would be in better agreement with the evaluators and as an aid for grading the assignments using rubrics developed for the course. The research methodology for this study begins with having two evaluators (instructors) grade student assignments independently, then comparing their evaluations in a norming exercise. Next, predefined assignment responses were created with the help of ChatGPT to build background information for the custom GPT. These responses are based on the existing rubric and help increase the reliability of the ratings against the rubric. The last step compares the LLM evaluations of student assignments to those provided by the evaluators. This is used to determine the degree of agreement between the faculty evaluators and the LLM using Krippendorff’s-α and Cohen’s Kappa test methods. The authors expect the results to show a strong level of agreement, with Cohen’s Kappa and Krippendorff’s-α greater than 0.6. It is believed that this will show an improvement in reliability in regard to the LLM’s usability for grading and feedback purposes. Further work will involve a review of the rubric used to determine its effect on Generative AI’s outputs.
Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.