The use of generative artificial intelligence has become commonplace among students, but less so among their faculty. In this paper, we present our experience using large language models (LLMs) to determine what misconceptions 89 undergraduate students have based on the incorrect responses to a four-part exam problem on CPU scheduling policies. If we are successful, our goal is to incorporate such feedback into a CPU scheduling simulation used by undergraduate Operating Systems students. To answer this question, we first manually analyze each of the 112 incorrect solutions to identify underlying misconceptions and random mistakes. We then ask three production-quality LLMs (ChatGPT, Claude, and Gemini) to do the same task and compare their responses to ours. The LLM responses varied widely in quality and accuracy. The LLMs tended to perform better when students made systematic errors and on simpler policies. Interestingly, the LLMs tended to perform better when the instructors initially disagreed on the misconception, and the error was systematic. Faculty may find this usage of artificial intelligence highly beneficial, but we recommend reviewing the responses before sharing them with students. We thus conclude that the feedback provided by these LLMs is not yet ready for use directly in the CPU simulator.
Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.