Large language models (LLMs) such as ChatGPT are widely used by engineering students to complete coursework, but current literature lacks empirical evidence on whether engineering students can recognize when LLMs generate confident but incorrect outputs (hallucinations). While LLMs can provide explanations, guidance, and feedback, they can also provide incorrect solutions, raising concerns for engineering education. This study addresses this gap through a sequential design consisting of (1) measurement of GPT-5.1 hallucination rates and illustrative examples using undergraduate mechanics and electricity and magnetism (E&M) problems, and (2) a survey of engineering students’ ability to detect these hallucinations in LLM-generated solutions. To our knowledge, this is the first study focused specifically on engineering students’ detection and confidence in LLM hallucinations in problem solving.
The authors generate benchmark data from GPT-5.1 to map failure modes of LLMs on quantitative tasks, quantify error frequencies, and develop examples of errors for use in the student survey. The authors survey engineering students’ ability to detect LLM hallucinations in engineering problem solving. Specifically, the survey presents GPT-5.1-generated solutions to mechanics and E&M problems—some correct, some containing hallucinations—and asks students to accept or reject each solution and rate their confidence. Primary outcomes include hallucination detection accuracy and corresponding confidence, while secondary analyses examine associations with self-reported artificial intelligence (AI) use, prior coursework, and background characteristics. Results indicate that higher AI usage frequency is significantly correlated with lower hallucination detection accuracy, and higher self-reported confidence is marginally correlated with lower hallucination detection accuracy.
By combining empirical LLM-generated benchmark data and a student survey, this work (1) quantifies how reliably engineering students detect LLM hallucinations, (2) identifies which error types are most frequent, and (3) offers practical insights for developing hallucination detection strategies in engineering education.
Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.