Since the release of ChatGPT 3.5, discussions of using AI in education within the educational community have gradually shifted from whether AI should be used in classrooms to how it can be effectively and responsibly integrated into STEM education. While the state-of-the-art AI models have shown remarkable capabilities, one of t most serious challenges in their application to STEM education is hallucination, where models generate plausible but incorrect or fabricated information. Since STEM learning requires precise understanding of theories, formulas, mathematical calculations, and scientific reasoning, the reliability of AI-generated responses becomes critical. Without an objective measure of correctness, students risk adopting misleading information that undermines conceptual understanding.
To address this challenge, this study introduces a “confidence index”, a quantitative measure representing the probability that an AI system provides accurate responses/solutions across different knowledge points within a specific domain. By systematically analyzing AI responses to structured STEM problems, the study constructs a confidence index map illustrating reliability levels across topics. This map serves as an evidence-based guide for both educators and learners, promoting critical engagement with AI tools rather than passive acceptance of their outputs. Ultimately, this framework aims to enhance AI literacy, support analytical thinking, and support responsible AI adoption in STEM education.
This research examined the reliability of two state-of-the-art AI models (Google’s Gemini 2.5 Pro and OpenAI’s GPT-5) in solving electrical circuit problems from Electric Circuits (12th Edition) by Nilsson and Riedel. A total of 360 questions were systematically selected from the book’s 18 chapters with 20 questions per chapter. The questions consisted of text-based questions and image-based questions containing circuit diagrams and plots. The confidence index was defined as the proportion of problems where an AI achieves correct solutions/answers across all applicable evaluation categories including diagram comprehension, stepwise problem decomposition, formula selection and calculation, circuit configuration analysis, equation use and numerical accuracy, and plot interpretation. We calculated confidence indices with logistic regression and chi-squared tests across three image conditions (circuit diagrams, plots/waveforms, no images) and all knowledge domains, including pairwise AI comparisons and AI×image interactions.
The preliminary results show that 1) GPT-5 achieved a 9.1% higher confidence score than Gemini 2.5 Pro with statistical significance; 2) Both models performed similarly (~90.7% confidence) on text-only problems; 3) Gemini 2.5 Pro showed slightly better performance in waveform and plot interpretation, though the difference was not statistically significant; 4) GPT-5 performed significantly better on problems involving Kirchhoff’s Laws, with a 64.3% confidence advantage; 5) GPT-5 showed moderate (~15%) advantages in fundamental circuit concepts such as transient response analysis; 6) Gemini 2.5 Pro performed about 20% better in topics involving operational amplifiers and frequency transforms, although these differences were not statistically significant.
In summary, GPT-5 achieved an overall accuracy rate of 80%, compared with 74.7% for Gemini 2.5 Pro. GPT-5 showed strength in circuit diagram interpretation and fundamental circuit analysis. Gemini demonstrated competitive performance for text-based problems and potential advantages in topics like operational amplifiers and frequency-domain analysis. The confidence indices and mapping support informed AI use in a class and promote responsible adoption in learning with critical thinking.
Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.