This study examines the effectiveness and efficiency of Pensive (formerly Pensieve Grader), an AI-assisted grading system that extends the capabilities of Gradescope for evaluating handwritten student work. Built on large language model (LLM) technology, Pensive automates the entire grading workflow—from transcription to rubric-based evaluation and feedback—within a human-in-the-loop interface. Although the system has been deployed at more than 20 institutions and has graded over 300,000 student responses across computer science, mathematics, physics, and chemistry, prior work shows that its adoption in physics-related fields remains limited (approximately 6.7%). This study investigates whether such limited use reflects concerns about effectiveness and efficiency or simply a lack of familiarity within subjects similar to physics.
The effectiveness of Pensive was evaluated by comparing its grading outcomes with instructor assessments across multiple problem types, including drawings (e.g., free-body diagrams (FBDs)), multiple-choice questions, and multi-step problem-solving tasks. Agreement metrics were analyzed to determine how well the system captured both handwriting recognition and grading accuracy. For diagrammatic and open-ended responses, qualitative rubrics were applied to assess Pensive’s recognition of key features such as correct force representation and labeling. For quantitative problems, accuracy was examined through AI-extracted expressions and numerical results relative to instructor solutions. The study also explored how the worksheet outline, rubric design, problem structure, and calibration procedures influence grading performance.
The efficiency of Pensive was assessed in terms of grading time reduction and feedback turnaround compared to traditional manual grading. Building on the findings of Yang et al.[1], who reported an average 65% reduction in grading time across STEM disciplines, this study investigates whether similar levels of efficiency can be achieved in the context of engineering dynamics. The analysis focuses on how factors such as the worksheet outlines, rubric design, calibration, and problem format may influence the degree of time savings and usability for instructors.
Findings indicate that grading effectiveness and efficiency varied by task type. Pensive achieved high agreement with instructor grading on multiple choice and short answer questions, and reached approximately 98% agreement on analytical problems after calibration. Drawing tasks showed the greatest variability, with agreement differing substantially across structurally similar problems. These results suggest that AI-assisted grading shifts rather than uniformly reduces instructor effort, with benefits depending on task type and the reusability of calibration across assessments. By providing empirical data from handwritten dynamics exams, this work contributes to understanding how LLM-assisted grading performs in symbol-heavy, spatially complex problem types rarely represented in prior Pensive studies. The results will clarify whether barriers to broader adoption stem from technical limitations or awareness gaps, offering evidence-based insights for integrating AI-assisted grading tools into STEM education.
http://orcid.org/0000-0002-9089-5746
Embry-Riddle Aeronautical University - Daytona Beach
[biography]
Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.