2026 ASEE Annual Conference & Exposition

A Rigorous Evaluation of Agentic Large Language Model Workflows for Multiple-Choice Question Generation in Advanced Engineering Courses

Presented at Electrical and Computer Engineering Division (ECE) Technical Session 4

Large language models (LLMs) present a transformative opportunity for automating multiple-choice question (MCQ) generation in engineering courses, where maintaining large, high-quality question banks aligned with Bloom’s Taxonomy remains a labor-intensive process. However, while recent work demonstrates that LLMs can generate pedagogically relevant MCQs, there remains limited understanding of how agentic workflows—those that allow models to reason and act iteratively—affect question quality, validity, and alignment with learning outcomes. This study rigorously evaluates a modular agentic workflow that applies an LLM (Gemini-2.5-Flash) to preprocess course materials, generate topics, create MCQs, and iteratively refine them through structured feedback loops.

An automated evaluation framework was developed to assess question quality across four pedagogically grounded dimensions: (1) item writing flaws (IWFs), using a 19-factor rubric adapted from nursing education research; (2) diversity, measured via Distinct-3 analysis; (3) content alignment, measured by semantic similarity to lecture transcripts; and (4) Bloom’s Taxonomy alignment, evaluated using a separate LLM-based classifier achieving 77.4% test accuracy. Experimental data were drawn from lecture transcripts of a third-year FPGA design course, with over 1,500 questions generated across all six Bloom levels.

Results indicate that agentic iteration markedly improves MCQ quality, with most gains achieved in the first two refinement passes. Topic-guided generation enhances quality for small batch sizes but degrades performance when too many topics are processed concurrently, likely due to attention diffusion. Preprocessing transcripts had negligible impact, suggesting modern LLMs are robust to minor transcription noise. Overall, 54% of generated questions met the pedagogical standard of fewer than two IWFs, with strong performance at lower Bloom levels but diminishing reliability for higher cognitive tasks.

These findings demonstrate that agentic LLM workflows can significantly enhance automated MCQ generation when carefully structured and limited to effective iterative stages. The study contributes an open, reproducible framework for systematically evaluating AI-generated assessment items and provides empirically grounded design principles for integrating LLMs into large-scale engineering education assessment pipelines.

Authors
  1. Eric Robert Tourigny University of Calgary
  2. Armando Miguel Acosta Simancas University of Calgary
  3. Dr. Denis Onen University of Calgary [biography]
  4. Miriam Nightingale University of Calgary
  5. Ethan MacDonald University of Calgary
Download paper (585 KB)

Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.