2026 ASEE Annual Conference & Exposition

Evaluating Creativity and Novelty in Large Language Model–Generated Design Assignments for Digital Systems Education

Presented at DSAI-Session 8: Generative AI in Assessment, Grading, and Accreditation

The increasing sophistication of large language models (LLMs) presents new opportunities to enhance curriculum design and project generation in engineering education. This research investigates how LLMs can support educators in creating diverse and original student design assignments, with a focus on digital logic systems courses that emphasize synchronous finite state machine (FSM) design. The study evaluates the creativity and novelty of assignment prompts generated by advanced LLMs—including GPT-4o, Gemini, LLaMA 3.2, and DeepSeek—for the EEE120: Digital Design Fundamentals course at Arizona State University. Each model was prompted to produce capstone-style FSM design projects incorporating Moore models with at least five states, constrained by standard course objectives and prior semester examples.

To quantify the distinctiveness of generated outputs, the study employs cosine similarity using Sentence-BERT embeddings to compare LLM-generated prompts with a baseline corpus of human-designed assignments. A novelty metric derived from Chakrabarti and Khadilkar’s framework is applied to assess structural and conceptual innovation, while qualitative evaluations by faculty assess logical consistency and pedagogical appropriateness. Both retrieval-augmented generation (RAG) and non-RAG configurations are tested to determine whether contextual grounding in course materials enhances creativity or coherence. Results demonstrate that RAG-enhanced models generally yield more pedagogically relevant and logically consistent prompts, whereas non-RAG setups tend to produce more conceptually divergent—and occasionally less feasible—design challenges.

Statistical and qualitative analyses reveal that GPT-4o and Gemini balance novelty with constraint adherence most effectively, producing assignments with moderate cosine similarity (0.45–0.60) to baseline data and strong faculty ratings for creativity and clarity. LLaMA 3.2 and DeepSeek exhibited higher novelty scores but lower pedagogical alignment, emphasizing the trade-off between innovation and instructional soundness. These findings suggest that hybrid prompting strategies, where educators iteratively refine AI outputs, can achieve optimal creativity without compromising curricular integrity.

This work contributes to ASEE community by presenting a reproducible, data-driven framework for evaluating the creativity of AI-generated educational materials. Beyond measuring performance, the study highlights how educators can leverage AI responsibly and effectively to create relevant and unique design prompts, decreasing the usage of repeated curriculum. The integration of cosine similarity, semantic embeddings, and novelty scoring establishes a scalable model for continuous curriculum improvement and instructor support in STEM education. Overall, this study underscores the role of AI as a catalyst for innovation in educational design and assessment.

Authors
  1. Mrs. Alicia Baumann Arizona State University [biography]
  2. Sai Vignesh Naragoni Arizona State University [biography]
Download paper (983 KB)

Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.