This study advances a multimodal analytic framework for investigating embodied cognition during oral explanations in statistics education. Grounded in theories of embodied, situated, and distributed cognition, gesture is conceived not just as an accessory to speech but also as reasoning that is enacted through coordinated movement and language. Using computer vision, gesture trajectories, rhythm, and spatial anchoring are captured, while speech-to-text large language models (LLMs) transcribe and semantically analyze verbal explanations. A central Inference Agent integrates these modalities to reveal how gesture and discourse converge or diverge as indicators of conceptual understanding. Rather than claiming autonomous interpretation, the system functions as an epistemic instrument that visualizes the coupling between gesture and conceptual understanding. By aligning theories of embodied cognition with computational observation, this work reframes oral assessment as a dialogic event in which knowing unfolds through motion, voice, and interpretation—positioning understanding itself as an embodied relation between cognition, expression, and inference.
http://orcid.org/0000-0001-5971-214X
Purdue University – West Lafayette
[biography]
http://orcid.org/https://0009-0004-6441-3805
Purdue University – West Lafayette (College of Engineering)
[biography]
Are you a researcher? Would you like to cite this paper? Visit the ASEE document repository at peer.asee.org for more tools and easy citations.