paper

Beyond Text: Probing K-12 Educators' Perspectives and Ideas for Learning Opportunities Leveraging Multimodal Large Language Models

arXiv:2507.20720

Abstract

Multimodal Large Language Models (MLLMs) are beginning to enable new user experiences from generated content across a range of media, including images, text, speech, and video. These capabilities have the potential to enrich learning by enabling users to interact with information using a variety of modalities, but little is known about how \textit{educators} envision how MLLMs might shape the future of learning, what challenges they encounter when using these models, and what practical needs should be considered for future implementation in educational contexts. We investigated educator perspectives through workshops with 12 K-12 educators, where participants brainstormed learning opportunities, discussed practical concerns, and prototyped MLLM learning applications using Claude 3.5 and its Artifacts feature. Through this work, we uncover how educators imagined MLLMs as a way for themselves and their students to author multimedia content, and how this could provide a pathway to support learning through iterative design. At the same time, educators anticipated challenges with younger students' ability to evaluate and refine model output to better meet their design goals. We end with implications for designing with and for MLLMs in future learning experiences.

Beyond Text: Probing K-12 Educators' Perspectives and Ideas for Learning Opportunities Leveraging Multimodal Large Language Models · wovepaper