3 papers
cs.HC2026
AutoCue: Multimodal LLM-Assisted Externalization of Implicit Inputs as Instructional Visual Cues in Screencast Tutorials
Shengyang Luo, Shengyao Luo, Xiaolei Guo +3
Tutorial videos are widely used for learning feature-rich software, yet following screencast tutorials often breaks down in practice. Through a survey and contextual inquiry, we fo…
cs.CV2026
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision
Tongyan Wang, Zhengyuan Li, Muhan Lin +5
Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form description…
cs.CV2026
MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks
Jui-Cheng Chiu, Yu-Chao Wang, Shengyang Luo +4
Appreciating multi-figure paintings requires understanding how characters relate through subtle cues like gaze alignment, gesture, and spatial arrangement. We present MIRAGE, an ev…