2 papers
cs.CY2026
EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content
Shuzhen Bi, Mingzi Zhang, Zhuoxuan Li +3
Large language models are increasingly used as educational assistants, yet evaluation of their educational capabilities remains concentrated on question-answering and tutoring task…
cs.CV2024
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
Chen Bao, Jiarui Xu, Xiaolong Wang +2
How can we predict future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language? In this paper, we exte…