5 papers
VISUALSKILL: Multimodal Skills for Computer-Use Agents
Ziyan Jiang, Li An, Yujian Liu +5
Computer-use agents (CUAs) approach human-level performance on standardised benchmarks but still struggle on long-horizon tasks and unseen software. Existing skill libraries addres…
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
Qiucheng Wu, Jing Shi, Simon Jenni +4
Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling in…
Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
Guanyu Yao, Qiucheng Wu, Yang Zhang +3
Multimodal large language models (MLLMs) have demonstrated strong capabilities on vision-and-language tasks. However, recent findings reveal an imbalance in their reasoning capabil…
VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
Qiucheng Wu, Handong Zhao, Zhixin Shu +3
Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they sti…
A Hierarchical Probabilistic Framework for Incremental Knowledge Tracing in Classroom Settings
Xinyi Gao, Qiucheng Wu, Yang Zhang +4
Knowledge tracing (KT) aims to estimate a student's evolving knowledge state and predict their performance on new exercises based on performance history. Many realistic classroom s…