15 papers
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
Chen Ling, Hanqian Li, Dongnan Liu +9
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models t…
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
Chen Ling, Tongwei Zhang, Hanqian Li +1
Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream visual understanding tasks, but their ability to process action scenes that contradi…
Explainable Detection of Machine Generated Music and Early Systematic Evaluation
Yupei Li, Qiyang Sun, Hanqian Li +2
Machine-generated music (MGM) has become a groundbreaking innovation with wide-ranging applications, such as music therapy, personalised editing, and creative inspiration within th…
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
Yibo Shi, Jungang Li, Linghao Zhang +25
Long-horizon GUI agents are a key step toward real-world deployment, yet effective interaction memory under prevailing paradigms remains under-explored. Replaying full interaction…
Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models
Linghao Zhang, Jungang Li, Yonghua Hei +12
Multimodal large language models (MLLMs) are typically trained in multiple stages, with video-based supervised fine-tuning (Video-SFT) serving as a key step for improving visual un…
EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next
Ye Pan, Chi Kit Wong, Yuanhuiyi Lyu +9
Egocentric video provides a natural modality for studying human behavior, but conventional visual understanding captures mainly observable scenes, objects, and actions rather than…