5 papers
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
Qinghongbing Xie, Zhaoyuan Xia, Feng Zhu +4
Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Exist…
Mitigating Instance Entanglement in Instance-Dependent Partial Label Learning
Rui Zhao, Bin Shi, Kai Sun +1
Partial label learning is a prominent weakly supervised classification task, where each training instance is ambiguously labeled with a set of candidate labels. In real-world scena…
On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
Xiaxu Chen, Wei Li, Chunxu Liu +5
Reinforcement Fine-Tuning (RFT) is proved to be greatly valuable for enhancing the reasoning ability of LLMs. Researchers have been starting to apply RFT to MLLMs, hoping it will a…
On the Robustness of Human-Object Interaction Detection against Distribution Shift
Chi Xie, Shuang Liang, Jie Li +4
Human-Object Interaction (HOI) detection has seen substantial advances in recent years. However, existing works focus on the standard setting with ideal images and natural distribu…
Re-Aligning Language to Visual Objects with an Agentic Workflow
Yuming Chen, Jiangyan Feng, Haodong Zhang +6
Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve LOD model generalizations. During…