From the 1 of 8 linked papers with an AI index.
8 papers
Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
Jianmin Chen, Jiaqi Tang, Wei Wei +9
Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthen, models may gradually rely…
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Jinjian Wu, Jiaqi Tang, Wei Wei +5
The paper introduces IQA-T1, a framework that combines multimodal large language models with specialized visual analysis tools to generate explicit evidence (e.g., noise residual m…
Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?
Jiaqi Tang, Jianmin Chen, Youyang Zhai +6
Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real-world visual corrupti…
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
Ke Ma, Jiaqi Tang, Bin Guo +8
Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall short due to their implicit, quer…
MedHorizon: Towards Long-context Medical Video Understanding in the Wild
Bodong Du, Bowen Liu, Yang Yu +8
Medical multimodal large language models (MLLMs) have advanced image understanding and short-video analysis, but real clinical review often requires full-procedure video understand…
LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
Jiaqi Tang, Yu Xia, Yi-Feng Wu +9
The advent of autonomous agents is transforming interactions with Graphical User Interfaces (GUIs) by employing natural language as a powerful intermediary. Despite the predominanc…