From the 1 of 16 linked papers with an AI index.
8 papers · 1 filter
TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models
Jinkun Zhao, Kui Zhang, Wenjun Wu
The paper introduces Tone‑Pressure Contrastive Decoding (TPCD), which subtracts logits from high‑pressure prompts from those of neutral prompts to reduce commitment bias in vision‑…
RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation
Sen Zhang, Runmei Li, Shizhuang Deng +7
As Automatic Train Operation (ATO) advances toward GoA4 and beyond, it increasingly depends on efficient, reliable cab-view visual perception and decision-oriented inference to ens…
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
Siwei Wen, Zhangcheng Wang, Xingjian Zhang +2
Online video understanding requires models to perform continuous perception and long-range reasoning within potentially infinite visual streams. Its fundamental challenge lies in t…
VoQA: Visual-only Question Answering
Jianing An, Luyang Jiang, Jie Luo +2
Visual understanding requires interpreting both natural scenes and the textual information that appears within them, motivating tasks such as Visual Question Answering (VQA). Howev…
What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt
Jinkun Zhao, Lei Huang, Haixin Ge +1
We introduce Embedded Stroop, a controlled diagnostic paradigm for measuring image-embedded prompt interference in Multimodal Large Language Models (MLLMs), where the query is rend…
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation
Siwei Wen, Junyan Ye, Peilin Feng +7
With the rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, synthetic images have become increasingly prevalent in everyday life, posing new challe…