6 papers
Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
Bingsen Chen, Boyan Li, Ping Nie +3
Existing benchmarks for Deep Research Agents (DRAs) treat report generation as a single-shot writing task, which fundamentally diverges from how human researchers iteratively draft…
Injecting Knowledge from Social Science Journals to Improve Indonesian Cultural Understanding by LLMs
Adimulya Kartiyasa, Bao Gia Cao, Boyang Li
Recently there have been intensifying efforts to improve the understanding of Indonesian cultures by large language models (LLMs). An attractive source of cultural knowledge that h…
Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs
Weihao Hong, Zhiyuan Jiang, Bingyu Shen +4
Vision-Language Models (VLMs) are increasingly used in safety-critical applications that require reliable visual grounding. However, these models often hallucinate details that are…
Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
Jin-Seop Lee, SungJoon Lee, SeongJun Jung +2
Video Temporal Grounding (VTG) aims to localize a temporal segment in a video corresponding to a natural language query. However, existing VTG models assume that a relevant segment…
Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
Zhuoyi Yang, Xu Guo, Tong Zhang +2
With this paper, we survey techniques for improving the predictive accuracy of pretrained large language models by allocating additional compute at inference time. In categorizing…
Two Causally Related Needles in a Video Haystack
Miaoyu Li, Qin Chao, Boyang Li
Properly evaluating the ability of Video-Language Models (VLMs) to understand long videos remains a challenge. We propose a long-context video understanding benchmark, Causal2Needl…