collaborators

5 papers

cs.LG2026

Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs

Zixuan Chen, Hao Lin, Zizhe Chen +6

LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply rather than correct. We term this…

cs.MM2026

AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction

Zixuan Chen, Depeng Wang, Hao Lin +6

We present AVID, the first large-scale benchmark for audio-visual inconsistency understanding in videos. While omni-modal large language models excel at temporally aligned tasks su…

cs.CV2025

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

Chao Gong, Depeng Wang, Zhipeng Wei +3

Audio-Visual Large Language Models (AV-LLMs) face prohibitive computational costs of processing massive, redundant audio-visual tokens. Existing unimodal compression techniques fai…

cs.AI2025

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning

Taihang Zhen, Jialiang Hong, Kai Chen +12

Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and obscure the underlying reasoning process.…

cs.CL2025

Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting

Yijie Hong, Xiaofei Yin, Xinzhong Wang +7

Large Vision Language Models have demonstrated impressive versatile capabilities through extensive multimodal pre-training, but face significant limitations when incorporating spec…