94 citations · 94 across the 14 of their papers we have counts for
6 papers · 1 filter
VideoResearcher: Self-Improving Tool Design for Long-Video Understanding
Dingqiang Ye, Dongdi Zhao, Kaishen Wang +11
Video agents have made substantial progress in long-video understanding. Yet effective video-agent systems require costly, time-consuming manual design and trial and error. Current…
Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding
Kaishen Wang, Dongdi Zhao, Yijun Liang +4
Vision-language models (VLMs) have made substantial progress in long-video understanding, with standard backbone models typically answering questions from frames sampled across the…
Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization
Kaishen Wang, Tong Zheng, Xuehao Cui +3
Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such…
Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention
Chaowen Yan, Kaishen Wang, Yong Wang +2
Oracle Bone Inscriptions (OBIs) recognition plays a crucial role in understanding ancient Chinese culture. However, accurately recognizing OBIs remains highly challenging due to th…
Unsafe by Reciprocity: How Generation-Understanding Coupling Undermines Safety in Unified Multimodal Models
Kaishen Wang, Heng Huang
Recent advances in Large Language Models (LLMs) and Text-to-Image (T2I) models have led to the emergence of Unified Multimodal Models (UMMs), where multimodal understanding and ima…
Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing
Tong Zheng, Chengsong Huang, Runpeng Dai +9
Parallel thinking has emerged as a promising paradigm for reasoning, yet it imposes significant computational burdens. Existing efficiency methods primarily rely on local, per-traj…