7 papers
HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents
Hongbo Jin, Rongpeng Zhu, Jiayu Ding +2
Large language model (LLM) agents have recently demonstrated strong capabilities in interactive decision-making, yet they remain fundamentally limited in long-horizon tasks that re…
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
Hongbo Jin, Siyi Xie, Jiayu Ding +2
While Large Video-Language Models (Video-LLMs) have achieved remarkable progress in perception, their reasoning capabilities remain a bottleneck. Existing solutions typically resor…
3D Instruction Ambiguity Detection
Jiayu Ding, Haoran Tang, Hongbo Jin +2
In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like "Pass me the vial" in a surgical setting could lead to catastrophic errors. Yet,…
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
Hongbo Jin, Kuanwei Lin, Wenhao Zhang +2
Reinforcement Learning (RL) is crucial for empowering VideoLLMs with complex spatiotemporal reasoning. However, current RL paradigms predominantly rely on random data shuffling or…
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
Hongbo Jin, Qingyuan Wang, Wenhao Zhang +2
Ultra long video understanding remains an open challenge, as existing vision language models (VLMs) falter on such content due to limited context length and inefficient long term m…
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
Xuyi Yang, Wenhao Zhang, Hongbo Jin +5
Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video…