most citedInfinite Video Understanding

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2025

Beyond Heuristics: A Decision-Theoretic Framework for Agent Memory Management

Changzhi Sun, Xiangyu Chen, Jixiang Luo +2

External memory is a key component of modern large language model (LLM) systems, enabling long-term interaction and personalization. Despite its importance, memory management is st…

cs.CV2025

Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding

Zhuoming Li, Aitong Liu, Mengxi Jia +5

Free-form gesture understanding is highly appealing for human-computer interaction, as it liberates users from the constraints of predefined gesture categories. However, the sole e…

cs.MM2025

EV-NVC: Efficient Variable bitrate Neural Video Compression

Yongcun Hu, Yingzhen Zhai, Jixiang Luo +4

Training neural video codec (NVC) with variable rate is a highly challenging task due to its complex training strategies and model structure. In this paper, we train an efficient v…

cs.CV20251 cited

Infinite Video Understanding

Dell Zhang, Xiangyu Chen, Jixiang Luo +6

The rapid advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have ushered in remarkable progress in video understanding. However, a fundamental ch…

cs.CV2025

Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations

Yizhen Li, Dell Zhang, Xuelong Li +1

Reasoning Segmentation (RS) is a multimodal vision-text task that requires segmenting objects based on implicit text queries, demanding both precise visual perception and vision-te…

cs.CV2025

SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding

Qianqian Sun, Jixiang Luo, Dell Zhang +1

Recent advancements in image editing have utilized large-scale multimodal models to enable intuitive, natural instruction-driven interactions. However, conventional methods still f…