4 papers · 1 filter
Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
Hongyu Qu, Guangming Yao, Ling Xing +7
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded m…
See the Text: From Tokenization to Visual Reading
Ling Xing, Rui Yan, Alex Jinpeng Wang +2
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle ty…
Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition
Hongyu Qu, Ling Xing, Jiachao Zhang +3
Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level representations for each video by desi…
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
Ling Xing, Hongyu Qu, Rui Yan +2
Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where even…