2 papers
cs.CV2026
Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models
Zhaoyang Li, Yanjun Li, Wangkai Li +2
Vision-Language Models (VLMs) are costly at inference time because they must process long sequences of visual tokens. Existing token pruning methods often degrade under high compre…
cs.CV2025
Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing
Yanjun Li, Zhaoyang Li, Honghui Chen +1
Video Scene Graph Generation (VidSGG) aims to capture dynamic relationships among entities by sequentially analyzing video frames and integrating visual and semantic information. H…