activity
20242026
most citedVL-Mamba: Exploring State Space Models for Multimodal Learning

10 citations · 16 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

TimeThink: Reasoning with Time for Video LLMs

Handong Li, Longteng Guo, Zikang Liu +8

Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…

cs.CV2026

Semantic-Enriched Latent Visual Reasoning

Tianrun Xu, Yue Sun, Qixun Wang +8

Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…

cs.CV2026

Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning

Longteng Guo, Yifan Wang, Pengkang Huo +4

Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning…

cs.CV2026

M-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

Jiatong Ma, Longteng Guo, Yuchen Liu +4

We present M-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multi…

cs.CV2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

Handong Li, Zikang Liu, Longteng Guo +10

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception throug…

cs.CV2025

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Tongtian Yue, Longteng Guo, Yepeng Tang +4

Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Curr…