activity
20242026
collaborators

5 papers

cs.CV2026

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

Jingqi Tian, Haoji Zhang, Lin Chen +7

Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large lan…

cs.CV2026

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

Jingqi Tian, Yiheng Du, Haoji Zhang +6

Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existing methods often struggle with m…

cs.CV2026

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

Shilin Ma, Chubin Zhang, Changyuan Wang +6

Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most…

cs.CV2025

Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation

Tianrui Zhu, Shiyi Zhang, Zhirui Sun +2

Frame-level autoregressive (frame-AR) models have achieved significant progress, enabling real-time video generation comparable to bidirectional diffusion models and serving as a f…

cs.CV2024

Ponder & Press: Advancing Visual GUI Agent towards General Computer Control

Yiqin Wang, Haoji Zhang, Jingqi Tian +1

Most existing GUI agents typically depend on non-vision inputs like HTML source code or accessibility trees, limiting their flexibility across diverse software environments and pla…