activity
20242026
most citedStreaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

1 citations · 1 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

Yuhao Wang, Mu Qiao, Xindong Zhang +3

GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usage. Training-free visual token pruning can reduce this cost,…

cs.CV2026

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience

Sitong Gong, Caixin Kang, Tianyu Yan +7

A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primari…

cs.CV2026

Revisiting Salient Object Detection from an Observer-Centric Perspective

Fuxi Zhang, Yifan Wang, Hengrun Zhao +7

Salient object detection is inherently a subjective problem, as observers with different priors may perceive different objects as salient. However, existing methods predominantly f…

cs.CV2026

AR-MOT: Autoregressive Multi-object Tracking

Lianjie Jia, Yuhan Wu, Binghao Ran +3

As multi-object tracking (MOT) tasks continue to evolve toward more general and multi-modal scenarios, the rigid and task-specific architectures of existing MOT methods increasingl…

cs.CV2026

Think3D: Thinking with Space for Spatial Reasoning

Zaibin Zhang, Yuhan Wu, Lianjie Jia +10

While Vision-Language Models (VLMs) excel at 2D visual understanding, they remain constrained by 2D-centric paradigm that severely limits genuine 3D spatial reasoning. To bridge th…

cs.CV2025

From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction

Zhida Zhao, Talas Fu, Yifan Wang +2

Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and d…