works on

From the 3 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

Zhuoran Zhang, Bowen Li, Jingcheng Ju +5

GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact solution by compressing multim…

cs.CV2026

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Wanshun Su, Yang Shi, Feihu Liu +10

Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio tok…

cs.CV2026

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Qixun Wang, Yang Shi, Letian Cheng +11

The paper introduces Beacon, an agentic visual reasoning system that learns when to invoke external tools and how to use them effectively, improving multimodal large language model…

cs.CV2026

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Tengfei Liu, Yang Shi, Yuran Wang +16

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-…

cs.CV2026

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Yuqi Tang, Tengfei Liu, Yizheng Lai +18

The paper introduces KeyFrame-Compass, a benchmark and evaluation framework for assessing how well video generation models follow supplied keyframes while preserving overall video…

cs.CV2026

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

Tengfei Liu, Yang Shi, Xuanyu Zhu +17

Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing b…