works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding

Haiyue Zhang, Yi Bin, Xun Jiang +5

VisualRouter is a training-free, plug‑and‑play framework that classifies queries as global or local and applies tailored visual sampling strategies to select informative frames, im…

cs.RO2026

Self-Correcting VLA: Online Action Refinement via Sparse World Imagination

Chenyv Liu, Wentao Tan, Lei Zhu +4

Standard vision-language-action (VLA) models rely on fitting statistical data priors, limiting their robust understanding of underlying physical dynamics. Reinforcement learning en…

cs.DB2026

RAC: Relation-Aware Cache Replacement for Large Language Models

Yuchong Wu, Zihuan Xu, Wangze Ni +5

The scaling of Large Language Model (LLM) services faces significant cost and latency challenges, making effective caching under tight capacity crucial. Existing cache replacement…

cs.RO2025

A Step Toward World Models: A Survey on Robotic Manipulation

Peng-Fei Zhang, Ying Cheng, Xiaofan Sun +4

Autonomous agents are increasingly expected to operate in complex, dynamic, and uncertain environments, performing tasks such as manipulation, navigation, and decision-making. Achi…

cs.AI2025

BLM: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

Wentao Tan, Bowen Wang, Heng Zhi +15

Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs ge…

cs.CV2025

Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation

Longzhen Yang, Zhangkai Ni, Ying Wen +3

Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and f…