5 papers
TGV-KV: Text-Grounded KV Eviction for Vision-Language Models
Jizhihui Liu, Ruizi Han, Miao Zhang +4
Vision-Language Models (VLMs) inherit the auto-regressive generation paradigm and cache the keys and values (KV) of all previous tokens to accelerate inference, resulting in memory…
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
Wei Li, Jizhihui Liu, Li Yixing +3
Current Vision-Language-Action (VLA) models primarily focus on mapping 2D observations to actions, but exhibit notable limitations in spatiotemporal perception and reasoning: 1) sp…
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
Jizhihui Liu, Feiyi Du, Guangdao Zhu +5
Vision-Language Models (VLMs) encode images and videos into abundant tokens, which contain substantial redundancy and computation cost. While visual token pruning mitigates the iss…
H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation
Yijie Zhu, Rui Shao, Ziyang Liu +4
Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how intera…
FM2S: Towards Spatially-Correlated Noise Modeling in Zero-Shot Fluorescence Microscopy Image Denoising
Jizhihui Liu, Qixun Teng, Qing Ma +1
Fluorescence microscopy image (FMI) denoising faces critical challenges due to the compound mixed Poisson-Gaussian noise with strong spatial correlation and the impracticality of a…