collaborators

14 papers

cs.RO2026

Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning

Shilin Shan, Chuhao Zhou, Ruize Wang +30

Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical…

cs.CV2026

YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models

Ting Chen, Geng Li, Guohao Chen +5

Contrastive decoding (CD) seeks to mitigate hallucinations in Large Vision-Language Models (LVLMs) by contrasting the output distributions of a standard model and a visually degrad…

cs.CV2026

OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning

Geng Li, Guohao Chen, Ting Chen +6

Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning…

cs.CV2026

Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion

Gang Dai, Yining Huang, Yiming Xia +2

The efficient Test-Time Scaling (TTS) paradigm offers a promising perspective for enhancing the generation performance of diffusion models. However, current solutions are limited t…

cs.LG2026

ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse

Guohao Chen, Shuaicheng Niu, Deyu Chen +5

Test-time entropy minimization helps adapt a model to novel environments and incentivize its reasoning capability, unleashing the model's potential during inference by allowing it…

cs.LG2026

EVA-0: Test-Time Model Evolution with Only Two Forward Passes per Sample

Guohao Chen, Shuaicheng Niu, Geng Li +4

Test-time model evolution offers a promising way for deployed models to improve from unlabeled test-time experience, yet most existing methods depend on backpropagation (BP), which…