collaborators

5 papers

cs.AI2026

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

Wencheng Ye, Yi Bin, Yujuan Ding +7

Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening eviden…

cs.AI2026

RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering

Wencheng Ye, Xiaoyang Yuan, Yi Bin +4

Recent work on domain-specific reasoning with large language models (LLMs) often relies on training-intensive approaches that require parameter updates. While activation steering h…

cs.AI2025

D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies

Sen Chen, Tong Zhao, Yi Bin +3

Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial Ge…

cs.CV2025

Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping

Yue Yang, Shuibai Zhang, Wenqi Shao +4

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across multimodal tasks such as visual perception and reasoning, leading to good performance on vario…

cs.LG2025

PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization

Mengzhao Chen, Yi Liu, Jiahao Wang +3

Existing weight-activation quantization methods for Large Language Models (LLMs) primarily address channel-wise outliers but often neglect token-wise outliers, which limits the acc…