collaborators

10 papers

cs.AI2026

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Yanchao Li, Wanhao Liu, Jiaqing Xie +4

Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produc…

cs.AI2026

Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT

Peng Sun, Yi Yang, Huawen Shen +4

Visual instruction tuning is crucial for improving vision-language large models (VLLMs). However, many samples can be solved via linguistic patterns or common-sense shortcuts, with…

cs.CY2026

On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Yue Huang, Chujie Gao, Siyuan Wu +63

Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions.…

cs.LG2026

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

Yue Huang, Haomin Zhuang, Jiayi Ye +6

Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yielding safer-on-paper yet less us…

cs.AI2025

AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models

Xiangqi Wang, Yue Huang, Yanbo Wang +4

LLMs often need effective configurations, like temperature and reasoning steps, to handle tasks requiring sophisticated reasoning and problem-solving, ranging from joke generation…

cs.CL2025

Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search

Yanbo Wang, Zixiang Xu, Yue Huang +6

Large Language Models (LLMs) often struggle to maintain their original performance when faced with semantically coherent but task-irrelevant contextual information. Although prior…