activity
20242026
most citedEchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models

1 citations · 1 across the 5 of their papers we have counts for

collaborators

12 papers

cs.AI2026

Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models

Shi Fu, Yingjie Wang, Shengchao Hu +2

Self-Rewarding Language Models (SRLMs) achieve notable success in iteratively improving alignment without external feedback. Yet, despite their striking empirical progress, the cor…

cs.CV20251 cited

EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models

Botai Yuan, Yutian Zhou, Yingjie Wang +9

Recent benchmarks for medical Large Vision-Language Models (LVLMs) emphasize leaderboard accuracy, overlooking reliability and safety. We study sycophancy -- models' tendency to un…

cs.LG2025

CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction

Yuzhu Chen, Yingjie Wang, Shunyu Liu +2

Autoregressive pre-trained models combined with decoding methods have achieved impressive performance on complex reasoning tasks. While mainstream decoding strategies such as beam…

cs.AI2025

Benchmarking Reasoning Robustness in Large Language Models

Tong Yu, Yongcheng Jing, Xikun Zhang +6

Despite the recent success of large language models (LLMs) in reasoning such as DeepSeek, we for the first time identify a key dilemma in reasoning robustness and generalization: s…

cs.AI2025

Graph-Augmented Reasoning: Evolving Step-by-Step Knowledge Graph Retrieval for LLM Reasoning

Wenjie Wu, Yongcheng Jing, Yingjie Wang +2

Recent large language model (LLM) reasoning, despite its success, suffers from limited domain knowledge, susceptibility to hallucinations, and constrained reasoning depth, particul…

cs.AI2025

Dynamic Parallel Tree Search for Efficient LLM Reasoning

Yifu Ding, Wentao Jiang, Shunyu Liu +9

Tree of Thoughts (ToT) enhances Large Language Model (LLM) reasoning by structuring problem-solving as a spanning tree. However, recent methods focus on search accuracy while overl…