collaborators

5 papers

cs.CL2025

AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment

Ruibo Deng, Duanyu Feng, Wenqiang Lei

Offline preference optimization offers a simpler and more stable alternative to RLHF for aligning language models. However, their effectiveness is critically dependent on ranking a…

cs.AI2025

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

Youcheng Huang, Bowen Qin, Chen Huang +3

Large Reasoning Models (LRMs) have demonstrated remarkable problem-solving abilities in mathematics, as evaluated by existing benchmarks exclusively on well-defined problems. Howev…

cs.CL2025

Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts

Youcheng Huang, Chen Huang, Duanyu Feng +2

Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior research has shown that a single LLM's concept representations can be captur…

cs.CL2024

Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets

Duanyu Feng, Bowen Qin, Chen Huang +3

The success of the reward model in distinguishing between responses with subtle safety differences depends critically on the high-quality preference dataset, which should capture t…

cs.CV2024

Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector

Youcheng Huang, Fengbin Zhu, Jingkun Tang +4

Visual Language Models (VLMs) are vulnerable to adversarial attacks, especially those from adversarial images, which is however under-explored in literature. To facilitate research…