9 papers · 1 filter
Preference Tuning as Spectral Update Reorganization
Peiyan Zhang, Haibo Jin, Liying Kang +1
Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opaque. We study RLHF and related…
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
Lucheng Fu, Ye Yu, Yiyang Wang +4
Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rew…
Controlling Output Rankings in Generative Engines for LLM-based Search
Haibo Jin, Ruoxi Chen, Peiyan Zhang +4
The way customers search for and choose products is changing with the rise of large language models (LLMs). LLM-based search, or generative engines, provides direct product recomme…
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
Ye Yu, Haibo Jin, Yaoning Yu +2
Large audio-language models increasingly operate on raw speech inputs, enabling more seamless integration across domains such as voice assistants, education, and clinical triage. T…
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
Gang Cheng, Haibo Jin, Wenbin Zhang +2
Large Language Models (LLMs) are increasingly deployed in finance, where unsafe behavior can lead to serious regulatory risks. However, most red-teaming research focuses on overtly…
GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs
Haibo Jin, Ruoxi Chen, Peiyan Zhang +3
As Large Language Models (LLMs) become increasingly integral to various domains, their potential to generate harmful responses has prompted significant societal and regulatory conc…