6 papers
MidSteer: Optimal Affine Framework for Steering Generative Models
Tatiana Gaintseva, Andrew Stepanov, Ziquan Liu +4
Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However,…
Confidence Should Be Calibrated More Than One Turn Deep
Zhaohan Zhang, Chengzhengxu Li, Xiaoming Liu +3
Large Language Models (LLMs) are increasingly applied in high-stakes domains such as finance, healthcare, and education, where reliable multi-turn interactions with users are essen…
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models
Zhaohan Zhang, Ziquan Liu, Ioannis Patras
Assessing the reliability of Large Language Models (LLMs) by confidence elicitation is a prominent approach to AI safety in high-stakes applications, such as healthcare and finance…
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework
Zhuo Zhi, Chen Feng, Adam Daneshmend +6
Multimodal large language models (MLLMs) show promise in tasks like visual question answering (VQA) but still face challenges in multimodal reasoning. Recent works adapt agentic fr…
Sebra: Debiasing Through Self-Guided Bias Ranking
Adarsh Kappiyath, Abhra Chaudhuri, Ajay Jaiswal +4
Ranking samples by fine-grained estimates of spuriosity (the degree to which spurious cues are present) has recently been shown to significantly benefit bias mitigation, over the t…
PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks
Chen Feng, Ziquan Liu, Zhuo Zhi +3
It is widely known that state-of-the-art machine learning models, including vision and language models, can be seriously compromised by adversarial perturbations. It is therefore i…