7 papers
Best-of-Better-: Generating Pre-Aligned Responses with In-Context Learning
Eric Lei, Hsiang Hsu, Chun-Fu Chen
Inference-time alignment methods, such as Best-of-, offer a flexible alternative to training-based alignment by using reward models to select high-quality responses generated by…
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
Arjun Nichani, Hsiang Hsu, Chun-Fu +2
Fairness and privacy are two vital pillars of trustworthy machine learning. Despite extensive research on these individual topics, their relationship has received significantly les…
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
Hsiang Hsu, Eric Lei, Chun-Fu Chen
Inference-time alignment effectively steers large language models (LLMs) by generating multiple candidates from a reference model and selecting among them with an imperfect reward…
The Unseen Threat: Residual Knowledge in Machine Unlearning under Perturbed Samples
Hsiang Hsu, Pradeep Niroula, Zichang He +3
Machine unlearning offers a practical alternative to avoid full model re-training by approximately removing the influence of specific user data. While existing methods certify unle…
Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
Seongmin Lee, Hsiang Hsu, Chun-Fu Chen +1
LLM hallucination, where unfaithful text is generated, presents a critical challenge for LLMs' practical applications. Current detection methods often resort to external knowledge,…
PASS: Private Attributes Protection with Stochastic Data Substitution
Yizhuo Chen, Chun-Fu, Chen +3
The growing Machine Learning (ML) services require extensive collections of user data, which may inadvertently include people's private information irrelevant to the services. Vari…