8 papers
From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models
Jiaxin Zhang, Wendi Cui, Zhuohang Li +4
While Large Language Models (LLMs) show remarkable capabilities, their unreliability remains a critical barrier to deployment in high-stakes domains. This survey charts a functiona…
Judging with Confidence: Calibrating Autoraters to Preference Distributions
Zhuohang Li, Xiaowei Li, Chengyu Huang +11
The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limite…
A Survey of Automatic Prompt Optimization with Instruction-focused Heuristic-based Search Algorithm
Wendi Cui, Zhuohang Li, Hao Sun +5
Recent advances in Large Language Models have led to remarkable achievements across a variety of Natural Language Processing tasks, making prompt engineering increasingly central t…
SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
Wendi Cui, Zhuohang Li, Hao Sun +5
Designing optimal prompts for Large Language Models (LLMs) is a complicated and resource-intensive task, often requiring substantial human expertise and effort. Existing approaches…
What Really is a Member? Discrediting Membership Inference via Poisoning
Neal Mangaokar, Ashish Hooda, Zhuohang Li +5
Membership inference tests aim to determine whether a particular data point was included in a language model's training set. However, recent works have shown that such tests often…
SCE: Scalable Consistency Ensembles Make Blackbox Large Language Model Generation More Reliable
Jiaxin Zhang, Zhuohang Li, Wendi Cui +3
Large language models (LLMs) have demonstrated remarkable performance, yet their diverse strengths and weaknesses prevent any single LLM from achieving dominance across all tasks.…