activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Judging with Confidence: Calibrating Autoraters to Preference Distributions

Zhuohang Li, Xiaowei Li, Chengyu Huang +11

The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limite…

cs.CL2025

A Survey of Automatic Prompt Optimization with Instruction-focused Heuristic-based Search Algorithm

Wendi Cui, Zhuohang Li, Hao Sun +5

Recent advances in Large Language Models have led to remarkable achievements across a variety of Natural Language Processing tasks, making prompt engineering increasingly central t…

cs.CL2025

SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization

Wendi Cui, Zhuohang Li, Hao Sun +5

Designing optimal prompts for Large Language Models (LLMs) is a complicated and resource-intensive task, often requiring substantial human expertise and effort. Existing approaches…

cs.CL2025

SCE: Scalable Consistency Ensembles Make Blackbox Large Language Model Generation More Reliable

Jiaxin Zhang, Zhuohang Li, Wendi Cui +3

Large language models (LLMs) have demonstrated remarkable performance, yet their diverse strengths and weaknesses prevent any single LLM from achieving dominance across all tasks.…

cs.CL2024

Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented Generation

Zhuohang Li, Jiaxin Zhang, Chao Yan +4

Language models (LMs) are known to suffer from hallucinations and misinformation. Retrieval augmented generation (RAG) that retrieves verifiable information from an external knowle…