collaborators

10 papers

cs.AI2025

GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models

Yue Zhang, Jiaxin Zhang, Qiuyu Ren +5

We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathe…

cs.CL2025

A Survey of Automatic Prompt Optimization with Instruction-focused Heuristic-based Search Algorithm

Wendi Cui, Zhuohang Li, Hao Sun +5

Recent advances in Large Language Models have led to remarkable achievements across a variety of Natural Language Processing tasks, making prompt engineering increasingly central t…

cs.CL2025

SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization

Wendi Cui, Zhuohang Li, Hao Sun +5

Designing optimal prompts for Large Language Models (LLMs) is a complicated and resource-intensive task, often requiring substantial human expertise and effort. Existing approaches…

cond-mat.mtrl-sci2025

Language Models for Materials Discovery and Sustainability: Progress, Challenges, and Opportunities

Zongrui Pei, Junqi Yin, Jiaxin Zhang

Significant advancements have been made in one of the most critical branches of artificial intelligence: natural language processing (NLP). These advancements are exemplified by th…

cs.CL2025

SCE: Scalable Consistency Ensembles Make Blackbox Large Language Model Generation More Reliable

Jiaxin Zhang, Zhuohang Li, Wendi Cui +3

Large language models (LLMs) have demonstrated remarkable performance, yet their diverse strengths and weaknesses prevent any single LLM from achieving dominance across all tasks.…

cs.LG2025

Towards Statistical Factuality Guarantee for Large Vision-Language Models

Zhuohang Li, Chao Yan, Nicholas J. Jackson +4

Advancements in Large Vision-Language Models (LVLMs) have demonstrated promising performance in a variety of vision-language tasks involving image-conditioned free-form text genera…