activity
20242026
collaborators
Showing cs.CLShow all

17 papers · 1 filter

cs.CL2025

AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy

Sebastian Antony Joseph, Syed Murtaza Husain, Stella S. R. Offner +7

Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate…

cs.CL2025

EvalAgent: Discovering Implicit Evaluation Criteria from the Web

Manya Wadhwa, Zayne Sprague, Chaitanya Malaviya +3

Evaluation of language model outputs on structured writing tasks is typically conducted with a number of desirable criteria presented to human evaluators or large language models (…

cs.CL2025

QUDsim: Quantifying Discourse Similarities in LLM-Generated Text

Ramya Namuduri, Yating Wu, Anshun Asher Zheng +3

As large language models become increasingly capable at various writing tasks, their weakness at generating unique and creative content becomes a major liability. Although LLMs hav…

cs.CL2025

Learning to Refine with Fine-Grained Natural Language Feedback

Manya Wadhwa, Xinyu Zhao, Junyi Jessy Li +1

Recent work has explored the capability of large language models (LLMs) to identify and correct errors in LLM-generated responses. These refinement approaches frequently evaluate w…

cs.CL2025

Using Natural Language Explanations to Rescale Human Judgments

Manya Wadhwa, Jifan Chen, Junyi Jessy Li +1

The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation. A common pr…

cs.CL2025

From Distributional to Overton Pluralism: Investigating Large Language Model Alignment

Thom Lake, Eunsol Choi, Greg Durrett

The alignment process changes several properties of a large language model's (LLM's) output distribution. We analyze two aspects of post-alignment distributional shift of LLM respo…