3 papers
cs.CL2026
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
Gaurav Srivastava, Aafiya Hussain, Zhenyu Bi +5
Evaluating language models fairly is increasingly difficult as static benchmarks risk contamination by training data, obscuring whether models truly reason or recall. We introduce…
cs.LG2026
Agentic Framework for Epidemiological Modeling
Rituparna Datta, Zihan Guan, Baltazar Espinoza +5
Epidemic modeling is essential for public health planning, yet traditional approaches rely on fixed model classes that require manual redesign as pathogens, policies, and scenario…
cs.HC2024
ArguMentor: Augmenting User Experiences with Counter-Perspectives
Priya Pitre, Kurt Luther
We encounter arguments everyday in the form of social media posts, presidential debates, news articles, and even advertisements. A ubiquitous, influential example is the opinion pi…