3 papers
cs.CL2026
Inducing Task Models from Computer-Use Traces
Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen +1
Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how…
cs.CL2026
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
Russell Yang, Ruishi Chen, Pierce Kelaita +6
Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise prefer…
cs.CL2026
Reflections and New Directions for Human-Centered Large Language Models
Caleb Ziems, Dora Zhao, Rose E. Wang +55
Large Language Models (LLMs) are increasingly shaping the private and professional lives of users, with numerous applications in business, education, finance, healthcare, law, and…