activity
20242026
collaborators

36 papers

cs.CL2026

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Ajay Patel, Kartik Hosanagar, Ramayya Krishnan +3

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answe…

cs.CY2026

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

Bowen Jiang, Yuan Yuan, Zhuoqun Hao +11

Personal intelligence is becoming a central frontier for user-facing AI agents. To be helpful in everyday life, agents must understand users across the digital contexts where their…

cs.CL2026

FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale

Ajay Patel, Colin Raffel, Chris Callison-Burch

Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised "predict the next word" objective on a vast amount of unstruct…

cs.CL2026

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

Delip Rao, Chris Callison-Burch

Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels. Yet the same verdicts can support wildly varying agreement num…

cs.CL2026

NSF-SciFy: Mining the NSF Awards Database for Scientific Claims

Delip Rao, Weiqiu You, Eric Wong +1

We introduce NSF-SciFy, a comprehensive dataset of scientific claims and investigation proposals extracted from National Science Foundation award abstracts. While previous scientif…

cs.CL2026

When Verification Fails: How Compositionally Infeasible Claims Escape Rejection

Muxin Liu, Delip Rao, Grace Kim +1

Scientific claim verification, the task of determining whether claims are entailed by scientific evidence, is fundamental to establishing discoveries in evidence while preventing m…