activity
20232026
collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Thomson: Continual Learning of Frontier Models for SovereignAI

Shengzhuang Chen, Jerrod Parker, Yejin Bang +23

The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetr…

cs.AI2026

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

Samuel J. Vincent, Daniel Calloway, Fangyi Yu +2

Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially deter…

cs.AI2026

ContractScrub: A benchmark for final review of legal contracts

Yejin Bang, Kirsty Fielding, Brandan Oliver +3

Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final r…

cs.AI2026

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…

cs.AI2026

To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands

Fangyi Yu, Nabeel Seedat, Jonathan Richard Schwarz +1

Language models deployed in high-stakes professional settings face conflicting demands from users, institutional authorities, and professional norms. How models act when these dema…

cs.AI2025

Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings

Andrew M. Bean, Nabeel Seedat, Shengzhuang Chen +1

The prohibitive cost of evaluating large language models (LLMs) on comprehensive benchmarks necessitates the creation of small yet representative data subsets (i.e., tiny benchmark…