6 papers · 1 filter
Thomson: Continual Learning of Frontier Models for SovereignAI
Shengzhuang Chen, Jerrod Parker, Yejin Bang +23
The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetr…
InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries
Samuel J. Vincent, Daniel Calloway, Fangyi Yu +2
Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially deter…
ContractScrub: A benchmark for final review of legal contracts
Yejin Bang, Kirsty Fielding, Brandan Oliver +3
Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final r…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands
Fangyi Yu, Nabeel Seedat, Jonathan Richard Schwarz +1
Language models deployed in high-stakes professional settings face conflicting demands from users, institutional authorities, and professional norms. How models act when these dema…
Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
Andrew M. Bean, Nabeel Seedat, Shengzhuang Chen +1
The prohibitive cost of evaluating large language models (LLMs) on comprehensive benchmarks necessitates the creation of small yet representative data subsets (i.e., tiny benchmark…