From the 1 of 11 linked papers with an AI index.
11 papers
Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
Kevin Miller, Arjun Chandra, Venkatesh Saligrama
Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduc…
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
Rahul Gupta, Abhinav Mohanty, Payal Motwani +8
The paper introduces a Threshold Exceedance Criteria (TEC) framework to systematically evaluate whether frontier language models increase a non‑expert's ability to plan chemical, b…
PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges
Swastik Roy, Rajkumar Pujari, Tharindu Kumarage +5
LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be `…
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
Narmeen Oozeer, Luke Marks, Shreyans Jain +2
Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of…
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
Yukun Huang, Leonardo F. R. Ribeiro, Momchil Hardalov +3
Search-augmented LLM agents can produce deep research reports (DRRs), but verifying claim-level factuality remains challenging. Existing fact-checkers are primarily designed for ge…
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
Shengao Wang, Wenqi Wang, Zecheng Wang +20
Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded…