5 papers
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Anka Reuel, Avijit Ghosh, Jenny Chim +32
Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risks and capabilities. Although general c…
The Trust Paradox: How CS Researchers Engage LLM Leaderboards
Pouya Sadeghi, Anamaria Crisan, Jimmy Lin
Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reli…
PinPoint: Prompting with Informative Interior Points
Pouya Sadeghi, Shawn He, Pedro Pablo Guerrero Vela +3
Modern referring image segmentation pipelines couple a vision-language model (VLM) for grounding with a promptable segmenter such as the Segment Anything Model (SAM) for mask gener…
Supermartingale Certificates for Quantitative Omega-regular Verification and Control
Thomas A. Henzinger, Kaushik Mallik, Pouya Sadeghi +1
We present the first supermartingale certificate for quantitative -regular properties of discrete-time infinite-state stochastic systems. Our certificate is defined on the prod…
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation
Samin Mahdizadeh Sani, Pouya Sadeghi, Thuy-Trang Vu +2
Large language models (LLMs) have made great progress in classification and text generation tasks. However, they are mainly trained on English data and often struggle with low-reso…