5 papers
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
Joshua Ong Jun Leang, Zheng Zhao, Aryo Pradipta Gema +7
The paper proposes PiCSAR, a training-free scoring method that uses the joint log-likelihood of reasoning steps and final answer to select the most reliable reasoning chain from mu…
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
GRADA: Graph-based Reranking against Adversarial Documents Attack
Jingjie Zheng, Aryo Pradipta Gema, Giwon Hong +4
Retrieval Augmented Generation (RAG) frameworks improve the accuracy of large language models (LLMs) by integrating external knowledge from retrieved documents, thereby overcoming…
An Auditing Test To Detect Behavioral Shift in Language Models
Leo Richter, Xuanli He, Pasquale Minervini +1
As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task perf…
Self-Training Large Language Models for Tool-Use Without Demonstrations
Ne Luo, Aryo Pradipta Gema, Xuanli He +3
Large language models (LLMs) remain prone to factual inaccuracies and computational errors, including hallucinations and mistakes in mathematical reasoning. Recent work augmented L…