activity
20242026
most citedDon't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search

1 citations · 1 across the 63 of their papers we have counts for

collaborators
Showing cs.CLShow all

112 papers · 1 filter

cs.CL2026

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah +11

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally…

cs.CL2026

One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

Minh Ngoc Ta, My Anh Tran Nguyen, Duong D. Nguyen +2

Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to…

cs.CL2026

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

Zhuohan Xie, Xueqing Peng, Georgi Georgiev +18

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and n…

cs.CL2026

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

Zhuohan Xie, Yuyang Dai, Rania Elbadry +18

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the corr…

cs.CL2026

Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG

Saadeldine Eletter, Owais Aijaz, Preslav Nakov

Multimodal retrieval-augmented generation (RAG) is often evaluated with clean evidence, yet real retrieval can return topically relevant but unreliable content: false text and misl…

cs.CL2026

MIRAGE: Defending Long-Form RAG Against Misinformation Pollution

Saadeldine Eletter, Ruihong Zeng, Yuxia Wang +3

Retrieval-Augmented Generation (RAG) improves factuality by grounding LLMs in external evidence, but real-world retrieval is often polluted: semantically relevant passages may cont…