activity
20242026
most citedDon't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search

1 citations · 1 across the 47 of their papers we have counts for

collaborators

136 papers

cs.AI2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Changjiang Jiang, Qiannian Zhao, Lei Xin +3

Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this reliance introduces significant inf…

cs.CL2026

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah +11

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally…

cs.AI2026

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

Yuyang Dai, Xueqing Peng, Yuxia Wang +2

Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings.…

cs.CL2026

One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

Minh Ngoc Ta, My Anh Tran Nguyen, Duong D. Nguyen +2

Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to…

cs.CL2026

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

Zhuohan Xie, Xueqing Peng, Georgi Georgiev +18

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and n…

cs.CL2026

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

Zhuohan Xie, Yuyang Dai, Rania Elbadry +18

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the corr…