activity
20242026
most citedDon't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search

1 citations · 1 across the 62 of their papers we have counts for

collaborators
Showing cs.AIShow all

15 papers · 1 filter

cs.AI2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Changjiang Jiang, Qiannian Zhao, Lei Xin +3

Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this reliance introduces significant inf…

cs.AI2026

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

Yuyang Dai, Xueqing Peng, Yuxia Wang +2

Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings.…

cs.AI2026

Can Agentic Trading Systems Pay for Their Own Intelligence?

Qiqi Duan, Changlun Li, Chen Wang +10

Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce tradin…

cs.AI2026

LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents

Jingpu Yang, Fengxian Ji, Zhengzhao Lai +8

Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challeng…

cs.AI2026

The FIL Hypothesis: Inductive Biases Help with Kernel Engineering

Nikolai Rozanov, Subhabrata Dutta, Preslav Nakov +1

The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in human knowledge, has become a dominant p…

cs.AI2026

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

Zhihao Zhang, Liting Huang, Guanghao Wu +3

Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal of benign queries or unsafe compl…