Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Chengyu Shen, Yujie Fu, Gangtao Xin +13
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…
cs.CL2025
Entropy-Guided Reasoning Compression
Hourun Zhu, Yang Gao, Wenlong Fei +2
Large reasoning models have demonstrated remarkable performance on complex reasoning tasks, yet the excessive length of their chain-of-thought outputs remains a major practical bot…