activity
20242026
most citedThe Bitter Lesson Learned from 2,000+ Multilingual Benchmarks

1 citations · 1 across the 27 of their papers we have counts for

collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI2026

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

Bo Zeng, Linfeng Gao, Peiqin Lin +9

Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere v…

cs.AI2026

CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning

Yongshi Ye, Tian Lan, Feihu Jiang +7

Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but…

cs.AI2026

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning

Yu Zhao, Zekun Zhang, Fan Jiang +6

Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm. Howeve…

cs.AI2026

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

Yuanyang Li, Xue Yang, Longyue Wang +2

Current LLM agents are proficient at calling isolated APIs but struggle with the "last mile" of commercial software automation. In real-world scenarios, tools are not independent;…

cs.AI2026

From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models

Ling Shi, Xinwei Wu, Xiaohu Zhao +7

While mechanistic interpretability tools like Sparse Autoencoders (SAEs) can uncover meaningful features within Large Language Models (LLMs), a critical gap remains in transforming…

cs.AI2026

Difficulty-Estimated Policy Optimization

Yu Zhao, Fan Jiang, Tianle Liu +4

Recent advancements in Large Reasoning Models (LRMs), exemplified by DeepSeek-R1, have underscored the potential of scaling inference-time compute through Group Relative Policy Opt…