1 citations · 1 across the 5 of their papers we have counts for
7 papers
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani +3
Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent work reports strong tool-calling perform…
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
Yihan Hong, Huaiyuan Yao, Bolin Shen +3
Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challengin…
LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning
Xudong Chen, Yixin Liu, Hua Wei +1
Large language models (LLMs) have become a strong foundation for multi-agent systems, but their effectiveness depends heavily on orchestration design. Across different tasks, role…
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Xiaolin Zhou, Aojie Yuan, Zheng Luo +12
Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user…
Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition
Tiejin Chen, Huaiyuan Yao, Jia Chen +2
While Large Language Model-based Multi-Agent Systems (MAS) consistently outperform single-agent systems on complex tasks, their intricate interactions introduce critical reliabilit…
AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
Chen Liang, Zhaoqi Huang, Haofen Wang +8
Large language models (LLMs), as a novel information technology, are seeing increasing adoption in the Architecture, Engineering, and Construction (AEC) field. They have shown thei…