13 papers
Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
Ying He, Zhouhong Gu, Zhecheng Hu +8
Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language…
The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models
Ying He, Sihang Jiang, Xingzhou Chen +6
Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural b…
Deep Research as Rubric for Reinforcement Learning
Wangyi Mei, Zhouhong Gu, Zhenhan Bai +9
Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a promising alternative, but ex…
Scaling Behavior of Single LLM-Driven Multi-Agent Systems
Jialing Li, Zhouhong Gu, Yin Cai +1
The burgeoning field of LLM-based Multi-Agent Systems (MAS) promises to tackle complex tasks through collaborative intelligence, yet fundamental questions regarding their scaling b…
CompBench: Benchmarking Complex Instruction-guided Image Editing
Bohan Jia, Wenxuan Huang, Yuntian Tang +14
While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack com…
MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments
Yin Cai, Zhouhong Gu, Zhaohan Du +5
Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly i…