10 papers
LLM-based Human Simulations Have Not Yet Been Reliable
Qian Wang, Jiaying Wu, Zichen Jiang +6
Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations rema…
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Qian Wang, Zhanzhi Lou, Zhenheng Tang +2
LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittl…
Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)
Zichen Tang, Zirui Zhang, Qian Wang +3
Current Large Language Models (LLMs) are gradually exploited in practically valuable agentic workflows such as Deep Research, E-commerce recommendation, and job recruitment. In the…
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
Zhenheng Tang, Xiang Liu, Qian Wang +3
As Large Language Models (LLMs) become more powerful and autonomous, they increasingly face conflicts and dilemmas in many scenarios. We first summarize and taxonomize these divers…
Towards Evaluting Fake Reasoning Bias in Language Models
Qian Wang, Zhenheng Tang, Zhanzhi Lou +3
Large Reasoning Models (LRMs), evolved from standard Large Language Models (LLMs), are increasingly utilized as automated judges because of their explicit reasoning processes. Yet…
MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs
Qian Wang, Tianyu Wang, Zhenheng Tang +4
LLM-based multi-agent systems (MAS) have shown promise in tackling complex tasks. However, existing solutions often suffer from limited agent coordination and heavy reliance on pre…