5 papers
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
Garvin Guo, Donglei Yu, Yu Chen +6
Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interpretation can be premature: a t…
Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark
Yangyang Liu, Dong Yu, Pengyuan Liu
Large language models are increasingly used in value-sensitive decision settings, where irrelevant demographic cues should not alter judgments. We construct the Realistic Value Dec…
THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models
Zhiqing Ma, Zhonghao Xu, Dong Yu +3
Multi-turn jailbreak attacks pose a growing threat to LLMs by exploiting conversational dynamics such as gradual escalation and cross-turn coordination. Existing defenses either re…
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models
Binxian Su, Haoye Lou, Shucheng Zhu +4
Large language models (LLMs) are being increasingly used in urban planning, but since gendered space theory highlights how gender hierarchies are embedded in spatial organization,…
Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action
Qing Dong, Pengyuan Liu, Dong Yu +1
Generative agents have made significant progress in simulating human behavior, but existing frameworks often simplify emotional modeling and focus primarily on specific tasks, limi…