5 papers
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
Jundong Xu, Qingchuan Li, Jiaying Wu +11
Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deploymen…
From Perception to Action: An Interactive Benchmark for Vision Reasoning
Yuhao Wu, Maojia Song, Yihuai Lan +8
Understanding the physical structure is essential for real-world applications such as embodied agents, interactive design, and long-horizon manipulation. Yet, prevailing Vision-Lan…
FedReFT: Federated Representation Fine-Tuning with All-But-Me Aggregation
Fatema Siddika, Md Anwar Hossen, J. Pablo Muñoz +3
Parameter-efficient fine-tuning (PEFT) adapts large pre-trained models by updating only a small subset of parameters. Recently, Representation Fine-Tuning (ReFT) has emerged as an…
An Empirical Study on Prompt Compression for Large Language Models
Zheng Zhang, Jinyi Li, Yihuai Lan +2
Prompt engineering enables Large Language Models (LLMs) to perform a variety of tasks. However, lengthy prompts significantly increase computational complexity and economic costs.…
DVM: Towards Controllable LLM Agents in Social Deduction Games
Zheng Zhang, Yihuai Lan, Yangsen Chen +3
Large Language Models (LLMs) have advanced the capability of game agents in social deduction games (SDGs). These games rely heavily on conversation-driven interactions and require…