activity
20232026
collaborators

6 papers

cs.AI2026

RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments

Linghua Zhang, Jun Wang, Jingtong Wu +1

Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments…

cs.CL2026

Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation

Shuaiyi Li, Zhisong Zhang, Yan Wang +5

Block attention, which processes the input as separate blocks that cannot attend to one another, offers significant potential to improve KV cache reuse in long-context scenarios su…

cs.CL2026

A Decomposition Perspective to Long-context Reasoning for LLMs

Yanling Xiao, Huaibing Xie, Guoliang Zhao +8

Long-context reasoning is essential for complex real-world applications, yet remains a significant challenge for Large Language Models (LLMs). Despite the rapid evolution in long-c…

cs.AI2026

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments

Linghua Zhang, Jun Wang, Jingtong Wu +1

Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments…

cs.CL2024

On the Transformations across Reward Model, Parameter Update, and In-Context Prompt

Deng Cai, Huayang Li, Tingchen Fu +11

Despite the general capabilities of pre-trained large language models (LLMs), they still need further adaptation to better serve practical applications. In this paper, we demonstra…

cs.CL2024

On the Worst Prompt Performance of Large Language Models

Bowen Cao, Deng Cai, Zhisong Zhang +2

The performance of large language models (LLMs) is acutely sensitive to the phrasing of prompts, which raises significant concerns about their reliability in real-world scenarios.…