Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
Luyu Chen, Zeyu Zhang, Haoran Tan +4
LLMs have emerged as powerful evaluators in the LLM-as-a-Judge paradigm, offering significant efficiency and flexibility compared to human judgments. However, previous methods prim…
cs.AI2025
MemEngine: A Unified and Modular Library for Developing Advanced Memory of LLM-based Agents
Zeyu Zhang, Quanyu Dai, Xu Chen +3
Recently, large language model based (LLM-based) agents have been widely applied across various fields. As a critical part, their memory capabilities have captured significant inte…
cs.AI2024
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
Zeyu Zhang, Quanyu Dai, Luyu Chen +7
LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lack…