1 citations · 1 across the 1 of their papers we have counts for
5 papers
Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
Yiming Du, Wenyu Huang, Danna Zheng +5
Memory is fundamental to large language model (LLM)-based agents, but existing surveys emphasize application-level use (e.g., personalized dialogue), while overlooking the atomic o…
Long-Form Information Alignment Evaluation Beyond Atomic Facts
Danna Zheng, Mirella Lapata, Jeff Z. Pan
Information alignment evaluators are vital for various NLG evaluation tasks and trustworthy LLM deployment, reducing hallucinations and enhancing user trust. Current fine-grained m…
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
Danna Zheng, Mirella Lapata, Jeff Z. Pan
Large Language Models (LLMs) are increasingly explored as knowledge bases (KBs), yet current evaluation methods focus too narrowly on knowledge retention, overlooking other crucial…
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
Danna Zheng, Danyang Liu, Mirella Lapata +1
Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, prompting a surge in their practical applications. However, concerns have arisen rega…
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
Danna Zheng, Mirella Lapata, Jeff Z. Pan
We present Archer, a challenging bilingual text-to-SQL dataset specific to complex reasoning, including arithmetic, commonsense and hypothetical reasoning. It contains 1,042 Englis…