10 citations · 10 across the 1 of their papers we have counts for
5 papers
LogicGraph : Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification
Yanrui Wu, Lingling Zhang, Xinyu Zhang +5
Evaluations of large language models (LLMs) primarily emphasize convergent logical reasoning, where success is defined by producing a single correct proof. However, many real-world…
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
Yifei Li, Weidong Guo, Lingling Zhang +6
Long-term conversational memory is a core capability for LLM-based dialogue systems, yet existing benchmarks and evaluation protocols primarily focus on surface-level factual recal…
From Detection to Diagnosis: Advancing Hallucination Analysis with Automated Data Synthesis
Yanyi Liu, Qingwen Yang, Tiezheng Guo +3
Hallucinations in Large Language Models (LLMs), defined as the generation of content inconsistent with facts or context, represent a core obstacle to their reliable deployment in c…
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
Shilin Sun, Wenbin An, Feng Tian +5
Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challe…
Scaffolded Language Models with Language Supervision for Mixed-Autonomy: A Survey
Matthieu Lin, Jenny Sheng, Andrew Zhao +7
This survey organizes the intricate literature on the design and optimization of emerging structures around post-trained LMs. We refer to this overarching structure as scaffolded L…