4 papers
FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
Xianfu Cheng, Shiwei Zhang, Jiyu Zhao +10
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. Howeve…
M: Every Task Deserves Its Own Memory Harness
Wenbo Pan, Shujie Liu, Xiangyang Zhou +4
Large language model agents rely on specialized memory systems to accumulate and reuse knowledge during extended interactions. Recent architectures typically adopt a fixed memory d…
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
Xianfu Cheng, Wei Zhang, Shiwei Zhang +16
The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly t…
Unleashing Potential of Evidence in Knowledge-Intensive Dialogue Generation
Xianjie Wu, Jian Yang, Tongliang Li +4
Incorporating external knowledge into dialogue generation (KIDG) is crucial for improving the correctness of response, where evidence fragments serve as knowledgeable snippets supp…