4 papers
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
Yitao Long, Yuru Jiang, Hongjun Liu +6
This work investigates the reasoning and planning capabilities of foundation models and their scalability in complex, dynamic environments. We introduce PuzzlePlex, a benchmark des…
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
Yitao Long, Tiansheng Hu, Yilun Zhao +2
Large Language Models (LLMs) frequently hallucinate to long-form questions, producing plausible yet factually incorrect answers. A common mitigation strategy is to provide attribut…
MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning
Hongjun Liu, Yinghao Zhu, Yuhui Wang +4
Recent progress in multimodal large language models (MLLMs) has demonstrated promising performance on medical benchmarks and in preliminary trials as clinical assistants. Yet, our…
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
Yilun Zhao, Yitao Long, Yuru Jiang +7
We introduce FinDVer, a comprehensive benchmark specifically designed to evaluate the explainable claim verification capabilities of LLMs in the context of understanding and analyz…