7 papers
ChartAct: A Benchmark for Dynamic Chart Understanding
Muye Huang, Lin Wu, Lingling Zhang +5
Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts, but real-world charts are of…
DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints
Zhitong Chen, Kai Yin, Weifeng Zhang +7
Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage assessment, into coherent mu…
MiRD: Reliable Set-Valued Prediction for Open-Ended Question Answering via Miscoverage Risk Decomposition
Anqi Hu, Zhiyuan Wang, Zijun Jia +1
Reliable set-valued prediction provides a principled way to mitigate hallucinations in open-ended question answering (QA), yet existing conformal approaches typically rely on a fra…
Towards Efficient and Robust Linguistic Emotion Diagnosis for Mental Health via Multi-Agent Instruction Refinement
Jian Zhang, Zhangqi Wang, Zhiyuan Wang +5
Linguistic expressions of emotions such as depression, anxiety, and trauma-related states are pervasive in clinical notes, counseling dialogues, and online mental health communitie…
-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation
Jian Zhang, Yu He, Zhiyuan Wang +5
Scientific reasoning relies not only on logical inference but also on activating prior knowledge and experiential structures. Memory can efficiently reuse knowledge and enhance rea…
MAXS: Meta-Adaptive Exploration with LLM Agents
Jian Zhang, Zhiyuan Wang, Zhangqi Wang +7
Large Language Model (LLM) Agents exhibit inherent reasoning abilities through the collaboration of multiple tools. However, during agent inference, existing methods often suffer f…