5 papers
CalBrief: A Pilot Diagnostic Benchmark for Evidence-Calibrated Scientific Briefing with Large Language Models
Yu Fu, Yongqi Kang, Yong Zhao
Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the sup…
X-MADAM-RAG: Diagnosing and Handling Chinese-English Evidence Conflict in Retrieval-Augmented Generation
Yongqi Kang, Yu Fu, Yong Zhao
Retrieval-augmented generation (RAG) systems may receive evidence that is not merely noisy but mutually contradictory. This issue becomes particularly salient in multilingual setti…
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
Tengda Guo, Jie Leng, Hanlei Li +6
Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual information to support semantic…
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement
Yu Fu, Yongqi Kang, Yujia Zhou +1
Authoring and refining presentation slides is a highly time-consuming core task in academic and business domains. While generative AI tools have lowered the barrier for creating in…
From "What to Eat?" to Perfect Recipe: ChefMind's Chain-of-Exploration for Ambiguous User Intent in Recipe Recommendation
Yu Fu, Linyue Cai, Ruoyu Wu +1
Personalized recipe recommendation faces challenges in handling fuzzy user intent, ensuring semantic accuracy, and providing sufficient detail coverage. We propose ChefMind, a hybr…