11 papers
WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts
Yuxin Meng, Yuhan Suo, Junjie Wang +9
Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions that determine whether a page…
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
Yinghao Wu, Zhuoyan Luo, Yiyao Yu +3
Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable performance degradation. Their emph…
Unified Data Selection for LLM Reasoning
Xiaoyuan Li, Yubo Ma, Chengpeng Li +6
Effectively training Large Language Models (LLMs) for complex, long-CoT reasoning is often bottlenecked by the need for massive high-quality reasoning data. Existing methods are ei…
When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering
Doeun Lee, Muge Zhang, Yi Yu +11
Across medical specialties, clinical practice is anchored in evidence-based guidelines that codify best studied diagnostic and treatment pathways. These pathways routinely fall sho…
Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation
Yuxuan Gao, Megan Wang, Yi Ling Yu
We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guarantees for forecasted quality…
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
Yuxuan Gao, Megan Wang, Yi Ling Yu +2
We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau-bench, BFCL multi-turn), a…