collaborators

11 papers

cs.CL2026

WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts

Yuxin Meng, Yuhan Suo, Junjie Wang +9

Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions that determine whether a page…

cs.CV2026

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding

Yinghao Wu, Zhuoyan Luo, Yiyao Yu +3

Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable performance degradation. Their emph…

cs.CL2026

Unified Data Selection for LLM Reasoning

Xiaoyuan Li, Yubo Ma, Chengpeng Li +6

Effectively training Large Language Models (LLMs) for complex, long-CoT reasoning is often bottlenecked by the need for massive high-quality reasoning data. Existing methods are ei…

cs.CL2026

When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering

Doeun Lee, Muge Zhang, Yi Yu +11

Across medical specialties, clinical practice is anchored in evidence-based guidelines that codify best studied diagnostic and treatment pathways. These pathways routinely fall sho…

cs.AI2026

Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation

Yuxuan Gao, Megan Wang, Yi Ling Yu

We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guarantees for forecasted quality…

cs.AI2026

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

Yuxuan Gao, Megan Wang, Yi Ling Yu +2

We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau-bench, BFCL multi-turn), a…