5 papers
EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction
Chengxuan Qin, Zhige Chen, Shu Peng +9
Electroencephalography (EEG) foundation models increasingly rely on multi-dataset training and evaluation, yet public EEG datasets still lack a shared task specification layer that…
SCOPE: Sequential Conformal Probing for Reliable OOD Rejection in LLM Services
Zhuoyun Li, Boxuan Wang, Changshun Wu +2
Rejecting inputs outside the defined in-distribution (IND) service scope is critical for large language model (LLM) services, where unsupported requests should be filtered before f…
ProCUA-SFT Technical Report
Jaehun Jung, Ximing Lu, Brandon Cui +11
Training computer-use agents (CUAs) -- models that interact with graphical desktops through screenshots and keyboard/mouse actions -- requires large-scale, diverse trajectory data…
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
Yuxuan Wan, Chaozheng Wang, Yi Dong +4
Websites are critical in today's digital world, with over 1.11 billion currently active and approximately 252,000 new sites launched daily. Converting website layout design into fu…
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
Yuxuan Wan, Yi Dong, Jingyu Xiao +3
Multi-page websites dominate modern web development. However, existing design-to-code methods rely on simplified assumptions, limiting to single-page, self-contained webpages witho…