8 papers
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Wei Dong, Tianyu Fu, Zhe Yu +9
As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion…
Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
Kaizhen Zhu, Mokai Pan, Zhechuan Yu +3
Diffusion Bridge and Flow Matching have both demonstrated compelling empirical performance in transformation between arbitrary distributions. However, there remains confusion about…
UniCA: Unified Covariate Adaptation for Time Series Foundation Model
Lu Han, Yu Liu, Lan Li +9
Time Series Foundation Models (TSFMs) have achieved remarkable success through large-scale pretraining. However, their design primarily targets real-valued series, limiting their a…
Comparative Separation: Evaluating Separation on Comparative Judgment Test Data
Xiaoyin Xi, Neeku Capak, Kate Stockwell +1
This research seeks to benefit the software engineering society by proposing comparative separation, a novel group fairness notion to evaluate the fairness of machine learning soft…
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
Yixin Tan, Zhe Yu, Jun Sakuma +1
Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications. However, its security implications remain unclear, parti…
FairReweighing: Density Estimation-Based Reweighing Framework for Improving Separation in Fair Regression
Xiaoyin Xi, Zhe Yu
There has been a prevalence of applying AI software in both high-stakes public-sector and industrial contexts. However, the lack of transparency has raised concerns about whether t…