6 papers
Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning
Chenghao Zhang, Canran Xiao, SaiSai Hu +1
Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible while hiding broken links, inconsistent s…
Prototype-Aligned Federated Soft-Prompts for Continual Web Personalization
Canran Xiao, Liwei Hou
Continual web personalization is essential for engagement, yet real-world non-stationarity and privacy constraints make it hard to adapt quickly without forgetting long-term prefer…
CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated Rewards
Zhiming Lin, Kai Zhao, Sophie Zhang +2
Large-scale Chinese spelling correction (CSC) remains critical for real-world text processing, yet existing LLMs and supervised methods lack robustness to novel errors and rely on…
From Points to Coalitions: Hierarchical Contrastive Shapley Values for Prioritizing Data Samples
Canran Xiao, Jiabao Dou, Zhiming Lin +2
How should we quantify the value of each training example when datasets are large, heterogeneous, and geometrically structured? Classical Data-Shapley answers in principle, but its…
Beyond More Context: Retrieval Diversity Boosts Multi-Turn Intent Understanding
Zhiming Lin
Multi turn intent understanding is central to task oriented chatbots, yet real deployments face tight token budgets and noisy contexts, and most retrieval pipelines emphasize relev…
ReviBranch: Deep Reinforcement Learning for Branch-and-Bound with Revived Trajectories
Dou Jiabao, Nie Jiayi, Yihang Cheng +5
The Branch-and-bound (B&B) algorithm is the main solver for Mixed Integer Linear Programs (MILPs), where the selection of branching variable is essential to computational efficienc…