Publications (28)
LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept Discovery
Difei Gu, Yunhe Gao, Gerasimos Chatzoudis +6
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
Can Jin, Yang Zhou, Qixin Zhang +8
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Jiaqi Liu, Shi Qiu, Mairui Li +33
Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces
Zihan Dong, Xinyu Fan, Zixiang Tang +1
RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
Haoxiang Jiang, Zihan Dong, Tianci Liu +5
RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
Pei Tian, Zihan Dong, Tianci Liu +2
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
Peng Xia, Jinglu Wang, Yibo Peng +10
ManuRAG: Multi-modal Retrieval Augmented Generation for Manufacturing Question Answering (Early Version)
Yunqing Li, Zihan Dong, Farhad Ameri +1
Students' Perceptions and Preferences of Generative Artificial Intelligence Feedback for Programming
Zhengdong Zhang, Zihan Dong, Yang Shi +3
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
Zihan Dong, Xiaotian Hou, Ruijia Wu +1
Enhancing Bloodstain Analysis Through AI-Based Segmentation: Leveraging Segment Anything Model for Crime Scene Investigation
Zihan Dong, ZhengDong Zhang
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
Zihan Dong, Zhixian Zhang, Yang Zhou +3
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Tianci Liu, Zihan Dong, Linjun Zhang +5
Self-Consolidating Language Models: Continual Knowledge Incorporation from Context
Zekun Wang, Anant Gupta, Zihan Dong +1
How Benchmarks Mis-Score Computer-Use Agents
Zihan Dong, Zhiyuan Ma, Zekun Wang +5
The paper examines how current benchmarks for computer-use agents often give inaccurate scores due to issues in task design, trajectory observation, scoring, and reporting, and pro…
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4
Mitigating Heterogeneous Token Overfitting in LLM Knowledge Editing
Tianci Liu, Ruirui Li, Zihan Dong +6
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
Zihan Dong, Rui Qian, Qishi Zhan +3
The paper introduces Adaptive Anticipatory Policy Trees (AAPT), a method that pre‑computes conditional action trees during idle screen time so GUI agents can react instantly to eve…
All-fiber highly efficient delivery of 2 kW laser over 2.45 km hollow-core fiber
Jing Shi, Binyu Rao, Zilun Chen +16
Evidence Over Plans: Online Trajectory Verification for Skill Distillation
Yang Zhou, Zihan Dong, Zhenting Wang +7
Avoid Catastrophic Forgetting with Rank-1 Fisher from Diffusion Models
Zekun Wang, Anant Gupta, Zihan Dong +1
FNSPID: A Comprehensive Financial News Dataset in Time Series
Zihan Dong, Xinyu Fan, Zhiyuan Peng
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
Ran Xu, Tianci Liu, Zihan Dong +6
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
Shuhang Lin, Chuhao Zhou, Xiao Lin +5
AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
Ran Xu, Yuchen Zhuang, Zihan Dong +7
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
Yang Zhou, Can Jin, Zihan Dong +7
Decompose Sparsely Where You Should, Absorb Densely Where You Should No
Ruixuan Deng, Zehao Jin, Zekun Wang +1
Contrastive Network Representation Learning
Zihan Dong, Xin Zhou, Ryumei Nakada +2