3 papers
cs.CL2026
Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning
Yongqi Tong, Zhenyu Zhang, Zimi Liu +8
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this set…
cs.CL2026
STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment
Yongqi Tong, Zhenyu Zhang, Ruirui Wang +6
Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference d…
cs.AI2026
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus +5
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This fle…