collaborators

7 papers

cs.CL2026

PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering

Xiangfeng Wang, Hangyu Guo, Yanlin Lai +11

While model-based verifiers are essential for scaling Reinforcement Learning with Verifiable Rewards (RLVR), current outcome-centric verification paradigms primarily focus on the c…

cs.CL2026

R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging

Yanlin Lai, Mitt Huang, Hangyu Guo +11

Reinforcement Learning from Human Feedback (RLHF) remains indispensable for aligning large language models (LLMs) in subjective domains. To enhance robustness, recent work shifts t…

cs.CL2026

ECR: Manifold-Guided Semantic Cues for Compact Language Models

Chung-Wei Victor Yuan

Compact models often lose the structure of their embedding space. The issue shows up when the capacity is tight or the data spans several languages. Such collapse makes it difficul…

cs.CV2025

Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search

Yunqi Zhou, Chengjie Jiang, Chun Yuan +1

With advances in satellite constellations, sensor technologies, and imaging pipelines, ultra-high-resolution (Ultra-HR) remote sensing imagery is becoming increasingly widespread.…

cs.AI2025

UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios

Haotian Luo, Huaisong Zhang, Xuelin Zhang +15

Autonomous agents have recently achieved remarkable progress across diverse domains, yet most evaluations focus on short-horizon, fully observable tasks. In contrast, many critical…

cs.LG2025

Preference Optimization for Combinatorial Optimization Problems

Mingjun Pan, Guanquan Lin, You-Wei Luo +4

Reinforcement Learning (RL) has emerged as a powerful tool for neural combinatorial optimization, enabling models to learn heuristics that solve complex problems without requiring…