4 papers
Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents
Liming Pu, Xiaoxia Li, Yifu Liu +2
Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-…
Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards
Xuefeng Jin, Jiashuo Zhang, Teng Cao +1
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt e…
Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes
Hikaru Shindo, Yu Deng, Teng Cao +5
Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed. This makes brittle behavior d…
SuperPose: Improved 6D Pose Estimation with Robust Tracking and Mask-Free Initialization
Yu Deng, Jiahong Xue, Teng Cao +3
We developed a robust solution for real-time 6D object detection in industrial applications by integrating FoundationPose, SAM2, and LightGlue, eliminating the need for retraining.…