4 papers
Beyond State Consistency: Behavior Consistency in Text-Based World Models
Youling Huang, Guanqiao Chen, Junchi Yao +8
World models have been emerging as critical components for assessing the consequences of actions generated by interactive agents in online planning and offline evaluation. In text-…
Triviality Corrected Endogenous Reward
Xinda Wang, Zhengxu Hou, Yangshijie Zhang +6
Reinforcement learning for open-ended text generation is constrained by the lack of verifiable rewards, necessitating reliance on judge models that require either annotated data or…
PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models
Chenzhuo Zhao, Ziqian Liu, Xinda Wang +2
Prompt optimization is a practical and widely applicable alternative to fine tuning for improving large language model performance. Yet many existing methods evaluate candidate pro…
TASE: Token Awareness and Structured Evaluation for Multilingual Language Models
Chenzhuo Zhao, Xinda Wang, Yue Huang +2
While large language models (LLMs) have demonstrated remarkable performance on high-level semantic tasks, they often struggle with fine-grained, token-level understanding and struc…