most citedReinforcement Learning with Rubric Anchors

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2026

Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation

Jia Li, Xiaomeng Fu, Xurui Peng +7

Autoregressive video diffusion models have emerged as a scalable paradigm for long video generation. However, they often suffer from severe extrapolation failure, where rapid error…

cs.CV2026

TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration

Haowei Zhu, Tingxuan Huang, Xing Wang +7

Diffusion models achieve strong generative performance but remain slow at inference due to the need for repeated full-model denoising passes. We present Token-Adaptive Predictor (T…

cs.CL2026

Fast-weight Product Key Memory

Tianyu Zhao, Llion Jones

Sequence modeling layers in modern language models typically face a trade-off between storage capacity and computational efficiency. While softmax attention offers unbounded storag…

cs.CE2025

Task-Specific Sparse Feature Masks for Molecular Toxicity Prediction with Chemical Language Models

Kwun Sy Lee, Jiawei Chen, Fuk Sheng Ford Chung +3

Reliable in silico molecular toxicity prediction is a cornerstone of modern drug discovery, offering a scalable alternative to experimental screening. However, the black-box nature…

cs.AI20251 cited

Reinforcement Learning with Rubric Anchors

Zenan Huang, Yihong Zhuang, Guoshan Lu +18

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing Large Language Models (LLMs), exemplified by the success of OpenAI's o-series…

cs.CL2025

TransEvalnia: Reasoning-based Evaluation and Ranking of Translations

Richard Sproat, Tianyu Zhao, Llion Jones

We present TransEvalnia, a prompting-based translation evaluation and ranking system that uses reasoning in performing its evaluations and ranking. This system presents fine-graine…