From the 2 of 7 linked papers with an AI index.
7 papers
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
Fangxu Yu, Tao Feng, Dehai Min +6
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs…
Weak-to-Strong On-Policy Distillation
Fangxu Yu, Zinan Lin, Xiaodong Liu +4
The paper proposes Weak-to-Strong On-Policy Distillation (W2S-OPD), a method that improves a large language model by distilling knowledge from multiple weaker models using a constr…
Rushes: A Human Preference Dataset for Pluralistic Alignment
Michael Xu, Jorge Leandro, Sudha Rao +5
We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface…
Test-Time Learning with an Evolving Library
Weijia Xu, Alessandro Sordoni, Chandan Singh +4
The paper introduces EvoLib, a test-time learning framework that lets large language models build, reuse, and evolve a shared library of knowledge abstractions across tasks without…
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
Chandan Singh, Yan Shuo Tan, Weijia Xu +4
Agentic data science (ADS) systems are rapidly improving their capability to autonomously analyze, fit, and interpret data, potentially moving towards a future where agents conduct…
Echoes in AI: Quantifying lack of plot diversity in LLM outputs
Weijia Xu, Nebojsa Jojic, Sudha Rao +2
With rapid advances in large language models (LLMs), there has been an increasing application of LLMs in creative content ideation and generation. A critical question emerges: can…