2 papers
cs.LG2025
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
Junkai Zhang, Zihao Wang, Lin Gui +7
Reinforcement fine-tuning (RFT) often suffers from reward over-optimization, where a policy model hacks the reward signals to achieve high scores while producing low-quality output…
cs.AI2024
Causal Understanding For Video Question Answering
Bhanu Prakash Reddy Guda, Tanmay Kulkarni, Adithya Sampath +1
Video Question Answering is a challenging task, which requires the model to reason over multiple frames and understand the interaction between different objects to answer questions…