6 papers
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
Shujin Wu, Cheng Qian, Xiusi Chen +1
Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of…
MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration
Zhitao He, Sandeep Polisetty, Zhiyuan Fan +3
In recent years, multimodal large language models (MLLMs) have made significant progress but continue to face inherent challenges in multimodal reasoning, which requires multi-leve…
Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization
Shujin Wu, Cheng Qian, Yi R. Fung +2
The growing capabilities of large language models (LLMs) present a key challenge of maintaining effective human oversight. Weak-to-strong generalization (W2SG) offers a promising f…
Aligning LLMs with Individual Preferences via Interaction
Shujin Wu, May Fung, Cheng Qian +3
As large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption.…
MACAROON: Training Vision-Language Models To Be Your Engaged Partners
Shujin Wu, Yi R. Fung, Sha Li +3
Large vision-language models (LVLMs), while proficient in following instructions and responding to diverse questions, invariably generate detailed responses even when questions are…
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
Tianyang Xu, Shujin Wu, Shizhe Diao +4
Large language models (LLMs) often generate inaccurate or fabricated information and generally fail to indicate their confidence, which limits their broader applications. Previous…