1 citations · 1 across the 3 of their papers we have counts for
24 papers
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
Heyang Jiang, Henry Liu, Baharan Mirzasoleiman
Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful i…
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
Yihao Xue, Kristjan Greenewald, Youssef Mroueh +1
Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications. In black-box settings, several self-consistency-based techniques have been pro…
Reasoning Quality Emerges Early: Data Curation for Reasoning Models
Hongyi Henry Jin, Wenhan Yang, Meysam Ghaffari +2
Supervised fine-tuning (SFT) on a small, high-quality set of long reasoning traces is an effective approach for eliciting strong reasoning capabilities in Large Language Models (LL…
Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models
Seongbin Park, Fan Zhang, Baharan Mirzasoleiman +2
Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these policies offer no guarantees…
ProbeAct: Probe-Guided Training-Free Failure Recovery in Vision-Language-Action Models
Fan Zhang, Seongbin Park, Baharan Mirzasoleiman +2
Vision-Language-Action (VLA) models demonstrate strong perfor-1 mance on language-conditioned robotic manipulation within their training dis-2 tribution, yet their generalization c…
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from -Parity
Jianhao Huang, Baharan Mirzasoleiman
Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understudied compared to their auto-regressive…