Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Thinking with Visual Grounding
Junkai Zhang, Yihe Deng, Kai-Wei Chang +1
Visual thinking should not only sound right; it should show its evidence. While recent vision-language models (VLMs) can produce natural-language reasoning traces, these traces oft…
cs.AI2025
More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment
Yifan Wang, Runjin Chen, Bolian Li +7
Aligning large language models (LLMs) with human values is an increasingly critical step in post-training. Direct Preference Optimization (DPO) has emerged as a simple, yet effecti…