5 papers
CORE: Context-Robust Remasking for Diffusion Language Models
Kevin Zhai, Sabbir Mollah, Zhenyi Wang +1
Standard decoding in Masked Diffusion Models (MDMs) is hindered by context rigidity: tokens are retained based on transient high confidence, often ignoring that early predictions l…
Curriculum-DPO++: Direct Preference Optimization via Data and Model Curricula for Text-to-Image Generation
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu +2
Direct Preference Optimization (DPO) has been proposed as an effective and efficient alternative to reinforcement learning from human feedback (RLHF). However, neither RLHF nor DPO…
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
Yimeng Zhang, Jiri Gesi, Ran Xue +10
LLMs have recently demonstrated strong potential in simulating online shopper behavior. Prior work has improved action prediction by applying SFT on action traces with LLM-generate…
STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery
David Robinson, Animesh Gupta, Rizwan Quershi +2
Despite advancements in rehabilitation protocols, clinical assessment of upper extremity (UE) function after stroke largely remains subjective, relying heavily on therapist observa…
Multi-Party Conversational Agents: A Survey
Sagar Sapkota, Mohammad Saqib Hasan, Mubarak Shah +1
Multi-party Conversational Agents (MPCAs) are systems designed to engage in dialogue with more than two participants simultaneously. Unlike traditional two-party agents, designing…