3 papers
cs.CV2025
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
Pha Nguyen, Sailik Sengupta, Girik Malik +2
The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step ta…
cs.LG2024
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
Subhojyoti Mukherjee, Anusha Lalitha, Sailik Sengupta +2
Multi-objective alignment from human feedback (MOAHF) in large language models (LLMs) is a challenging problem as human preferences are complex, multifaceted, and often conflicting…
cs.LG2024
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
Jongwoo Ko, Saket Dingliwal, Bhavana Ganesh +3
Direct alignment algorithms (DAAs), such as direct preference optimization (DPO), have become popular alternatives for Reinforcement Learning from Human Feedback (RLHF) due to thei…