2 papers
cs.AI2026
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs
Darsh Kachroo, Arjun Prasaath Anbazhagan, Adriana Caraeni +2
Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex reasoning tasks. DPO optimiz…
cs.SD2025
Probing Audio-Generation Capabilities of Text-Based Language Models
Arjun Prasaath Anbazhagan, Parteek Kumar, Ujjwal Kaur +3
How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prom…