4 papers
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
Chia-Yu Hung, Navonil Majumder, Haoyuan Deng +7
Vision--language--action (VLA) models have recently shown promising performance on a variety of embodied tasks, yet they still fall short in reliability and generalization, especia…
Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
Maojia Song, Renhang Liu, Xinyu Wang +6
RAG (Retrieval-Augmented Generation) systems and web agents are increasingly evaluated on multi-hop deep search tasks, yet current practice suffers from two major limitations. Firs…
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
Renhang Liu, Chia-Yu Hung, Navonil Majumder +5
Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and fait…
JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
Abhinaba Roy, Renhang Liu, Tongyu Lu +1
We introduce JamendoMaxCaps, a large-scale music-caption dataset featuring over 362,000 freely licensed instrumental tracks from the renowned Jamendo platform. The dataset includes…