6 papers
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
Chia-Yu Hung, Navonil Majumder, Haoyuan Deng +7
Vision--language--action (VLA) models have recently shown promising performance on a variety of embodied tasks, yet they still fall short in reliability and generalization, especia…
10 Open Challenges Steering the Future of Vision-Language-Action Models
Soujanya Poria, Navonil Majumder, Chia-Yu Hung +7
Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly prevalent in the embodied AI arena, following the widespread succ…
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
Renhang Liu, Chia-Yu Hung, Navonil Majumder +5
Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and fait…
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
Chia-Yu Hung, Qi Sun, Pengfei Hong +5
Existing Visual-Language-Action (VLA) models have shown promising performance in zero-shot scenarios, demonstrating impressive task execution and reasoning capabilities. However, a…
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
Chia-Yu Hung, Navonil Majumder, Zhifeng Kong +6
We introduce TangoFlux, an efficient Text-to-Audio (TTA) generative model with 515M parameters, capable of generating up to 30 seconds of 44.1kHz audio in just 3.7 seconds on a sin…
Inference Time Alignment with Reward-Guided Tree Search
Chia-Yu Hung, Navonil Majumder, Ambuj Mehrish +1
Inference-time computation methods enhance the performance of Large Language Models (LLMs) by leveraging additional computational resources to achieve superior results. Common tech…