1 citations · 1 across the 11 of their papers we have counts for
14 papers
Iterative Refinement Improves Compositional Image Generation
Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj +5
Text-to-image (T2I) models have achieved remarkable progress, yet they continue to struggle with complex prompts that require simultaneously handling multiple objects, relations, a…
Chat with UAV -- Human-UAV Interaction Based on Large Language Models
Haoran Wang, Zhuohang Chen, Guang Li +2
The future of UAV interaction systems is evolving from engineer-driven to user-driven, aiming to replace traditional predefined Human-UAV Interaction designs. This shift focuses on…
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
Chia-Yu Hung, Navonil Majumder, Haoyuan Deng +7
Vision--language--action (VLA) models have recently shown promising performance on a variety of embodied tasks, yet they still fall short in reliability and generalization, especia…
10 Open Challenges Steering the Future of Vision-Language-Action Models
Soujanya Poria, Navonil Majumder, Chia-Yu Hung +7
Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly prevalent in the embodied AI arena, following the widespread succ…
From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs
Haonan Wang, Weida Liang, Zihang Fu +8
Recent reasoning LLMs (RLMs), especially those trained with verifier-based reinforcement learning, often perform worse with few-shot CoT than with direct answering. We revisit this…
FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models
Chuan Li, Qianyi Zhao, Fengran Mo +1
Efficiently enhancing the reasoning capabilities of large language models (LLMs) in federated learning environments remains challenging, particularly when balancing performance gai…