5 papers
Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training
Yuanhao Chiang, Hongbo Duan, Chunru Yang +3
Autoregressive text-to-image (T2I) generation has recently advanced rapidly, yet aligning generated images with human preferences remains challenging. GRPO-style online reinforceme…
IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions
Kai Golan Hashiloni, Daniel Fadlon, Lior Livyatan +3
Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, therefore, requires semantic a…
OVAL: Open-Vocabulary Augmented Memory Model for Lifelong Object Goal Navigation
Jiahua Pei, Yi Liu, Guoping Pan +3
Object Goal Navigation (ObjectNav) refers to an agent navigating to an object in an unseen environment, which is an ability often required in the accomplishment of complex tasks. W…
Belief-Calibrated Multi-Agent Consensus Seeking for Complex NLP Tasks
Wentao Deng, Jiahuan Pei, Zhiwei Xu +3
A multi-agent system (MAS) enhances its capacity to solve complex natural language processing (NLP) tasks through collaboration among multiple agents, where consensus-seeking serve…
MELoRA: Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-Tuning
Pengjie Ren, Chengshun Shi, Shiguang Wu +5
Parameter-efficient fine-tuning (PEFT) is a popular method for tailoring pre-trained large language models (LLMs), especially as the models' scale and the diversity of tasks increa…