7 papers
AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models
Jintao Lin, Bowen Dong, Weikang Shi +4
The capability of Unified Multimodal Models (UMMs) to apply world knowledge across diverse tasks remains a critical, unresolved challenge. Existing benchmarks fall short, offering…
MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
Yingying Wang, Xuanhua He, Chen Wu +5
Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispe…
VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
Meng Chu, Senqiao Yang, Haoxuan Che +8
Generative models can now produce photorealistic imagery, yet they still struggle with the long, multi-goal prompts that professional designers issue. To expose this gap and better…
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
Kecheng Chen, Ziru Liu, Xijia Tao +7
Diffusion Language Models (DLMs) have recently achieved significant success due to their any-order generation capabilities. However, existing inference methods typically rely on lo…
Efficient Reasoning via Reward Model
Yuhao Wang, Xiaopeng Li, Cheng Gong +4
Reinforcement learning with verifiable rewards (RLVR) has been shown to enhance the reasoning capabilities of large language models (LLMs), enabling the development of large reason…
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
Ziru Liu, Cheng Gong, Xinyu Fu +7
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a powerful paradigm for facilitating the self-improvement of large language models (LLMs), particularl…