11 papers
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu +1
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is th…
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
Zitong Huang, Kaidong Zhang, Yukang Ding +4
Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization (DPO) methods rely on multi-sam…
CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning
Zhenquan Yao, Zitong Huang, Yihan Zeng +5
Graphical User Interface (GUI) Agents, benefiting from recent advances in multimodal large language models (MLLM), have achieved significant development. However, due to the freque…
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
Mi Zheng, Guanglei Yang, Zitong Huang +3
With the emergence of transformer-based architectures and large language models (LLMs), the accuracy of road scene perception has substantially advanced. Nonetheless, current road…
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
Bowen Dong, Minheng Ni, Zitong Huang +3
Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse…
MR-GDINO: Efficient Open-World Continual Object Detection
Bowen Dong, Zitong Huang, Guanglei Yang +2
Open-world (OW) recognition and detection models show strong zero- and few-shot adaptation abilities, inspiring their use as initializations in continual learning methods to improv…