4 citations · 11 across the 34 of their papers we have counts for
18 papers · 1 filter
Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
Jianmin Chen, Jiaqi Tang, Wei Wei +9
Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthen, models may gradually rely…
Unveiling the Unknown: Open Vocabulary Object Detection with Scene Graphs
Yi Chen, Yinghao Lu, Zhehao Li +4
Open-vocabulary object detection seeks to identify novel object categories that were not part of the training data. Many knowledge distillation-based approaches have shown promisin…
HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing
Ruyi Chen, Lu Zhou, Xiaogang Xu +3
Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation meth…
AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers
Ruibin Min, Yexin Liu, Aimin Pan +5
Representation alignment has recently emerged as an effective paradigm for accelerating Diffusion Transformer training. Despite their success, existing alignment methods typically…
Low-Light Video Enhancement with An Effective Spatial-Temporal Decomposition Paradigm
Xiaogang Xu, Kun Zhou, Tao Hu +4
Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition s…
Class Incremental Medical Image Segmentation via Prototype-Guided Calibration and Dual-Aligned Distillation
Shengqian Zhu, Chengrong Yu, Qiang Wang +6
Class incremental medical image segmentation (CIMIS) aims to preserve knowledge of previously learned classes while learning new ones without relying on old-class labels. However,…