papers

Publications (46)

cs.CV2025

Discriminator-Free Direct Preference Optimization for Video Diffusion

Haoran Cheng, Qide Dong, Liang Peng +7

Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. Howe…

cs.IR2026

DeGRe: Dense-supervised Generative Reranking for Recommendation

Chaotian Song, Jingyao Zhang, Chenghao Chen +6

In multi-stage recommender systems, reranking optimizes overall utility by capturing intra-list contextual dependencies, yet its central challenge lies in exploring optimal sequenc…

cs.LG2023

Exploring the Relationship between Architecture and Adversarially Robust Generalization

Aishan Liu, Shiyu Tang, Siyuan Liang +4

Adversarial training has been demonstrated to be one of the most effective remedies for defending adversarial examples, yet it often suffers from the huge robustness generalization…

cs.CV2024

SmoothVideo: Smooth Video Synthesis with Noise Constraints on Diffusion Models for One-shot Video Tuning

Liang Peng, Haoran Cheng, Zheng Yang +6

Recent one-shot video tuning methods, which fine-tune the network on a specific video based on pre-trained text-to-image models (e.g., Stable Diffusion), are popular in the communi…

cs.CV2023

CrossFormer++: A Versatile Vision Transformer Hinging on Cross-scale Attention

Wenxiao Wang, Wei Chen, Qibo Qiu +5

While features of different scales are perceptually important to visual inputs, existing vision transformers do not yet take advantage of them explicitly. To this end, we first pro…

cs.CV2023

APPT : Asymmetric Parallel Point Transformer for 3D Point Cloud Understanding

Hengjia Li, Tu Zheng, Zhihao Chi +5

Transformer-based networks have achieved impressive performance in 3D point cloud understanding. However, most of them concentrate on aggregating local features, but neglect to dir…

cs.CV2024

Local Conditional Controlling for Text-to-Image Diffusion Models

Yibo Zhao, Liang Peng, Yang Yang +9

Diffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the genera…

cs.CV2025

MagicView: Multi-View Consistent Identity Customization via Priors-Guided In-Context Learning

Hengjia Li, Jianjin Xu, Keli Cheng +5

Recent advances in personalized generative models have demonstrated impressive capabilities in producing identity-consistent images of the same individual across diverse scenes. Ho…

cs.CL2025

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +1340

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consist…

cs.CV2025

HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment

Lifan Jiang, Boxi Wu, Jiahui Zhang +2

With the rapid development of AIGC technology, significant progress has been made in diffusion model-based technologies for text-to-image (T2I) and text-to-video (T2V). In recent y…

cs.CV2022

Towards Efficient Adversarial Training on Vision Transformers

Boxi Wu, Jindong Gu, Zhifeng Li +3

Vision Transformer (ViT), as a powerful alternative to Convolutional Neural Network (CNN), has received much attention. Recent work showed that ViTs are also vulnerable to adversar…

cs.CV2022

Towards In-distribution Compatibility in Out-of-distribution Detection

Boxi Wu, Jie Jiang, Haidong Ren +7

Deep neural network, despite its remarkable capability of discriminating targeted in-distribution samples, shows poor performance on detecting anomalous out-of-distribution data. T…

cs.CV2026

SNR-Edit: Structure-Aware Noise Rectification for Inversion-Free Flow-Based Editing

Lifan Jiang, Boxi Wu, Yuhang Pei +5

Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing approaches rely on fixed Gaussian noise to co…

cs.CV2024

ConsistencyTrack: A Robust Multi-Object Tracker with a Generation Strategy of Consistency Model

Lifan Jiang, Zhihui Wang, Siqi Yin +3

Multi-object tracking (MOT) is a critical technology in computer vision, designed to detect multiple targets in video sequences and assign each target a unique ID per frame. Existe…

cs.CV2025

RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence

Xuming He, Zehao Fan, Hengjia Li +7

Recent advances in video generation have enabled the synthesis of videos with strong temporal consistency and impressive visual quality, marking a crucial step toward vision founda…

cs.CV2024

UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation

Hengjia Li, Yang Liu, Yuqi Lin +8

Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt…

cs.CV2026

UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs

Lifan Jiang, Tianrun Wu, Yuhang Pei +3

The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross-paradigm c…

cs.CV2024

GCA-3D: Towards Generalized and Consistent Domain Adaptation of 3D Generators

Hengjia Li, Yang Liu, Yibo Zhao +9

Recently, 3D generative domain adaptation has emerged to adapt the pre-trained generator to other domains without collecting massive datasets and camera pose distributions. Typical…

cs.CV2026

Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion

Yang Yang, Tianyi Zhang, Wei Huang +6

Interactive long video generation requires prompt switching to introduce new subjects or events, while maintaining perceptual fidelity and coherent motion over extended horizons. R…

cs.CV2025

Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking

Haonan Zhang, Xinyao Wang, Boxi Wu +3

3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters,…

cs.CV2026

PhyRPR: Training-Free Physics-Constrained Video Generation

Yibo Zhao, Hengjia Li, Xiaofei He +1

Recent diffusion-based video generation models can synthesize visually plausible videos, yet they often struggle to satisfy physical constraints. A key reason is that most existing…

cs.CV2024

Object Detectors in the Open Environment: Challenges, Solutions, and Outlook

Siyuan Liang, Wei Wang, Ruoyu Chen +5

With the emergence of foundation models, deep learning-based object detectors have shown practical usability in closed set scenarios. However, for real-world tasks, object detector…

cs.CV2025

SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization

Liang Peng, Boxi Wu, Haoran Cheng +2

Previous text-to-image diffusion models typically employ supervised fine-tuning (SFT) to enhance pre-trained base models. However, this approach primarily minimizes the loss of mea…

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…

cs.CV2023

Learning Occupancy for Monocular 3D Object Detection

Liang Peng, Junkai Xu, Haoran Cheng +6

Monocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth estimates to fac…

cs.LG2022

Improving alignment of dialogue agents via targeted human judgements

Amelia Glaese, Nat McAleese, Maja Trębacz +31

We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…

cs.CV2022

WeakM3D: Towards Weakly Supervised Monocular 3D Object Detection

Liang Peng, Senbo Yan, Boxi Wu +3

Monocular 3D object detection is one of the most challenging tasks in 3D scene understanding. Due to the ill-posed nature of monocular imagery, existing monocular 3D detection meth…

cs.LG2021

Attacking Adversarial Attacks as A Defense

Boxi Wu, Heng Pan, Li Shen +6

It is well known that adversarial attacks can fool deep neural networks with imperceptible perturbations. Although adversarial training significantly improves model robustness, fai…

cs.CV2024

Searching Priors Makes Text-to-Video Synthesis Better

Haoran Cheng, Liang Peng, Linxuan Xia +5

Significant advancements in video diffusion models have brought substantial progress to the field of text-to-video (T2V) synthesis. However, existing T2V synthesis model struggle t…

cs.CV2019

Improving Semantic Segmentation via Dilated Affinity

Boxi Wu, Shuai Zhao, Wenqing Chu +2

Introducing explicit constraints on the structural predictions has been an effective way to improve the performance of semantic segmentation models. Existing methods are mainly bas…

cs.CV2025

MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization

Hengjia Li, Lifan Jiang, Xi Xiao +4

Video identity customization seeks to produce high-fidelity videos that maintain consistent identity and exhibit significant dynamics based on users' reference images. However, exi…

cs.CV2023

NormKD: Normalized Logits for Knowledge Distillation

Zhihao Chi, Tu Zheng, Hengjia Li +4

Logit based knowledge distillation gets less attention in recent years since feature based methods perform better in most cases. Nevertheless, we find it still has untapped potenti…

cs.CV2023

CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

Yuqi Lin, Minghao Chen, Wenxiao Wang +5

Weakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training cos…

cs.CV2025

VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control

Lifan Jiang, Shuang Chen, Boxi Wu +2

With the advancement of generative artificial intelligence, previous studies have achieved the task of generating aesthetic images from hand-drawn sketches, fulfilling the public's…

cs.CV2024

LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion Models

Yang Yang, Wen Wang, Liang Peng +8

Customization generation techniques have significantly advanced the synthesis of specific concepts across varied contexts. Multi-concept customization emerges as the challenging ta…

cs.CV2023

Self-supervised and Weakly Supervised Contrastive Learning for Frame-wise Action Representations

Minghao Chen, Renbo Tu, Chenxi Huang +3

Previous work on action representation learning focused on global representations for short video clips. In contrast, many practical applications, such as video alignment, strongly…

cs.CV2023

GD-MAE: Generative Decoder for MAE Pre-training on LiDAR Point Clouds

Honghui Yang, Tong He, Jiaheng Liu +5

Despite the tremendous progress of Masked Autoencoders (MAE) in developing vision tasks such as image and video, exploring MAE in large-scale 3D point clouds remains challenging du…

cs.CV2025

PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation

Hengjia Li, Haonan Qiu, Shiwei Zhang +6

The current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video g…

cs.LG2021

Do Wider Neural Networks Really Help Adversarial Robustness?

Boxi Wu, Jinghui Chen, Deng Cai +2

Adversarial training is a powerful type of defense against adversarial examples. Previous empirical results suggest that adversarial training requires wider networks for better per…

cs.CV2025

Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model

Yang Yang, Siming Zheng, Qirui Yang +6

Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based boke…

cs.CV2026

ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing

Hengjia Li, Liming Jiang, Qing Yan +6

Instruction-driven image editing with unified multimodal generative models has advanced rapidly, yet their underlying visual reasoning remains limited, leading to suboptimal perfor…

cs.LG2026

RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization

Linxuan Xia, Xiaolong Yang, Yongyuan Chen +4

Aligning large language models (LLMs) on domain-specific data remains a fundamental challenge. Supervised fine-tuning (SFT) offers a straightforward way to inject domain knowledge…

cs.CV2019

Correlation Maximized Structural Similarity Loss for Semantic Segmentation

Shuai Zhao, Boxi Wu, Wenqing Chu +2

Most semantic segmentation models treat semantic segmentation as a pixel-wise classification task and use a pixel-wise classification error as their optimization criterions. Howeve…

cs.CV2023

One-shot Implicit Animatable Avatars with Model-based Priors

Yangyi Huang, Hongwei Yi, Weiyang Liu +6

Existing neural rendering methods for creating human avatars typically either require dense input signals such as video or multi-view images, or leverage a learned prior from large…

cs.CV2024

TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation

Xiaopei Wu, Yuenan Hou, Xiaoshui Huang +8

Training deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the sparsity p…

cs.CV2024

Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-dataset 3D Object Detection

Zhanwei Zhang, Minghao Chen, Shuai Xiao +7

Recent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels,…