Publications (46)
Discriminator-Free Direct Preference Optimization for Video Diffusion
Haoran Cheng, Qide Dong, Liang Peng +7
Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. Howe…
DeGRe: Dense-supervised Generative Reranking for Recommendation
Chaotian Song, Jingyao Zhang, Chenghao Chen +6
In multi-stage recommender systems, reranking optimizes overall utility by capturing intra-list contextual dependencies, yet its central challenge lies in exploring optimal sequenc…
Exploring the Relationship between Architecture and Adversarially Robust Generalization
Aishan Liu, Shiyu Tang, Siyuan Liang +4
Adversarial training has been demonstrated to be one of the most effective remedies for defending adversarial examples, yet it often suffers from the huge robustness generalization…
SmoothVideo: Smooth Video Synthesis with Noise Constraints on Diffusion Models for One-shot Video Tuning
Liang Peng, Haoran Cheng, Zheng Yang +6
Recent one-shot video tuning methods, which fine-tune the network on a specific video based on pre-trained text-to-image models (e.g., Stable Diffusion), are popular in the communi…
CrossFormer++: A Versatile Vision Transformer Hinging on Cross-scale Attention
Wenxiao Wang, Wei Chen, Qibo Qiu +5
While features of different scales are perceptually important to visual inputs, existing vision transformers do not yet take advantage of them explicitly. To this end, we first pro…
APPT : Asymmetric Parallel Point Transformer for 3D Point Cloud Understanding
Hengjia Li, Tu Zheng, Zhihao Chi +5
Transformer-based networks have achieved impressive performance in 3D point cloud understanding. However, most of them concentrate on aggregating local features, but neglect to dir…
Local Conditional Controlling for Text-to-Image Diffusion Models
Yibo Zhao, Liang Peng, Yang Yang +9
Diffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the genera…
MagicView: Multi-View Consistent Identity Customization via Priors-Guided In-Context Learning
Hengjia Li, Jianjin Xu, Keli Cheng +5
Recent advances in personalized generative models have demonstrated impressive capabilities in producing identity-consistent images of the same individual across diverse scenes. Ho…
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud +1340
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consist…
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
Lifan Jiang, Boxi Wu, Jiahui Zhang +2
With the rapid development of AIGC technology, significant progress has been made in diffusion model-based technologies for text-to-image (T2I) and text-to-video (T2V). In recent y…
Towards Efficient Adversarial Training on Vision Transformers
Boxi Wu, Jindong Gu, Zhifeng Li +3
Vision Transformer (ViT), as a powerful alternative to Convolutional Neural Network (CNN), has received much attention. Recent work showed that ViTs are also vulnerable to adversar…
Towards In-distribution Compatibility in Out-of-distribution Detection
Boxi Wu, Jie Jiang, Haidong Ren +7
Deep neural network, despite its remarkable capability of discriminating targeted in-distribution samples, shows poor performance on detecting anomalous out-of-distribution data. T…
SNR-Edit: Structure-Aware Noise Rectification for Inversion-Free Flow-Based Editing
Lifan Jiang, Boxi Wu, Yuhang Pei +5
Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing approaches rely on fixed Gaussian noise to co…
ConsistencyTrack: A Robust Multi-Object Tracker with a Generation Strategy of Consistency Model
Lifan Jiang, Zhihui Wang, Siqi Yin +3
Multi-object tracking (MOT) is a critical technology in computer vision, designed to detect multiple targets in video sequences and assign each target a unique ID per frame. Existe…
RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
Xuming He, Zehao Fan, Hengjia Li +7
Recent advances in video generation have enabled the synthesis of videos with strong temporal consistency and impressive visual quality, marking a crucial step toward vision founda…
UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation
Hengjia Li, Yang Liu, Yuqi Lin +8
Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt…
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
Lifan Jiang, Tianrun Wu, Yuhang Pei +3
The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross-paradigm c…
GCA-3D: Towards Generalized and Consistent Domain Adaptation of 3D Generators
Hengjia Li, Yang Liu, Yibo Zhao +9
Recently, 3D generative domain adaptation has emerged to adapt the pre-trained generator to other domains without collecting massive datasets and camera pose distributions. Typical…
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
Yang Yang, Tianyi Zhang, Wei Huang +6
Interactive long video generation requires prompt switching to introduce new subjects or events, while maintaining perceptual fidelity and coherent motion over extended horizons. R…
Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
Haonan Zhang, Xinyao Wang, Boxi Wu +3
3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters,…
PhyRPR: Training-Free Physics-Constrained Video Generation
Yibo Zhao, Hengjia Li, Xiaofei He +1
Recent diffusion-based video generation models can synthesize visually plausible videos, yet they often struggle to satisfy physical constraints. A key reason is that most existing…
Object Detectors in the Open Environment: Challenges, Solutions, and Outlook
Siyuan Liang, Wei Wang, Ruoyu Chen +5
With the emergence of foundation models, deep learning-based object detectors have shown practical usability in closed set scenarios. However, for real-world tasks, object detector…
SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
Liang Peng, Boxi Wu, Haoran Cheng +2
Previous text-to-image diffusion models typically employ supervised fine-tuning (SFT) to enhance pre-trained base models. However, this approach primarily minimizes the loss of mea…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
Learning Occupancy for Monocular 3D Object Detection
Liang Peng, Junkai Xu, Haoran Cheng +6
Monocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth estimates to fac…
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja TrÄbacz +31
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…
WeakM3D: Towards Weakly Supervised Monocular 3D Object Detection
Liang Peng, Senbo Yan, Boxi Wu +3
Monocular 3D object detection is one of the most challenging tasks in 3D scene understanding. Due to the ill-posed nature of monocular imagery, existing monocular 3D detection meth…
Attacking Adversarial Attacks as A Defense
Boxi Wu, Heng Pan, Li Shen +6
It is well known that adversarial attacks can fool deep neural networks with imperceptible perturbations. Although adversarial training significantly improves model robustness, fai…
Searching Priors Makes Text-to-Video Synthesis Better
Haoran Cheng, Liang Peng, Linxuan Xia +5
Significant advancements in video diffusion models have brought substantial progress to the field of text-to-video (T2V) synthesis. However, existing T2V synthesis model struggle t…
Improving Semantic Segmentation via Dilated Affinity
Boxi Wu, Shuai Zhao, Wenqing Chu +2
Introducing explicit constraints on the structural predictions has been an effective way to improve the performance of semantic segmentation models. Existing methods are mainly bas…
MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization
Hengjia Li, Lifan Jiang, Xi Xiao +4
Video identity customization seeks to produce high-fidelity videos that maintain consistent identity and exhibit significant dynamics based on users' reference images. However, exi…
NormKD: Normalized Logits for Knowledge Distillation
Zhihao Chi, Tu Zheng, Hengjia Li +4
Logit based knowledge distillation gets less attention in recent years since feature based methods perform better in most cases. Nevertheless, we find it still has untapped potenti…
CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation
Yuqi Lin, Minghao Chen, Wenxiao Wang +5
Weakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training cos…
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control
Lifan Jiang, Shuang Chen, Boxi Wu +2
With the advancement of generative artificial intelligence, previous studies have achieved the task of generating aesthetic images from hand-drawn sketches, fulfilling the public's…
LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion Models
Yang Yang, Wen Wang, Liang Peng +8
Customization generation techniques have significantly advanced the synthesis of specific concepts across varied contexts. Multi-concept customization emerges as the challenging ta…
Self-supervised and Weakly Supervised Contrastive Learning for Frame-wise Action Representations
Minghao Chen, Renbo Tu, Chenxi Huang +3
Previous work on action representation learning focused on global representations for short video clips. In contrast, many practical applications, such as video alignment, strongly…
GD-MAE: Generative Decoder for MAE Pre-training on LiDAR Point Clouds
Honghui Yang, Tong He, Jiaheng Liu +5
Despite the tremendous progress of Masked Autoencoders (MAE) in developing vision tasks such as image and video, exploring MAE in large-scale 3D point clouds remains challenging du…
PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation
Hengjia Li, Haonan Qiu, Shiwei Zhang +6
The current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video g…
Do Wider Neural Networks Really Help Adversarial Robustness?
Boxi Wu, Jinghui Chen, Deng Cai +2
Adversarial training is a powerful type of defense against adversarial examples. Previous empirical results suggest that adversarial training requires wider networks for better per…
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
Yang Yang, Siming Zheng, Qirui Yang +6
Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based boke…
ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing
Hengjia Li, Liming Jiang, Qing Yan +6
Instruction-driven image editing with unified multimodal generative models has advanced rapidly, yet their underlying visual reasoning remains limited, leading to suboptimal perfor…
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
Linxuan Xia, Xiaolong Yang, Yongyuan Chen +4
Aligning large language models (LLMs) on domain-specific data remains a fundamental challenge. Supervised fine-tuning (SFT) offers a straightforward way to inject domain knowledge…
Correlation Maximized Structural Similarity Loss for Semantic Segmentation
Shuai Zhao, Boxi Wu, Wenqing Chu +2
Most semantic segmentation models treat semantic segmentation as a pixel-wise classification task and use a pixel-wise classification error as their optimization criterions. Howeve…
One-shot Implicit Animatable Avatars with Model-based Priors
Yangyi Huang, Hongwei Yi, Weiyang Liu +6
Existing neural rendering methods for creating human avatars typically either require dense input signals such as video or multi-view images, or leverage a learned prior from large…
TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation
Xiaopei Wu, Yuenan Hou, Xiaoshui Huang +8
Training deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the sparsity p…
Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-dataset 3D Object Detection
Zhanwei Zhang, Minghao Chen, Shuai Xiao +7
Recent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels,…