Publications (57)
MM-VID: Advancing Video Understanding with GPT-4V(ision)
Kevin Lin, Faisal Ahmed, Linjie Li +9
Learning the Depths of Moving People by Watching Frozen People
Zhengqi Li, Tali Dekel, Forrester Cole +4
ReCo: Region-Controlled Text-to-Image Generation
Zhengyuan Yang, Jianfeng Wang, Zhe Gan +8
LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling
Linjie Li, Zhe Gan, Kevin Lin +4
DicFace: Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face Restoration
Yan Chen, Hanlin Shang, Ce Liu +6
GIT: A Generative Image-to-text Transformer for Vision and Language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu +6
Deep 3D-to-2D Watermarking: Embedding Messages in 3D Meshes and Extracting Them from 2D Renderings
Innfarn Yoo, Huiwen Chang, Xiyang Luo +4
Dynamic PDB: A New Dataset and a SE(3) Model Extension by Integrating Dynamic Behaviors and Physical Properties in Protein Structures
Ce Liu, Jun Wang, Zhiqiang Cai +10
Single Image Depth Prediction Made Better: A Multivariate Gaussian Take
Ce Liu, Suryansh Kumar, Shuhang Gu +2
COMISR: Compression-Informed Video Super-Resolution
Yinxiao Li, Pengchong Jin, Feng Yang +3
DepthTransfer: Depth Extraction from Video Using Non-parametric Sampling
Kevin Karsch, Ce Liu, Sing Bing Kang
On Conservative Stable Standard of Behavior and Perfect Coalitional Equilibrium
S. Nageeb Ali, Ce Liu
Adaptively Learning the Crowd Kernel
Omer Tamuz, Ce Liu, Serge Belongie +2
Regularizing Generative Adversarial Networks under Limited Data
Hung-Yu Tseng, Lu Jiang, Ce Liu +2
AutoFlow: Learning a Better Training Set for Optical Flow
Deqing Sun, Daniel Vlasic, Charles Herrmann +6
Self-Enforced Job Matching
Ce Liu, Ziwei Wang, Hanzhe Zhang
Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone
Zi-Yi Dou, Aishwarya Kamath, Zhe Gan +9
NeRD: Neural Reflectance Decomposition from Image Collections
Mark Boss, Raphael Braun, Varun Jampani +3
Unified Contrastive Learning in Image-Text-Label Space
Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4
Microwave Photonic Imaging Radar with a Millimeter-level Resolution
Cong Ma, Yue Yang, Ce Liu +5
Visual Clues: Bridging Vision and Language Foundations for Image Paragraph Captioning
Yujia Xie, Luowei Zhou, Xiyang Dai +4
Florence: A New Foundation Model for Computer Vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen +20
Stereo Risk: A Continuous Modeling Approach to Stereo Matching
Ce Liu, Suryansh Kumar, Shuhang Gu +3
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Marah Abdin, Jyoti Aneja, Hany Awadalla +126
DVMark: A Deep Multiscale Framework for Video Watermarking
Xiyang Luo, Yinxiao Li, Huiwen Chang +3
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +556
Credible Persuasion
Xiao Lin, Ce Liu
LASR: Learning Articulated Shape Reconstruction from a Monocular Video
Gengshan Yang, Deqing Sun, Varun Jampani +6
IPRU: Input-Perturbation-based Radio Frequency Fingerprinting Unlearning for LAWNs
Ce Liu, Rui Meng, Yinqiu Liu +4
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Zhengyuan Yang, Linjie Li, Jianfeng Wang +7
LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results
Xiang Chen, Hao Li, Jiangxin Dong +54
X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusion
Hanqing Zhao, Dianmo Sheng, Jianmin Bao +9
Robust image stitching with multiple registrations
Charles Herrmann, Chen Wang, Richard Strong Bowen +4
SLIDE: Single Image 3D Photography with Soft Layering and Depth-aware Inpainting
Varun Jampani, Huiwen Chang, Kyle Sargent +8
Neural-PIL: Neural Pre-Integrated Lighting for Reflectance Decomposition
Mark Boss, Varun Jampani, Raphael Braun +3
Supervised Contrastive Learning
Prannay Khosla, Piotr Teterwak, Chen Wang +6
Simulating the Real World: A Unified Survey of Multimodal Generative Models
Yuqi Hu, Longguang Wang, Xian Liu +7
Pyramid Adversarial Training Improves ViT Performance
Charles Herrmann, Kyle Sargent, Lu Jiang +5
Coalitions in Repeated Games
S. Nageeb Ali, Ce Liu
Depth Extraction from Video Using Non-parametric Sampling
Kevin Karsch, Ce Liu, Sing Bing Kang
Smart, Sparse Contours to Represent and Edit Images
Tali Dekel, Chuang Gan, Dilip Krishnan +2
AlphaFolding: 4D Diffusion for Dynamic Protein Structure Prediction with Reference and Motion Guidance
Kaihui Cheng, Ce Liu, Qingkun Su +6
ViTGAN: Training GANs with Vision Transformers
Kwonjoon Lee, Huiwen Chang, Lu Jiang +3
NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose Annotations
Varun Jampani, Kevis-Kokitsi Maninis, Andreas Engelhardt +13
Movie Gen: A Cast of Media Foundation Models
Adam Polyak, Amit Zohar, Andrew Brown +85
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Bin Xiao, Haiping Wu, Weijian Xu +6
Learning Customized Visual Models with Retrieval-Augmented Knowledge
Haotian Liu, Kilho Son, Jianwei Yang +4
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
Tingshu Mou, Jiabo He, Renying Wang +5
TDGCN-Based Mobile Multiuser Physical-Layer Authentication for EI-Enabled IIoT
Rui Meng, Hangyu Zhao, Liang Jin +3
K-LITE: Learning Transferable Visual Models with External Knowledge
Sheng Shen, Chunyuan Li, Xiaowei Hu +11
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
Mingwang Xu, Hui Li, Qingkun Su +6
VA-DepthNet: A Variational Approach to Single Image Depth Prediction
Ce Liu, Suryansh Kumar, Shuhang Gu +2
Boundless: Generative Adversarial Networks for Image Extension
Piotr Teterwak, Aaron Sarna, Dilip Krishnan +4
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
Junke Wang, Dongdong Chen, Zuxuan Wu +7
DroneSR: Rethinking Few-shot Thermal Image Super-Resolution from Drone-based Perspective
Zhipeng Weng, Xiaopeng Liu, Ce Liu +3
MaskGIT: Masked Generative Image Transformer
Huiwen Chang, Han Zhang, Lu Jiang +2
Stability in Repeated Matching Markets
Ce Liu