papers
Publications (11)
cs.CV2022
Exploring Plain Vision Transformer Backbones for Object Detection
Yanghao Li, Hanzi Mao, Ross Girshick +1
cs.LG2025
Data-regularized Reinforcement Learning for Diffusion Models at Scale
Haotian Ye, Kaiwen Zheng, Jiashu Xu +15
cs.CV2025
Describe Anything: Detailed Localized Image and Video Captioning
Long Lian, Yifan Ding, Yunhao Ge +8
cs.CL2021
Entailment as Few-Shot Learner
Sinong Wang, Han Fang, Madian Khabsa +2
cs.CV2025
Cosmos World Foundation Model Platform for Physical AI
NVIDIA, :, Niket Agarwal +76
cs.CV2023
Segment Anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi +9
cs.LG2025
Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling
Kaiwen Zheng, Yongxin Chen, Hanzi Mao +3
cs.CV2026
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
cs.RO2025
Scalable Policy Evaluation with Video World Models
Wei-Cheng Tseng, Jinwei Gu, Qinsheng Zhang +4
cs.CV2022
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu +3
cs.CV2025
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
NVIDIA, :, Hassan Abu Alhaija +38