10 citations · 10 across the 6 of their papers we have counts for
7 papers · 1 filter
Learning Sparse Visual Representations via Spatial-Semantic Factorization
Theodore Zhengde Zhao, Sid Kiblawi, Jianwei Yang +6
Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens th…
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
Leon Liangyu Chen, Haoyu Ma, Zhipeng Fan +11
Unified models can handle both multimodal understanding and generation within a single architecture, yet they typically operate in a single pass without iteratively refining their…
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
Jiaqi Wang, Yuhang Zang, Pan Zhang +31
Detecting objects in real-world scenes is a complex task due to various challenges, including the vast range of object categories, and potential encounters with previously unknown…
Learning to Generate Scene Graph from Natural Language Supervision
Yiwu Zhong, Jing Shi, Jianwei Yang +2
Learning from image-text data has demonstrated recent success for many recognition tasks, yet is currently limited to visual features or individual visual concepts such as objects.…
Object-Centric Diagnosis of Visual Reasoning
Jianwei Yang, Jiayuan Mao, Jiajun Wu +4
When answering questions about an image, it not only needs knowing what -- understanding the fine-grained contents (e.g., objects, relationships) in the image, but also telling why…
Graph R-CNN for Scene Graph Generation
Jianwei Yang, Jiasen Lu, Stefan Lee +2
We propose a novel scene graph generation model called Graph R-CNN, that is both effective and efficient at detecting objects and their relations in images. Our model contains a Re…