activity
20232026
most citedGrounded SAM: Assembling Open-World Models for Diverse Visual Tasks

93 citations · 119 across the 11 of their papers we have counts for

collaborators

13 papers

cs.AI2026

QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge

Yuanzhe Zhou, Zhaoyang Zeng

Post-click conversion rate (pCVR) prediction requires jointly modeling feature interactions and sequential user behaviors. The KDD Cup 2026 Tencent UniRec Challenge calls for a uni…

cs.CV2025

Detect Anything via Next Point Prediction

Qing Jiang, Junan Huo, Xingyu Chen +6

Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts have attempted to levera…

cs.CV2025

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

Qing Jiang, Xingyu Chen, Zhaoyang Zeng +2

Object referring aims to detect all objects in an image that match a given natural language description. We argue that a robust object referring model should be grounded, meaning i…

cs.CV2025

Referring to Any Person

Qing Jiang, Lin Wu, Zhaoyang Zeng +5

Humans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referr…

cs.CV2024

TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video

Jinyuan Qu, Hongyang Li, Shilong Liu +3

In this paper, built upon TAPTRv2, we present TAPTRv3. TAPTRv2 is a simple yet effective DETR-like point tracking framework that works fine in regular videos but tends to fail in l…

cs.CV2024

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Qing Jiang, Gen Luo, Yuqin Yang +5

Perception and understanding are two pillars of computer vision. While multimodal large language models (MLLM) have demonstrated remarkable visual understanding capabilities, they…