93 citations · 119 across the 11 of their papers we have counts for
13 papers
QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge
Yuanzhe Zhou, Zhaoyang Zeng
Post-click conversion rate (pCVR) prediction requires jointly modeling feature interactions and sequential user behaviors. The KDD Cup 2026 Tencent UniRec Challenge calls for a uni…
Detect Anything via Next Point Prediction
Qing Jiang, Junan Huo, Xingyu Chen +6
Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts have attempted to levera…
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
Qing Jiang, Xingyu Chen, Zhaoyang Zeng +2
Object referring aims to detect all objects in an image that match a given natural language description. We argue that a robust object referring model should be grounded, meaning i…
Referring to Any Person
Qing Jiang, Lin Wu, Zhaoyang Zeng +5
Humans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referr…
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
Jinyuan Qu, Hongyang Li, Shilong Liu +3
In this paper, built upon TAPTRv2, we present TAPTRv3. TAPTRv2 is a simple yet effective DETR-like point tracking framework that works fine in regular videos but tends to fail in l…
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Qing Jiang, Gen Luo, Yuqin Yang +5
Perception and understanding are two pillars of computer vision. While multimodal large language models (MLLM) have demonstrated remarkable visual understanding capabilities, they…