activity
20232026
most citedRankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners

2 citations · 2 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation

Chao Li, Tianhong Li, Sai Vidyaranya Nuthalapati +9

Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes:…

cs.IR2026

Reason to Contrast: A Cascaded Multimodal Retrieval Framework

Xuanming Cui, Hong-You Chen, Hao Yu +10

Traditional multimodal retrieval systems rely primarily on bi-encoder architectures, where performance is closely tied to embedding dimensionality. Recent work, Think-Then-Embed (T…

cs.CV2026

Xray-Visual Models: Scaling Vision models on Industry Scale Data

Shlok Mishra, Tsung-Yu Lin, Linda Wang +24

We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…

cs.CL20242 cited

RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners

Chi Hu, Yuan Ge, Xiangnan Ma +5

Large Language Models (LLMs) have achieved impressive performance across various reasoning tasks. However, even state-of-the-art LLMs such as ChatGPT are prone to logical errors du…

cs.CV2023

Towards the Unification of Generative and Discriminative Visual Foundation Model: A Survey

Xu Liu, Tong Zhou, Yuanxin Wang +7

The advent of foundation models, which are pre-trained on vast datasets, has ushered in a new era of computer vision, characterized by their robustness and remarkable zero-shot gen…