activity
20192026
most citedPolyhistor: Parameter-Efficient Multi-Task Adaptation for Dense Vision Tasks

15 citations · 38 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

Lei Zhang, Junjiao Tian, Zhipeng Fan +9

Humans paint images incrementally: they plan a global layout, sketch a coarse draft, inspect, and refine details, and most importantly, each step is grounded in the evolving visual…

cs.CV2025

Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models

Xu Ma, Peize Sun, Haoyu Ma +22

Autoregressive (AR) models, long dominant in language generation, are increasingly applied to image synthesis but are often considered less competitive than Diffusion-based models.…

cs.CV2024

Grounding Descriptions in Images informs Zero-Shot Visual Recognition

Shaunak Halbe, Junjiao Tian, K J Joseph +4

Vision-language models (VLMs) like CLIP have been cherished for their ability to perform zero-shot visual recognition on open-vocabulary concepts. This is achieved by selecting the…

cs.CV2023★ 1 cited

Fast Trainable Projection for Robust Fine-Tuning

Junjiao Tian, Yen-Cheng Liu, James Seale Smith +1

Robust fine-tuning aims to achieve competitive in-distribution (ID) performance while maintaining the out-of-distribution (OOD) robustness of a pre-trained model when transferring…

cs.CV2023★ 9 cited

Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion

Junjiao Tian, Lavisha Aggarwal, Andrea Colaco +2

Producing quality segmentation masks for images is a fundamental problem in computer vision. Recent research has explored large-scale supervised training to enable zero-shot segmen…

cs.CV2023★ 1 cited

Continual Adaptation of Vision Transformers for Federated Learning

Shaunak Halbe, James Seale Smith, Junjiao Tian +1

In this paper, we focus on the important yet understudied problem of Continual Federated Learning (CFL), where a server communicates with a set of clients to incrementally learn ne…