collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Test-Time Registers as Global Priors for Tokenized Image Generation

Cheng-Yao Hong, Yifan Wang, Yuewei Lin +1

Attention-based models often develop attention sinks, where a small number of tokens repeatedly attract attention and accumulate unusually large activations. In vision transformers…

cs.CV2026

HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles

Yifan Wang, Francesco Pittaluga, Zaid Tasneem +3

Controllable driving scene generation is critical for realistic and scalable autonomous driving simulation, yet existing approaches struggle to jointly achieve photorealism and pre…

cs.CV2025

Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting

Wen Zhang, Qin Ren, Wenjing Liu +2

Accurate nuclear instance segmentation is a pivotal task in computational pathology, supporting data-driven clinical insights and facilitating downstream translational applications…

cs.CV2025

Scale Where It Matters: Training-Free Localized Scaling for Diffusion Models

Qin Ren, Yufei Wang, Lanqing Guo +3

Diffusion models have become the dominant paradigm in text-to-image generation, and test-time scaling (TTS) improves sample quality by allocating additional computation at inferenc…

cs.CV2025

Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis

Wenjing Liu, Qin Ren, Wen Zhang +2

Multimodal survival analysis aims to improve cancer prognosis using heterogeneous biomedical data, such as histopathology images and genomic profiles. A common strategy is to align…

cs.CV2025

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

Lingxiao Li, Yifan Wang, Xinyan Gao +3

Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language mo…