activity
20242026
most citedSegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Yunlong Lin, Zixu Lin, Zhaohu Xing +23

Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, aud…

cs.CV2026

GEAR: Guided End-to-End AutoRegression for Image Synthesis

Bin Lin, Zheyuan Liu, Chenguo Lin +8

Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator is trained on its discrete in…

cs.CV20261 cited

SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation

Sicheng Yang, Hongqiu Wang, Zhaohu Xing +5

Self-supervised DINO models provide strong transferable visual representations, yet applying them directly to image segmentation remains challenging. Existing approaches commonly r…

cs.CV2026

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

Sixiang Chen, Zhaohu Xing, Tian Ye +7

Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal generative ability with ext…

cs.CV2026

Latent Action Control for Reasoning-Guided Unified Image Generation

Fuxiang Zhai, Sixiang Chen, Yingjin Li +4

Unified multimodal models can encode visual understanding and image generation within a shared backbone, yet understanding does not automatically translate into control: models may…

cs.CV2026

LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution

Song Fei, Tian Ye, Sixiang Chen +3

Generative real-world image super-resolution (Real-ISR) can synthesize visually convincing details from severely degraded low-resolution (LR) inputs, yet its stochastic sampling ma…