collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection

Manwen Yang, Leqian Ding, Yu Guo +1

Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to object-centric bias, normal and anomalous text prototypes exhibit a high s…

cs.CV2026

TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models

Leqian Ding, Junning Qiu, Manwen Yang +2

Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progres…

cs.CV2026

Pixel-Space Diffusion via Observation Operators

Shaojie Guo, Lichen Ma, Haoyang Tong +8

Pixel-space diffusion models directly model image distributions but remain difficult to optimize. Recent methods alleviate this challenge through target reparameterization, while s…

cs.CV2026

PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster

Xiaoan Liu, Lichen Ma, Zipeng Guo +12

Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs. However, most existing methods either target end-to-end po…

cs.CV2026

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

Xiaoan Liu, Lichen Ma, Zipeng Guo +13

Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing metho…

cs.CV2025

UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space

Yong Liu, Jinshan Pan, Yinchuan Li +4

Diffusion models have shown great potential in generating realistic image detail. However, adapting these models to video super-resolution (VSR) remains challenging due to their in…