activity
20242026
collaborators

12 papers

cs.CV2026

VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection

Huilin Deng, Hongchen Luo, Wei Zhai +2

Zero-shot anomaly detection (ZSAD) recognizes and localizes anomalies in previously unseen objects by establishing feature mapping between textual prompts and inspection images, de…

cs.CV2026

Gloria: Consistent Character Video Generation via Content Anchors

Yuhang Yang, Fan Zhang, Huaijin Pi +5

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Ex…

cs.CV2025

MatE: Material Extraction from Single-Image via Geometric Prior

Zeyu Zhang, Wei Zhai, Jian Yang +1

The creation of high-fidelity, physically-based rendering (PBR) materials remains a bottleneck in many graphics pipelines, typically requiring specialized equipment and expert-driv…

cs.CV2025

Event Stream Filtering via Probability Flux Estimation

Jinze Chen, Wei Zhai, Yang Cao +2

Event cameras asynchronously capture brightness changes with microsecond latency, offering exceptional temporal precision but suffering from severe noise and signal inconsistencies…

cs.CV2025

MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking

Han Han, Wei Zhai, Yang Cao +2

Tracking Any Point (TAP) plays a crucial role in motion analysis. Video-based approaches rely on iterative local matching for tracking, but they assume linear motion during the bli…

cs.CV2025

MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

Jian Yang, Dacheng Yin, Yizhou Zhou +4

Recent advancements in multi-modal large language models have propelled the development of joint probabilistic models capable of both image understanding and generation. However, w…