14 papers
PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
Satoshi Hashimoto, Yanan Wang, Hitoshi Nishimura +1
Deploying video anomaly detection (VAD) in the real world is often constrained by the scarcity, privacy, and cost of collecting real abnormal footage. We propose PA-VAD, a novel ps…
OlfactProfile: Profile-Conditioned Odor Prediction from Audiovisual Content
Zhengyu Lou, Bosheng Qin, Yanan Wang +3
Automated video-odor matching predicts scents aligned with audiovisual content for scent-enhanced media. Existing methods usually treat odor labels as determined only by scene cont…
Qwen-Image-2.0 Technical Report
Bing Zhao, Chenfei Wu, Deqing Li +72
We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite rece…
DAPE: Dynamic Non-uniform Alignment and Progressive Detail Enhancement Techniques for Improving the Performance of Efficient Visual Language Models
Mengyuan Tian, Qiyan Zhao, Yanan Wang +1
In recent years, pre-trained visual-linguistic models have demonstrated tremendous potential, becoming a crucial foundational framework for numerous downstream tasks. However, the…
Logics-Parsing-Omni Technical Report
Xin An, Jingyi Cai, Xiangyang Chen +22
Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This fram…
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation
Motonari Kambara, Koki Seno, Tomoya Kaichi +2
We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only…