activity
20242026
collaborators

7 papers

cs.AI2026

Ostrakon-VL: Towards Domain-Expert MLLM for Food-Service and Retail Stores

Zhiyong Shen, Gongpeng Zhao, Jun Zhou +10

Multimodal Large Language Models (MLLMs) have recently achieved substantial progress in general-purpose perception and reasoning. Nevertheless, their deployment in Food-Service and…

cs.CV2025

DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion

Haoteng Li, Zhao Yang, Zezhong Qian +5

Accurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding…

cs.CV2025

DualDiff+: Dual-Branch Diffusion for High-Fidelity Video Generation with Reward Guidance

Zhao Yang, Zezhong Qian, Xiaofan Li +5

Accurate and high-fidelity driving scene reconstruction demands the effective utilization of comprehensive scene information as conditional inputs. Existing methods predominantly r…

cs.CV2024

AUD-TGN: Advancing Action Unit Detection with Temporal Convolution and GPT-2 in Wild Audiovisual Contexts

Jun Yu, Zerui Zhang, Zhihong Wei +6

Leveraging the synergy of both audio data and visual data is essential for understanding human emotions and behaviors, especially in in-the-wild setting. Traditional methods for in…

cs.CV2024

Multimodal Fusion Method with Spatiotemporal Sequences and Relationship Learning for Valence-Arousal Estimation

Jun Yu, Gongpeng Zhao, Yongqi Wang +7

This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio seg…

cs.CV2024

Exploring Facial Expression Recognition through Semi-Supervised Pretraining and Temporal Modeling

Jun Yu, Zhihong Wei, Zhongpeng Cai +6

Facial Expression Recognition (FER) plays a crucial role in computer vision and finds extensive applications across various fields. This paper aims to present our approach for the…