3 papers
cs.CV2025
Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource
Lei Yang, Junshan Jin, Mingyuan Zhang +3
Visual speech recognition is a technique to identify spoken content in silent speech videos, which has raised significant attention in recent years. Advancements in data-driven dee…
cs.CV2025
BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
Xingyu Zheng, Xianglong Liu, Haotong Qin +7
With the advancement of diffusion models (DMs) and the substantially increased computational requirements, quantization emerges as a practical solution to obtain compact and effici…
eess.IV2024
PtychoFormer: A Transformer-based Model for Ptychographic Phase Retrieval
Ryuma Nakahata, Shehtab Zaman, Mingyuan Zhang +2
Ptychography is a computational method of microscopy that recovers high-resolution transmission images of samples from a series of diffraction patterns. While conventional phase re…