3 citations · 3 across the 4 of their papers we have counts for
13 papers
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
Huanlin Gao, Ping Chen, Fuyuan Shi +5
We present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local he…
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
Ruijia Wu, Ping Chen, Fei Shen +8
Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. Howeve…
PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
Xiang liu, Zhaoxiang Liu, Huan Hu +5
Recent advancements in personalized image generation have significantly improved facial identity preservation, particularly in fields such as entertainment and social media. Howeve…
MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
Kaikai Zhao, Zhaoxiang Liu, Peng Wang +7
General-domain large multimodal models (LMMs) have achieved significant advances in various image-text tasks. However, their performance in the Intelligent Traffic Surveillance (IT…
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
Kohou Wang, ZhaoXiang Liu, Lin Bai +5
It is crucial that robots' performance can be improved after deployment, as they are inherently likely to encounter novel scenarios never seen before. This paper presents an innova…
Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
Enbo Zhao, Yi Shen, Shuming Shi +7
Recently, there is a high demand for deploying DeepSeek-R1 and V3 locally, possibly because the official service often suffers from being busy and some organizations have data priv…