6 citations · 54 across the 65 of their papers we have counts for
5 papers · 2 filters
DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal Reasoning
Wan Xu, Yuanfan Guo, Kevin Han +2
Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural-language expression space. C…
Improving Complex Moiré Removal with Generative Supervision
Xinyang Gu, Zhilu Zhang, Honglei Xu +3
The availability of high-quality paired data is essential for training learning-based image demoiréing models. However, it remains challenging for existing datasets to encompass th…
GS-RealBlur: A Flexible Data Acquisition Framework for Real-World Image Deblurring
Mingyang Chen, Zhilu Zhang, Honglei Xu +3
High-quality, large-scale paired data is essential for training learning-based image deblurring models. However, synthetic blurry images generally lack realism, while real-world ca…
FedSmoothLoRA: Toward Smoother and Faster Convergence in Federated Low-Rank Adaptation
Zehao Wang, Guanglei Yang, Yihan Zeng +4
Federated fine-tuning of foundation models with Low-Rank Adaptation (LoRA) provides an efficient solution for reducing communication and computation costs while preserving data loc…
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
Zitong Huang, Kaidong Zhang, Yukang Ding +4
Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization (DPO) methods rely on multi-sam…