4 papers
CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning
Zhenxuan Fan, Jie Cao, Yang Dai +5
Chain-of-thought (CoT) prompting improves LLM reasoning but incurs high latency and memory cost due to verbose traces, motivating CoT compression with preserved correctness. Existi…
Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
Yang Dai, Jianxiang An, Tianwei Lin +6
Multimodal Large Language Models (MLLMs) have achieved success across various domains. However, their applicability tends to degrade when confronted with different types of data in…
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
Haoyu Zheng, Qifan Yu, Binghe Yu +5
Diffusion models have achieved remarkable progress in image and video stylization. However, most existing methods focus on single-style transfer, while video stylization involving…
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
Haoyu Zheng, Wenqiao Zhang, Zheqi Lv +8
Diffusion-based text-to-image (T2I) models have demonstrated remarkable results in global video editing tasks. However, their focus is primarily on global video modifications, and…