6 papers
Dynamic Training-Free Fusion of Subject and Style LoRAs
Qinglong Cao, Yuntian Chen, Chao Ma +1
Recent studies have explored the combination of multiple LoRAs to simultaneously generate user-specified subjects and styles. However, most existing approaches fuse LoRA weights us…
Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning
Qinglong Cao, Yuntian Chen, Chao Ma +1
Multimodal large language models (MLLMs) have shown remarkable capabilities in multimodal perception and understanding tasks. However, their effectiveness in specialized domains, s…
Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial Frame
Qinglong Cao, Xirui Li, Ding Wang +3
Video diffusion models have achieved impressive results in natural scene generation, yet they struggle to generalize to scientific phenomena such as fluid simulations and meteorolo…
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
Bohan Li, Xin Jin, Jianan Wang +8
Recent diffusion models have demonstrated remarkable performance in both 3D scene generation and perception tasks. Nevertheless, existing methods typically separate these two proce…
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
Wenxi Li, Yuchen Guo, Jilai Zheng +4
Recent years have seen an increase in the use of gigapixel-level image and video capture systems and benchmarks with high-resolution wide (HRW) shots. However, unlike close-up shot…
Cross-View Consistency Regularisation for Knowledge Distillation
Weijia Zhang, Dongnan Liu, Weidong Cai +1
Knowledge distillation (KD) is an established paradigm for transferring privileged knowledge from a cumbersome model to a lightweight and efficient one. In recent years, logit-base…