3 citations · 4 across the 9 of their papers we have counts for
7 papers · 1 filter
Dynamic Training-Free Fusion of Subject and Style LoRAs
Qinglong Cao, Yuntian Chen, Chao Ma +1
Recent studies have explored the combination of multiple LoRAs to simultaneously generate user-specified subjects and styles. However, most existing approaches fuse LoRA weights us…
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
Wenxi Li, Yuchen Guo, Jilai Zheng +4
Recent years have seen an increase in the use of gigapixel-level image and video capture systems and benchmarks with high-resolution wide (HRW) shots. However, unlike close-up shot…
Cross-View Consistency Regularisation for Knowledge Distillation
Weijia Zhang, Dongnan Liu, Weidong Cai +1
Knowledge distillation (KD) is an established paradigm for transferring privileged knowledge from a cumbersome model to a lightweight and efficient one. In recent years, logit-base…
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
Bohan Li, Xin Jin, Jianan Wang +8
Recent diffusion models have demonstrated remarkable performance in both 3D scene generation and perception tasks. Nevertheless, existing methods typically separate these two proce…
Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial Frame
Qinglong Cao, Xirui Li, Ding Wang +3
Video diffusion models have achieved impressive results in natural scene generation, yet they struggle to generalize to scientific phenomena such as fluid simulations and meteorolo…
Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics
Junyi Cao, Shanyan Guan, Yanhao Ge +3
While humans effortlessly discern intrinsic dynamics and adapt to new scenarios, modern AI systems often struggle. Current methods for visual grounding of dynamics either use pure…