7 papers
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
Kai Liu, Wei Li, Lai Chen +8
This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer…
Structure-aware Domain Knowledge Injection for Large Language Models
Kai Liu, Ze Chen, Zhihang Fu +6
This paper introduces a pioneering methodology, termed StructTuning, to efficiently transform foundation Large Language Models (LLMs) into domain specialists. It significantly redu…
ESOD: Efficient Small Object Detection on High-Resolution Images
Kai Liu, Zhihang Fu, Sheng Jin +5
Enlarging input images is a straightforward and effective approach to promote small object detection. However, simple image enlargement is significantly expensive on both computati…
Category-Extensible Out-of-Distribution Detection via Hierarchical Context Descriptions
Kai Liu, Zhihang Fu, Chao Chen +5
The key to OOD detection has two aspects: generalized feature representation and precise category description. Recently, vision-language models such as CLIP provide significant adv…
Enhancing LLM's Cognition via Structurization
Kai Liu, Zhihang Fu, Chao Chen +6
When reading long-form text, human cognition is complex and structurized. While large language models (LLMs) process input contexts through a causal and sequential perspective, thi…
Rethinking Out-of-Distribution Detection on Imbalanced Data Distribution
Kai Liu, Zhihang Fu, Sheng Jin +6
Detecting and rejecting unknown out-of-distribution (OOD) samples is critical for deployed neural networks to void unreliable predictions. In real-world scenarios, however, the eff…