activity
20242026
collaborators

7 papers

cs.CV2026

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

Kai Liu, Wei Li, Lai Chen +8

This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer…

cs.CL2025

Structure-aware Domain Knowledge Injection for Large Language Models

Kai Liu, Ze Chen, Zhihang Fu +6

This paper introduces a pioneering methodology, termed StructTuning, to efficiently transform foundation Large Language Models (LLMs) into domain specialists. It significantly redu…

cs.CV2024

ESOD: Efficient Small Object Detection on High-Resolution Images

Kai Liu, Zhihang Fu, Sheng Jin +5

Enlarging input images is a straightforward and effective approach to promote small object detection. However, simple image enlargement is significantly expensive on both computati…

cs.CV2024

Category-Extensible Out-of-Distribution Detection via Hierarchical Context Descriptions

Kai Liu, Zhihang Fu, Chao Chen +5

The key to OOD detection has two aspects: generalized feature representation and precise category description. Recently, vision-language models such as CLIP provide significant adv…

cs.CL2024

Enhancing LLM's Cognition via Structurization

Kai Liu, Zhihang Fu, Chao Chen +6

When reading long-form text, human cognition is complex and structurized. While large language models (LLMs) process input contexts through a causal and sequential perspective, thi…

cs.CV2024

Rethinking Out-of-Distribution Detection on Imbalanced Data Distribution

Kai Liu, Zhihang Fu, Sheng Jin +6

Detecting and rejecting unknown out-of-distribution (OOD) samples is critical for deployed neural networks to void unreliable predictions. In real-world scenarios, however, the eff…