3 papers
cs.CV2026
Temporal Aware Pruning for Efficient Diffusion-based Video Generation
Sheng Li, Yang Sui, Junhao Ran +3
Video diffusion models have recently enabled high-quality video generation with ViT-based architectures, but remain computationally intensive because generation requires attention…
cs.IR2024
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
Shi Yu, Chaoyue Tang, Bokai Xu +8
Retrieval-augmented generation (RAG) is an effective technique that enables large language models (LLMs) to utilize external knowledge sources for generation. However, current RAG…
cs.CE2024
Assessing and Enhancing Large Language Models in Rare Disease Question-answering
Guanchu Wang, Junhao Ran, Ruixiang Tang +6
Despite the impressive capabilities of Large Language Models (LLMs) in general medical domains, questions remain about their performance in diagnosing rare diseases. To answer this…