5 papers
SD-GS: Structured Deformable 3D Gaussians for Efficient Dynamic Scene Reconstruction
Wei Yao, Shuzhao Xie, Letian Li +5
Current 4D Gaussian frameworks for dynamic scene reconstruction deliver impressive visual fidelity and rendering speed, however, the inherent trade-off between storage costs and th…
Multimodal Pragmatic Jailbreak on Text-to-image Models
Tong Liu, Zhixin Lai, Jiawen Wang +6
Diffusion models have recently achieved remarkable advancements in terms of image quality and fidelity to textual prompts. Concurrently, the safety of such generative models has be…
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
Xupeng Chen, Zhixin Lai, Kangrui Ruan +3
Artificial intelligence has made significant strides in medical visual question answering (Med-VQA), yet prevalent studies often interpret images holistically, overlooking the visu…
Visual Large Language Models for Generalized and Specialized Applications
Yifan Li, Zhixin Lai, Wentao Bao +7
Visual-language models (VLM) have emerged as a powerful tool for learning a unified embedding space for vision and language. Inspired by large language models, which have demonstra…
Confidence Trigger Detection: Accelerating Real-time Tracking-by-detection Systems
Zhicheng Ding, Zhixin Lai, Siyang Li +3
Real-time object tracking necessitates a delicate balance between speed and accuracy, a challenge exacerbated by the computational demands of deep learning methods. In this paper,…