collaborators

5 papers

cs.GR2025

SD-GS: Structured Deformable 3D Gaussians for Efficient Dynamic Scene Reconstruction

Wei Yao, Shuzhao Xie, Letian Li +5

Current 4D Gaussian frameworks for dynamic scene reconstruction deliver impressive visual fidelity and rendering speed, however, the inherent trade-off between storage costs and th…

cs.CV2025

Multimodal Pragmatic Jailbreak on Text-to-image Models

Tong Liu, Zhixin Lai, Jiawen Wang +6

Diffusion models have recently achieved remarkable advancements in terms of image quality and fidelity to textual prompts. Concurrently, the safety of such generative models has be…

cs.CV2025

R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest

Xupeng Chen, Zhixin Lai, Kangrui Ruan +3

Artificial intelligence has made significant strides in medical visual question answering (Med-VQA), yet prevalent studies often interpret images holistically, overlooking the visu…

cs.CV2025

Visual Large Language Models for Generalized and Specialized Applications

Yifan Li, Zhixin Lai, Wentao Bao +7

Visual-language models (VLM) have emerged as a powerful tool for learning a unified embedding space for vision and language. Inspired by large language models, which have demonstra…

cs.CV2024

Confidence Trigger Detection: Accelerating Real-time Tracking-by-detection Systems

Zhicheng Ding, Zhixin Lai, Siyang Li +3

Real-time object tracking necessitates a delicate balance between speed and accuracy, a challenge exacerbated by the computational demands of deep learning methods. In this paper,…