5 papers
Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
Suyang Xi, Chenxi Yang, Hong Ding +4
Multimodal large language models (MLLMs) often fail in fine-grained visual question answering, producing hallucinations about object identities, positions, and relations because te…
Multimodal Medical Image Binding via Shared Text Embeddings
Yunhao Liu, Suyang Xi, Shiqi Liu +6
Medical image analysis increasingly relies on the integration of multiple imaging modalities to capture complementary anatomical and functional information, enabling more accurate…
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
Hong Ding, Chia Chao Kang, SuYang Xi +3
This research introduces an FPGA-based hardware accelerator to optimize the Singular Value Decomposition (SVD) and Fast Fourier transform (FFT) operations in AI models. The propose…
Enhancing Autonomous Driving Safety through World Model-Based Predictive Navigation and Adaptive Learning Algorithms for 5G Wireless Applications
Hong Ding, Ziming Wang, Yi Ding +3
Addressing the challenge of ensuring safety in ever-changing and unpredictable environments, particularly in the swiftly advancing realm of autonomous driving in today's 5G wireles…
Planning-Aware Diffusion Networks for Enhanced Motion Forecasting in Autonomous Driving
Liu Yunhao, Ding Hong, Zhang Ziming +3
Autonomous driving technology has seen significant advancements, but existing models often fail to fully capture the complexity of multi-agent environments, where interactions betw…