3 papers
cs.CV2025
Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
Thanh-Dat Truong, Huu-Thien Tran, Tran Thai Son +2
Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some f…
cs.CV2025
BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
Huu-Thien Tran, Thanh-Dat Truong, Khoa Luu
Large vision-language models have become widely adopted to advance in various domains. However, developing a trustworthy system with minimal interpretable characteristics of large-…
cs.CV2025
MEX: Memory-efficient Approach to Referring Multi-Object Tracking
Huu-Thien Tran, Phuoc-Sang Pham, Thai-Son Tran +1
Referring Multi-Object Tracking (RMOT) is a relatively new concept that has rapidly gained traction as a promising research direction at the intersection of computer vision and nat…