3 citations · 3 across the 3 of their papers we have counts for
5 papers · 1 filter
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produce…
FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
Trong Thang Pham, Ngoc-Vuong Ho, Nhat-Tan Bui +8
Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiolo…
NeIn: Telling What You Don't Want
Nhat-Tan Bui, Dinh-Hieu Hoang, Quoc-Huy Trinh +3
Negation is a fundamental linguistic concept used by humans to convey information that they do not desire. Despite this, minimal research has focused on negation within text-guided…
CattleFace-RGBT: RGB-T Cattle Facial Landmark Benchmark
Ethan Coffman, Reagan Clark, Nhat-Tan Bui +5
To address this challenge, we introduce CattleFace-RGBT, a RGB-T Cattle Facial Landmark dataset consisting of 2,300 RGB-T image pairs, a total of 4,600 images. Creating a landmark…
PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-Identification
Quoc-Huy Trinh, Nhat-Tan Bui, Dinh-Hieu Hoang +6
Person Re-Identification (Re-ID) task seeks to enhance the tracking of multiple individuals by surveillance cameras. It supports multimodal tasks, including text-based person retri…