7 papers
DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction
Cuong Tran Van, Trong-Thang Pham, Ngoc-Son Nguyen +2
Sparse-view Cone-Beam Computed Tomography reconstruction from limited X-ray projections remains a challenging problem in medical imaging due to the inherent undersampling of fine-g…
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Khoa Vo, Taisei Hanyu, Yuki Ikebe +8
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action…
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
Trong Thang Pham, Hien Nguyen, Ngan Le
Current multimodal large language models (MLLMs) cannot effectively utilize eye-gaze information for video understanding, even when gaze cues are supplied via visual overlays or te…
PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
Trong-Thang Pham, Loc Nguyen, Anh Nguyen +3
Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training data. Re-prompting rerolls the entire gen…
TolerantECG: A Foundation Model for Imperfect Electrocardiogram
Huynh Dang Nguyen, Trong-Thang Pham, Ngan Le +1
The electrocardiogram (ECG) is an essential and effective tool for diagnosing heart diseases. However, its effectiveness can be compromised by noise or unavailability of one or mor…
GazeSearch: Radiology Findings Search Benchmark
Trong Thang Pham, Tien-Phat Nguyen, Yuki Ikebe +5
Medical eye-tracking data is an important information source for understanding how radiologists visually interpret medical images. This information not only improves the accuracy o…