3 papers
cs.CV2026
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
Hao Vo, Khoa Vo, Thinh Phan +5
Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performa…
cs.CV2026
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
Trong Thang Pham, Hien Nguyen, Ngan Le
Current multimodal large language models (MLLMs) cannot effectively utilize eye-gaze information for video understanding, even when gaze cues are supplied via visual overlays or te…
cs.CV2026
PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
Trong-Thang Pham, Loc Nguyen, Anh Nguyen +3
Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training data. Re-prompting rerolls the entire gen…