most citedReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization

2 citations · 2 across the 7 of their papers we have counts for

collaborators

13 papers

cs.CV2025

MasHeNe: A Benchmark for Head and Neck CT Mass Segmentation using Window-Enhanced Mamba with Frequency-Domain Integration

Thao Thi Phuong Dao, Tan-Cong Nguyen, Nguyen Chi Thanh +7

Head and neck masses are space-occupying lesions that can compress the airway and esophagus and may affect nerves and blood vessels. Available public datasets primarily focus on ma…

cs.CV2025

GenKOL: Modular Generative AI Framework For Scalable Virtual KOL Generation

Tan-Hiep To, Duy-Khang Nguyen, Tam V. Nguyen +2

Key Opinion Leader (KOL) play a crucial role in modern marketing by shaping consumer perceptions and enhancing brand credibility. However, collaborating with human KOLs often invol…

cs.CV20252 cited

ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization

Thinh-Phuc Nguyen, Thanh-Hai Nguyen, Gia-Huy Dinh +3

Image captioning systems often produce generic descriptions that fail to capture event-level semantics which are crucial for applications like news reporting and digital archiving.…

cs.CV2025

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran +1

Event-based image retrieval from free-form captions presents a significant challenge: models must understand not only visual features but also latent event semantics, context, and…

cs.CV2025

TaleForge: Interactive Multimodal System for Personalized Story Creation

Minh-Loi Nguyen, Quang-Khai Le, Tam V. Nguyen +2

Storytelling is a deeply personal and creative process, yet existing methods often treat users as passive consumers, offering generic plots with limited personalization. This under…

cs.CV2025

SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran +1

Retrieving 3D objects in complex indoor environments using only a masked 2D image and a natural language description presents significant challenges. The ROOMELSA challenge limits…