collaborators

5 papers

cs.CV2026

One Model, Two Worlds: Bidirectional Sonar-Optical Translation

Shengji Jin, Trung Tien Dong, Ahmed Lamidi +3

Translating between imaging sonar and optical cameras is valuable for underwater perception, but supporting both directions with separate models duplicates storage and computation.…

cs.RO2026

AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision

Trung Tien Dong, Shengji Jin, Chen Chen +2

Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigation depends on understanding the surrounding free and occupied space. Bi…

cs.CV2026

BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation

Satvik Praveen, Shengji Jin, Ahmed Lamidi +2

Multi-organ ultrasound segmentation remains challenging when anatomically adjacent structures must be delineated jointly, as localized boundary errors can persist even when Dice sc…

cs.CV2026

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms

Shengji Jin, Yuanhao Zou, Victor Zhu +2

While Multimodal Large Language Models (MLLMs) have advanced Video Temporal Grounding (VTG), existing methods often couple output paradigms with different backbones, datasets, and…

cs.CV2025

A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering

Yuanhao Zou, Shengji Jin, Andong Deng +3

Effectively applying Vision-Language Models (VLMs) to Video Question Answering (VideoQA) hinges on selecting a concise yet comprehensive set of frames, as processing entire videos…