10 papers
LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
Jiajun Cheng, Subarna Tripathi, Sainan Liu +2
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders…
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
Shubhashis Roy Dipta, Tz-Ying Wu, Subarna Tripathi
We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus on factual accuracy. Unlike existing…
Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
Sainan Liu, Tz-Ying Wu, Hector A Valdez +1
We present Search2Motion, a training-free framework for object-level motion editing in image-to-video generation. Unlike prior methods requiring trajectories, bounding boxes, masks…
TrajPred: Trajectory-Conditioned Joint Embedding Prediction for Surgical Instrument-Tissue Interaction Recognition in Vision-Language Models
Jiajun Cheng, Xiaofan Yu, Subarna Tripathi +2
Recognizing instruments' interactions with tissues is essential for building context-aware AI assistants in robotic surgery. Vision-language models (VLMs) have opened a new avenue…
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
Julia Lee Romero, Kyle Min, Subarna Tripathi +1
Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We…
Harnessing Object Grounding for Time-Sensitive Video Understanding
Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi
We propose to improve the time-sensitive video understanding (TSV) capability of video large language models (Video-LLMs) with grounded objects (GO). We hypothesize that TSV tasks…