4 papers
LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
Jiajun Cheng, Subarna Tripathi, Sainan Liu +2
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders…
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
Mingxu Tao, Jiawei Hu, Xian Zhou +5
Legal case retrieval remains challenging due to the complexity of legal language and the need for precise lexical alignment between queries and relevant cases. Although dense retri…
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
Xuanzhao Dong, Wenhui Zhu, Xiwen Chen +13
The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to support clinical diagnosis. However,…
TrajPred: Trajectory-Conditioned Joint Embedding Prediction for Surgical Instrument-Tissue Interaction Recognition in Vision-Language Models
Jiajun Cheng, Xiaofan Yu, Subarna Tripathi +2
Recognizing instruments' interactions with tissues is essential for building context-aware AI assistants in robotic surgery. Vision-language models (VLMs) have opened a new avenue…