5 papers
Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins
Yiqing Shen, Hao Ding, Mathias Unberath
Text-to-video retrieval in operating rooms (OR) is an enabling technology for OR safety, as it allows stakeholders to retrieve and inspect recordings of specific events. However, b…
Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA
Yiqing Shen, Han Zhang, Mathias Unberath
Surgical video question answering requires multi-step reasoning across semantic, spatial, and temporal dimensions. Existing methods architecturally compress videos into discrete to…
FinSphere, a Real-Time Stock Analysis Agent Powered by Instruction-Tuned LLMs and Domain Tools
Shijie Han, Jingshu Zhang, Yiqing Shen +2
Current financial large language models (FinLLMs) struggle with two critical limitations: the absence of objective evaluation metrics to assess the quality of stock analysis report…
Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning
Jinzheng Li, Sibo Ju, Yanzhou Su +2
Existing large language models (LLMs) driven search agents typically rely on prompt engineering to decouple the user queries into search plans, limiting their effectiveness in comp…
An Agent Framework for Real-Time Financial Information Searching with Large Language Models
Jinzheng Li, Jingshu Zhang, Hongguang Li +1
Financial decision-making requires processing vast amounts of real-time information while understanding their complex temporal relationships. While traditional search engines excel…