From the 1 of 8 linked papers with an AI index.
8 papers
SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting
Yu Song, Hao Sun, Shiyu Teng +2
Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or require substantial architectur…
DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
Shiyu Teng, Haichen Yu, Jiaqing Liu +6
The paper introduces DynaBridge, a framework that combines acoustic, visual, and textual signals with LLM‑generated DASS‑aware summaries to predict depression, anxiety, and stress…
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
Hao Sun, Yu Song, Shiyu Teng +2
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suff…
Dynamic Summary Generation for Interpretable Multimodal Depression Detection
Shiyu Teng, Jiaqing Liu, Hao Sun +6
Depression remains widely underdiagnosed and undertreated because stigma and subjective symptom ratings hinder reliable screening. To address this challenge, we propose a coarse-to…
Retrieval-Augmented Multimodal Depression Detection
Ruibo Hou, Shiyu Teng, Jiaqing Liu +4
Multimodal deep learning has shown promise in depression detection by integrating text, audio, and video signals. Recent work leverages sentiment analysis to enhance emotional unde…
A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation
Shurong Chai, Rahul Kumar JAIN, Rui Xu +6
Deep learning relies heavily on data augmentation to mitigate limited data, especially in medical imaging. Recent multimodal learning integrates text and images for segmentation, k…