works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.CV2026

SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting

Yu Song, Hao Sun, Shiyu Teng +2

Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or require substantial architectur…

cs.LG2026

DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment

Shiyu Teng, Haichen Yu, Jiaqing Liu +6

The paper introduces DynaBridge, a framework that combines acoustic, visual, and textual signals with LLM‑generated DASS‑aware summaries to predict depression, anxiety, and stress…

cs.RO2026

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

Hao Sun, Yu Song, Shiyu Teng +2

VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suff…

cs.AI2026

Dynamic Summary Generation for Interpretable Multimodal Depression Detection

Shiyu Teng, Jiaqing Liu, Hao Sun +6

Depression remains widely underdiagnosed and undertreated because stigma and subjective symptom ratings hinder reliable screening. To address this challenge, we propose a coarse-to…

cs.LG2025

Retrieval-Augmented Multimodal Depression Detection

Ruibo Hou, Shiyu Teng, Jiaqing Liu +4

Multimodal deep learning has shown promise in depression detection by integrating text, audio, and video signals. Recent work leverages sentiment analysis to enhance emotional unde…

cs.CV2025

A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation

Shurong Chai, Rahul Kumar JAIN, Rui Xu +6

Deep learning relies heavily on data augmentation to mitigate limited data, especially in medical imaging. Recent multimodal learning integrates text and images for segmentation, k…