2 papers
cs.CL2026
Knowing When to Trust Images: Reliability-Aware Multi-modal Entity Alignment
Chenxiao Li, Yunhe Feng, Dongfang Liu +3
The visual modality, i.e., images, plays a key role in multi-modal entity alignment (MMEA). Existing approaches often directly fuse the image with other modalities to align differe…
cs.CV2026
LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation
Zan Wang, Yunhe Feng, Dong Nie +4
Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models…