2 papers
cs.MM2025
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
Jinyuan Li, Ziyan Li, Han Li +4
Grounded Multimodal Named Entity Recognition (GMNER) task aims to identify named entities, entity types and their corresponding visual regions. GMNER task exhibits two challenging…
cs.CL2025
Explainable Multimodal Depression Recognition in Clinical Interviews via PHQ-Aligned Symptom Summarization
Wenjie Zheng, Qiming Xie, Zengzhi Wang +7
Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues.…