2 papers
cs.CV2026
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
Meng Zhang, Jinzhong Ning, Xiaolong Wu +2
Grounded Multimodal Named Entity Recognition (GMNER) aims to jointly identify named entity mentions in text, predict their semantic types, and ground each entity to a corresponding…
cs.CL2025
CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
Jinzhong Ning, Paerhati Tulajiang, Yingying Le +4
Speech Relation Extraction (SpeechRE) aims to extract relation triplets directly from speech. However, existing benchmark datasets rely heavily on synthetic data, lacking sufficien…