5 papers
HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
Chen Li, Peiji Yang, Yicheng Zhong +5
Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion…
ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER
Jielong Tang, Shuang Wang, Zhenxing Wang +2
Grounded Multimodal Named Entity Recognition (GMNER) extends traditional NER by jointly detecting textual mentions and grounding them to visual regions. While existing supervised m…
Geography-Aware Large Language Models for Next POI Recommendation
Zhao Liu, Wei Liu, Huajie Zhu +4
The next Point-of-Interest (POI) recommendation task aims to predict users' next destinations based on their historical movement data and plays a key role in location-based service…
Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition
Jielong Tang, Zhenxing Wang, Ziyang Gong +3
Grounded Multimodal Named Entity Recognition (GMNER) is an emerging information extraction (IE) task, aiming to simultaneously extract entity spans, types, and corresponding visual…
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning
Ziqi Qiu, Jianxing Yu, Yufeng Zhang +4
This paper focuses on sarcasm detection, which aims to identify whether given statements convey criticism, mockery, or other negative sentiment opposite to the literal meaning. To…