8 papers
A Systematic Survey of Semantic Role Labeling in the Era of Pretrained Language Models
Huiyao Chen, Meishan Zhang, Jing Li +4
Semantic role labeling (SRL) is a central natural language processing task for understanding predicate-argument structures within texts and enabling downstream applications. Despit…
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
Pengxin Xu, Xincheng Lin, Luping Xiao +5
General scene perception has progressed from object recognition toward open-vocabulary grounding, part localization, and affordance prediction. Yet these capabilities are often rea…
From Pre-trained Models to Large Language Models: A Comprehensive Survey of AI-Driven Psychological Computing
Huiyao Chen, Ruimeng Liu, Yan Luo +4
The intersection of artificial intelligence and psychological science has experienced remarkable growth, with annual publications expanding from 859 papers in 2000 to 29,979 by 202…
VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
Haidong Xu, Guangwei Xu, Zhedong Zheng +7
This paper introduces VimoRAG, a novel video-based retrieval-augmented motion generation framework for motion large language models (LLMs). As motion LLMs face severe out-of-domain…
When Words Smile: Generating Diverse Emotional Facial Expressions from Text
Haidong Xu, Meishan Zhang, Hao Ju +4
Enabling digital humans to express rich emotions has significant applications in dialogue systems, gaming, and other interactive scenarios. While recent advances in talking head sy…
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
Yu Zhao, Hao Fei, Xiangtai Li +6
In the visual spatial understanding (VSU) area, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods f…