43 citations · 44 across the 5 of their papers we have counts for
5 papers
Towards Unified Facial Action Unit Recognition Framework by Large Language Models
Guohong Hu, Xing Lan, Hanyu Jiang +2
Facial Action Units (AUs) are of great significance in the realm of affective computing. In this paper, we propose AU-LLaVA, the first unified AU recognition framework based on the…
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
Xing Lan, Jian Xue, Ji Qi +3
Facial expression recognition (FER) is a critical task in multimedia with significant implications across various domains. However, analyzing the causes of facial expressions is es…
Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation
Sara Papi, Peidong Wang, Junkun Chen +4
The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering transl…
FoodSAM: Any Food Segmentation
Xing Lan, Jiayi Lyu, Hanyu Jiang +4
In this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generat…
Pre-training End-to-end ASR Models with Augmented Speech Samples Queried by Text
Eric Sun, Jinyu Li, Jian Xue +1
In end-to-end automatic speech recognition system, one of the difficulties for language expansion is the limited paired speech and text training data. In this paper, we propose a n…