36 citations · 58 across the 5 of their papers we have counts for
5 papers
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
Jielin Qiu, Andrea Madotto, Zhaojiang Lin +7
Vision-extended LLMs have made significant strides in Visual Question Answering (VQA). Despite these advancements, VLLMs still encounter substantial difficulties in handling querie…
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin +10
We present Any-Modality Augmented Language Model (AnyMAL), a unified model that reasons over diverse input modality signals (i.e. text, image, video, audio, IMU motion sensor), and…
Joint Audio-Text Model for Expressive Speech-Driven 3D Facial Animation
Yingruo Fan, Zhaojiang Lin, Jun Saito +2
Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely be…
FaceFormer: Speech-Driven 3D Facial Animation with Transformers
Yingruo Fan, Zhaojiang Lin, Jun Saito +2
Speech-driven 3D facial animation is challenging due to the complex geometry of human faces and the limited availability of 3D audio-visual data. Prior works typically focus on lea…
Language Models as Few-Shot Learner for Task-Oriented Dialogue Systems
Andrea Madotto, Zihan Liu, Zhaojiang Lin +1
Task-oriented dialogue systems use four connected modules, namely, Natural Language Understanding (NLU), a Dialogue State Tracking (DST), Dialogue Policy (DP) and Natural Language…