6 papers
Location-Aware Pretraining for Medical Difference Visual Question Answering
Denis Musinguzi, Caren Han, Prasenjit Mitra
Differential medical VQA models compare multiple images to identify clinically meaningful changes and rely on vision encoders to capture fine-grained visual differences that reflec…
MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning
Eileen Wang, Hiba Arnaout, Dhita Pratama +6
We present MMCOMET, the first multimodal commonsense knowledge graph (MMKG) that integrates physical, social, and eventive knowledge. MMCOMET extends the ATOMIC2020 knowledge graph…
FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance
Hyunsuk Chung, Soyeon Caren Han, Caren Han +5
Multimodal foundation models integrate heterogeneous signals across modalities, yet it remains unclear whether their predictions can be controlled by explicitly modulating reliance…
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
Yan Li, Soyeon Caren Han, Yue Dai +1
Transformer-based models have achieved remarkable success in various Natural Language Processing (NLP) tasks, yet their ability to handle long documents is constrained by computati…
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
Eileen Wang, Caren Han, Josiah Poon
Video Paragraph Captioning (VPC) aims to generate paragraph captions that summarises key events within a video. Despite recent advancements, challenges persist, notably in effectiv…
Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer
Zichen Geng, Caren Han, Zeeshan Hayder +3
Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally exp…