4 papers
Beyond Single-Modal Analytics: A Framework for Integrating Heterogeneous LLM-Based Query Systems for Multi-Modal Data
Ruyu Li, Tinghui Zhang, Haodi Ma +2
With the increasing use of multi-modal data, semantic query has become more and more demanded in data management systems, which is an important way to access and analyze multi-moda…
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
Haodi Ma, Vyom Pathak, Daisy Zhe Wang
Video Question Answering (VQA) requires models to reason over spatial, temporal, and causal cues in videos. Recent vision language models (VLMs) achieve strong results but often re…
Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts
Haodi Ma, Dzmitry Kasinets, Daisy Zhe Wang
Multimodal knowledge graph completion (MMKGC) aims to predict missing links in multimodal knowledge graphs (MMKGs) by leveraging information from various modalities alongside struc…
LaPuda: LLM-Enabled Policy-Based Query Optimizer for Multi-modal Data
Yifan Wang, Haodi Ma, Daisy Zhe Wang
Large language model (LLM) has marked a pivotal moment in the field of machine learning and deep learning. Recently its capability for query planning has been investigated, includi…