5 papers
Hierarchical Conditional Relation Networks for Multimodal Video Question Answering
Thao Minh Le, Vuong Le, Svetha Venkatesh +1
Video QA challenges modelers in multiple fronts. Modeling video necessitates building not only spatio-temporal models for the dynamic visual channel but also multimodal structures…
GEFA: Early Fusion Approach in Drug-Target Affinity Prediction
Tri Minh Nguyen, Thin Nguyen, Thao Minh Le +1
Predicting the interaction between a compound and a target is crucial for rapid drug repurposing. Deep learning has been successfully applied in drug-target affinity (DTA) problem.…
Dynamic Language Binding in Relational Visual Reasoning
Thao Minh Le, Vuong Le, Svetha Venkatesh +1
We present Language-binding Object Graph Network, the first neural reasoning method with dynamic relational structures across both visual and textual domains with applications in v…
Hierarchical Conditional Relation Networks for Video Question Answering
Thao Minh Le, Vuong Le, Svetha Venkatesh +1
Video question answering (VideoQA) is challenging as it requires modeling capacity to distill dynamic visual artifacts and distant relations and to associate them with linguistic c…
Neural Reasoning, Fast and Slow, for Video Question Answering
Thao Minh Le, Vuong Le, Svetha Venkatesh +1
What does it take to design a machine that learns to answer natural questions about a video? A Video QA system must simultaneously understand language, represent visual content ove…