1 citations · 1 across the 2 of their papers we have counts for
6 papers · 1 filter
ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation
Ahmed Abul Hasanaath, Bicheng Xu, Mir Rayat Imtiaz Hossain +2
Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-gra…
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh +6
Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI a…
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
Bicheng Xu, Qi Yan, Renjie Liao +2
We introduce a framework for joint grounded scene graph - image generation, a challenging task involving high-dimensional, multi-modal structured data. To effectively model this co…
Consistent Multiple Sequence Decoding
Bicheng Xu, Leonid Sigal
Sequence decoding is one of the core components of most visual-lingual models. However, typical neural decoders when faced with decoding multiple, possibly correlated, sequences of…
Watch, Listen and Tell: Multi-modal Weakly Supervised Dense Event Captioning
Tanzila Rahman, Bicheng Xu, Leonid Sigal
Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from lang…
Time Perception Machine: Temporal Point Processes for the When, Where and What of Activity Prediction
Yatao Zhong, Bicheng Xu, Guang-Tong Zhou +2
Numerous powerful point process models have been developed to understand temporal patterns in sequential data from fields such as health-care, electronic commerce, social networks,…