3 papers
cs.CV2026
ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation
Ahmed Abul Hasanaath, Bicheng Xu, Mir Rayat Imtiaz Hossain +2
Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-gra…
cs.CV2026
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh +6
Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI a…
cs.CV2025
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
Bicheng Xu, Qi Yan, Renjie Liao +2
We introduce a framework for joint grounded scene graph - image generation, a challenging task involving high-dimensional, multi-modal structured data. To effectively model this co…