activity
20192022
most citedFine-Grained Semantically Aligned Vision-Language Pre-Training

29 citations · 44 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

Auto-Encoding Morph-Tokens for Multimodal LLM

Kaihang Pan, Siliang Tang, Juncheng Li +6

For multimodal LLMs, the synergy of visual comprehension (textual output) and generation (visual output) presents an ongoing challenge. This is due to a conflicting objective: for…

cs.CV202229 cited

Fine-Grained Semantically Aligned Vision-Language Pre-Training

Juncheng Li, Xin He, Longhui Wei +6

Large-scale vision-language pre-training has shown impressive advances in a wide range of downstream tasks. Existing methods mainly model the cross-modal alignment by the similarit…

cs.CV20225 cited

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Juncheng Li, Junlin Xie, Long Qian +6

Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language…

cs.CV2021

Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language Inference

Juncheng Li, Siliang Tang, Linchao Zhu +5

Video-and-Language Inference is a recently proposed task for joint video-and-language understanding. This new task requires a model to draw inference on whether a natural language…

cs.CV2019

Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation

Juncheng Li, Xin Wang, Siliang Tang +4

Visual navigation is a task of training an embodied agent by intelligently navigating to a target object (e.g., television) using only visual observations. A key challenge for curr…

cs.CV2019

Walking with MIND: Mental Imagery eNhanceD Embodied QA

Juncheng Li, Siliang Tang, Fei Wu +1

The EmbodiedQA is a task of training an embodied agent by intelligently navigating in a simulated environment and gathering visual information to answer questions. Existing approac…