4 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CV2024
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models
Mukai Li, Lei Li, Shansan Gong +1
Visual Language Models (VLMs) demonstrate impressive capabilities in processing multimodal inputs, yet applications such as visual agents, which require handling multiple images an…
cs.CL2023★ 2 cited
Can Language Models Understand Physical Concepts?
Lei Li, Jingjing Xu, Qingxiu Dong +4
Language models~(LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite.…
cs.CV2023★ 4 cited
NIKI: Neural Inverse Kinematics with Invertible Neural Networks for 3D Human Pose and Shape Estimation
Jiefeng Li, Siyuan Bian, Qi Liu +3
With the progress of 3D human pose and shape estimation, state-of-the-art methods can either be robust to occlusions or obtain pixel-aligned accuracy in non-occlusion cases. Howeve…