most citedTem-adapter: Adapting Image-Text Pretraining for Video Question Answer

3 citations · 4 across the 7 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning

Sungjin Park, Xiao Liu, Yeyun Gong +1

Despite recent advances in large language models, open-source models often struggle to consistently perform well on complex reasoning tasks. Existing ensemble methods, whether appl…

hep-ex2024

Search for lepton number violating decays of

BESIII Collaboration, M. Ablikim, M. N. Achasov +655

Based on 7.33 fb of collision data collected by the BESIII detector operating at the BEPCII collider at center-of-mass energies from 4.128 to 4.226 GeV, a search fo…

cs.CL2024

Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning

Yiming Huang, Xiao Liu, Yeyun Gong +4

Large language models (LLMs) have shown great potential in complex reasoning tasks, yet their performance is often hampered by the scarcity of high-quality and reasoning-focused tr…

cs.RO2023

Multimodal Learning of Soft Robot Dynamics using Differentiable Filters

Xiao Liu, Yifan Zhou, Shuhei Ikemoto +1

Differentiable Filters, as recursive Bayesian estimators, possess the ability to learn complex dynamics by deriving state transition and measurement models exclusively from data. T…

cs.CV20231 cited

DreamSpace: Dreaming Your Room Space with Text-Driven Panoramic Texture Propagation

Bangbang Yang, Wenqi Dong, Lin Ma +4

Diffusion-based methods have achieved prominent success in generating 2D media. However, accomplishing similar proficiencies for scene-level mesh texturing in 3D spatial applicatio…

cs.CV20233 cited

Tem-adapter: Adapting Image-Text Pretraining for Video Question Answer

Guangyi Chen, Xiao Liu, Guangrun Wang +4

Video-language pre-trained models have shown remarkable success in guiding video question-answering (VideoQA) tasks. However, due to the length of video sequences, training large-s…