5 citations · 8 across the 3 of their papers we have counts for
3 papers
cs.CL2023★ 5 cited
Tree-of-Mixed-Thought: Combining Fast and Slow Thinking for Multi-hop Visual Reasoning
Pengbo Hu, Ji Qi, Xingyu Li +5
There emerges a promising trend of using large language models (LLMs) to generate code-like plans for complex inference tasks such as visual reasoning. This paradigm, known as LLM-…
cs.CV2023★ 3 cited
Boosting Multi-modal Model Performance with Adaptive Gradient Modulation
Hong Li, Xingyu Li, Pengbo Hu +3
While the field of multi-modal learning keeps growing fast, the deficiency of the standard joint training paradigm has become clear through recent studies. They attribute the sub-o…
cs.LG2022
SHAPE: An Unified Approach to Evaluate the Contribution and Cooperation of Individual Modalities
Pengbo Hu, Xingyu Li, Yi Zhou
As deep learning advances, there is an ever-growing demand for models capable of synthesizing information from multi-modal resources to address the complex tasks raised from real-l…