5 citations · 8 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 5 cited
Tree-of-Mixed-Thought: Combining Fast and Slow Thinking for Multi-hop Visual Reasoning
Pengbo Hu, Ji Qi, Xingyu Li +5
There emerges a promising trend of using large language models (LLMs) to generate code-like plans for complex inference tasks such as visual reasoning. This paradigm, known as LLM-…
cs.CV2023★ 3 cited
Boosting Multi-modal Model Performance with Adaptive Gradient Modulation
Hong Li, Xingyu Li, Pengbo Hu +3
While the field of multi-modal learning keeps growing fast, the deficiency of the standard joint training paradigm has become clear through recent studies. They attribute the sub-o…