5 citations · 8 across the 2 of their papers we have counts for
3 papers
cs.LG2024
ATLAS: Adapter-Based Multi-Modal Continual Learning with a Two-Stage Learning Strategy
Hong Li, Zhiquan Tan, Xingyu Li +1
While vision-and-language models significantly advance in many fields, the challenge of continual learning is unsolved. Parameter-efficient modules like adapters and prompts presen…
cs.CL2023★ 5 cited
Tree-of-Mixed-Thought: Combining Fast and Slow Thinking for Multi-hop Visual Reasoning
Pengbo Hu, Ji Qi, Xingyu Li +5
There emerges a promising trend of using large language models (LLMs) to generate code-like plans for complex inference tasks such as visual reasoning. This paradigm, known as LLM-…
cs.CV2023★ 3 cited
Boosting Multi-modal Model Performance with Adaptive Gradient Modulation
Hong Li, Xingyu Li, Pengbo Hu +3
While the field of multi-modal learning keeps growing fast, the deficiency of the standard joint training paradigm has become clear through recent studies. They attribute the sub-o…