3 papers
cs.AI2025
Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
Zixian Guo, Ming Liu, Qilong Wang +4
Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end tra…
cs.CV2024
LLM as a Complementary Optimizer to Gradient Descent: A Case Study in Prompt Tuning
Zixian Guo, Ming Liu, Zhilong Ji +3
Mastering a skill generally relies on both hands-on experience from doers and insightful, high-level guidance by mentors. Will this strategy also work well for solving complex non-…
cs.CV2023
Black-Box Tuning of Vision-Language Models with Effective Gradient Approximation
Zixian Guo, Yuxiang Wei, Ming Liu +4
Parameter-efficient fine-tuning (PEFT) methods have provided an effective way for adapting large vision-language models to specific tasks or scenarios. Typically, they learn a very…