1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
Right this way: Can VLMs Guide Us to See More to Answer Questions?
Li Liu, Diji Yang, Sijia Zhong +4
In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answ…
cs.CL2023★ 1 cited
Tackling Vision Language Tasks Through Learning Inner Monologues
Diji Yang, Kezhen Chen, Jinmeng Rao +4
Visual language tasks require AI models to comprehend and reason with both visual and textual content. Driven by the power of Large Language Models (LLMs), two prominent methods ha…