2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2024
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
Bo Tong, Bokai Lai, Yiyi Zhou +5
Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent effor…
cs.CV2024★ 2 cited
Deep Instruction Tuning for Segment Anything Model
Xiaorui Huang, Gen Luo, Chaoyang Zhu +4
Recently, Segment Anything Model (SAM) has become a research hotspot in the fields of multimedia and computer vision, which exhibits powerful yet versatile capabilities on various…