activity
20142023
most citedLearning Fine-grained Image Similarity with Deep Ranking

95 citations · 97 across the 5 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.CV2024

Instruction Tuning-free Visual Token Complement for Multimodal LLMs

Dongsheng Wang, Jiequan Cui, Miaoge Li +3

As the open community of large language models (LLMs) matures, multimodal LLMs (MLLMs) have promised an elegant bridge between vision and language. However, current research is inh…

cs.CV20243 cited

HICEScore: A Hierarchical Metric for Image Captioning Evaluation

Zequn Zeng, Jianqiao Sun, Hao Zhang +5

Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to eval…

cs.IR2024

All Roads Lead to Rome: Unveiling the Trajectory of Recommender Systems Across the LLM Era

Bo Chen, Xinyi Dai, Huifeng Guo +9

Recommender systems (RS) are vital for managing information overload and delivering personalized content, responding to users' diverse information needs. The emergence of large lan…

cs.CV20242 cited

MeaCap: Memory-Augmented Zero-shot Image Captioning

Zequn Zeng, Yan Xie, Hao Zhang +3

Zero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods…

cs.CR2024

Privacy-Preserving State Estimation in the Presence of Eavesdroppers: A Survey

Xinhao Yan, Guanzhong Zhou, Daniel E. Quevedo +3

Networked systems are increasingly the target of cyberattacks that exploit vulnerabilities within digital communications, embedded hardware, and software. Arguably, the simplest cl…

cs.CV2024

SnapCap: Efficient Snapshot Compressive Video Captioning

Jianqiao Sun, Yudi Su, Hao Zhang +5

Video Captioning (VC) is a challenging multi-modal task since it requires describing the scene in language by understanding various and complex videos. For machines, the traditiona…