3 papers
cs.CV2025
MOSABench: Multi-Object Sentiment Analysis Benchmark for Evaluating Multimodal Large Language Models Understanding of Complex Image
Shezheng Song, Chengxiang He, Shan Zhao +4
Multimodal large language models (MLLMs) have shown remarkable progress in high-level semantic tasks such as visual question answering, image captioning, and emotion recognition. H…
cs.NI2025
Camel: Energy-Aware LLM Inference on Resource-Constrained Devices
Hao Xu, Long Peng, Shezheng Song +5
Most Large Language Models (LLMs) are currently deployed in the cloud, with users relying on internet connectivity for access. However, this paradigm faces challenges such as netwo…
cs.CL2025
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
Shezheng Song, Xiaopeng Li, Shasha Li +5
We explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilit…