most citedA Survey on Remote Sensing Foundation Models: From Vision to Multimodality

2 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.DC2025

SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices

Will Chow

Large Language Models (LLMs), as the foundational architecture for next-generation interactive AI applications, not only power intelligent dialogue systems but also drive the evolu…

cs.CL2025

GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art

Yiming Lei, Chenkai Zhang, Zeming Liu +5

Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural…

cs.CV2025

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding

Chenkai Zhang, Yiming Lei, Zeming Liu +5

With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities o…

cs.CL2025

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

Yiming Lei, Zhizheng Yang, Zeming Liu +5

Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal mode…

cs.CV20252 cited

A Survey on Remote Sensing Foundation Models: From Vision to Multimodality

Ziyue Huang, Hongxi Yan, Qiqi Zhan +7

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data int…

cs.CL2025

KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus

Xiaoming Shi, Zeming Liu, Yiming Lei +8

Video-based dialogue systems, such as education assistants, have compelling application value, thereby garnering growing interest. However, the current video-based dialogue systems…