From the 1 of 15 linked papers with an AI index.
15 papers
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
Sihan Chen, Jiale Li, Jianghang Lin +1
The paper introduces MoHallBench, a large benchmark designed to evaluate and diagnose motion hallucination—incorrectly inferred human motions—in video large language models, coveri…
Active-SAOOD: Active Sparsely Annotated Oriented Object Detection in Remote Sensing Images
Yu Lin, Jianghang Lin, Kai Ye +2
Reducing the annotation cost of oriented object detection in remote sensing remains a major challenge. Recently, sparse annotation has gained attention for effectively reducing ann…
Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs
Jianghang Lin, Haihua Yang, Deli Yu +6
Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional data curation strategies tha…
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
Yansong Guo, Chaoyang Zhu, Jiayi Ji +2
Video Large Language Models (VideoLLMs) have demonstrated impressive capabilities in video understanding, yet the massive number of input video tokens incurs a significant computat…
SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia
Pengfei Yue, Xingran Zhao, Juntao Chen +6
Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on…
Referring Industrial Anomaly Segmentation
Pengfei Yue, Xiaokang Jiang, Yilin Lu +3
Industrial Anomaly Detection (IAD) is vital for manufacturing, yet traditional methods face significant challenges: unsupervised approaches yield rough localizations requiring manu…