most citedTextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

2 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2025

OpenUS: A Fully Open-Source Foundation Model for Ultrasound Image Analysis via Self-Adaptive Masked Contrastive Learning

Xiaoyu Zheng, Xu Chen, Awais Rauf +6

Ultrasound (US) is one of the most widely used medical imaging modalities, thanks to its low cost, portability, real-time feedback, and absence of ionizing radiation. However, US i…

cs.CL2025

DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models

Xiao Zheng

Large Language Model (LLM) hallucination is a significant barrier to their reliable deployment. Current methods like Retrieval-Augmented Generation (RAG) are often reactive. We int…

cs.LG2025

Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching

Yanhao Dong, Yubo Miao, Weinan Li +4

Large Language Models (LLMs) exhibit pronounced memory-bound characteristics during inference due to High Bandwidth Memory (HBM) bandwidth constraints. In this paper, we propose an…

cs.CV2025

XFMamba: Cross-Fusion Mamba for Multi-View Medical Image Classification

Xiaoyu Zheng, Xu Chen, Shaogang Gong +2

Compared to single view medical image classification, using multiple views can significantly enhance predictive accuracy as it can account for the complementarity of each view whil…

cs.CV20242 cited

TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Ya-Qi Yu, Minghui Liao, Jihao Wu +3

Multimodal Large Language Models (MLLMs) have shown impressive results on various multimodal tasks. However, most existing MLLMs are not well suited for document-oriented tasks, wh…