8 papers
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
Zhijiang Tang, Jiaxin Qi, Kaihua Tang +2
Image captioning is a primary task in vision--language research, yet assessing how faithfully a caption preserves image semantics without relying on reference captions remains unse…
Spatial Transcriptomics as Images for Large-Scale Pretraining
Yishun Zhu, Jiaxin Qi, Jian Wang +2
Spatial Transcriptomics (ST) profiles thousands of gene expression values at discrete spots with precise coordinates on tissue sections, preserving spatial context essential for cl…
In Search of Lost DNA Sequence Pretraining
Zhijiang Tang, Jiaxin Qi, Yan Cui +3
DNA sequence encoding is fundamental to gene function prediction, protein synthesis, and diverse downstream biological tasks. Despite the substantial progress achieved by large-sca…
Scaling Test-Time Robustness of Vision-Language Models via Self-Critical Inference Framework
Kaihua Tang, Jiaxin Qi, Jinli Ou +2
The emergence of Large Language Models (LLMs) has driven rapid progress in multi-modal learning, particularly in the development of Large Vision-Language Models (LVLMs). However, e…
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
Keda Tao, Yuhua Zheng, Jia Xu +13
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily fo…
MobileKernelBench: Can LLMs Write Efficient Kernels for Mobile Devices?
Xingze Zou, Jing Wang, Yuhua Zheng +8
Large language models (LLMs) have demonstrated remarkable capabilities in code generation, yet their potential for generating kernels specifically for mobile devices remains largel…