8 papers
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
Zhijiang Tang, Jiaxin Qi, Kaihua Tang +2
Image captioning is a primary task in vision--language research, yet assessing how faithfully a caption preserves image semantics without relying on reference captions remains unse…
In Search of Lost DNA Sequence Pretraining
Zhijiang Tang, Jiaxin Qi, Yan Cui +3
DNA sequence encoding is fundamental to gene function prediction, protein synthesis, and diverse downstream biological tasks. Despite the substantial progress achieved by large-sca…
Scaling Test-Time Robustness of Vision-Language Models via Self-Critical Inference Framework
Kaihua Tang, Jiaxin Qi, Jinli Ou +2
The emergence of Large Language Models (LLMs) has driven rapid progress in multi-modal learning, particularly in the development of Large Vision-Language Models (LVLMs). However, e…
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
Keda Tao, Yuhua Zheng, Jia Xu +13
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily fo…
MobileKernelBench: Can LLMs Write Efficient Kernels for Mobile Devices?
Xingze Zou, Jing Wang, Yuhua Zheng +8
Large language models (LLMs) have demonstrated remarkable capabilities in code generation, yet their potential for generating kernels specifically for mobile devices remains largel…
Spatial Transcriptomics as Images for Large-Scale Pretraining
Yishun Zhu, Jiaxin Qi, Jian Wang +2
Spatial Transcriptomics (ST) profiles thousands of gene expression values at discrete spots with precise coordinates on tissue sections, preserving spatial context essential for cl…