collaborators

7 papers

cs.CV2026

OrganLens: Organ-Specific Representation Learning for CT Foundation Models

Zhixuan Ge, Anqi Li, Sadeer Al-Kindi +2

A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a sep…

cs.IR2026

UniNote: A Unified Embedding Model for Multimodal Representation and Ranking

Jinghan Zhao, Wenwei Jin, Anqi Li +5

Item-to-Item (I2I) retrieval is a fundamental part of modern content platforms, supporting critical industrial workflows from recommendation engines to content auditing. While mult…

cs.CV2026

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

Xinrui Shi, Kai Liu, Ziqing Zhang +3

Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations…

cs.LG2026

MedMIX: Modality-Internal Expert Fusion for Multimodal Medical Diagnosis

Seungik Cho, Anqi Li, Wei Qiu

Multimodal clinical prediction faces three challenges: multiple foundation models (FMs) with complementary strengths per modality, pervasive missing modalities at training and test…

cs.CV2026

Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention

Junhao Du, Jialong Xue, Anqi Li +2

Video large language models (Video-LLMs) face high computational costs due to large volumes of visual tokens. Existing token compression methods typically adopt a two-stage spatiot…

cs.CL2026

Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling

Anqi Li, Wenwei Jin, Jintao Tong +3

Social platforms have revolutionized information sharing, but also accelerated the dissemination of harmful and policy-violating content. To ensure safety and compliance at scale,…