From the 1 of 18 linked papers with an AI index.
18 papers
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
Jiahao Huo, Wenjie Qu, Yibo Yan +5
The paper introduces SAMark, a text watermarking method that remains detectable even after paragraph‑level paraphrasing by removing reliance on sentence order and using a hyperboli…
Aligned but Stereotypical? How System Prompts Shape Demographic Bias in LLM-Based Text-to-Image Models
NaHyeon Park, Na Min An, Kunhee Kim +3
Text-to-image (T2I) systems increasingly rely on Large Language Model (LLM)-based text conditioning to interpret and expand user prompts. While this improves prompt understanding a…
MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework
Haowen Xiang, Yibo Yan, Jiahao Huo +4
Multi-vector visual document retrievers achieve strong fine-grained matching by representing each page with multiple vectors from deep Vision-Language Models (VLMs), but this desig…
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
Yibo Yan, Mingdong Ou, Yi Cao +5
Visual Document Retrieval (VDR), which aims to retrieve relevant pages within vast corpora of visually-rich documents, is of significance in current multimodal retrieval applicatio…
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
Yibo Yan, Shen Wang, Jiahao Huo +7
Scientific reasoning, the process through which humans apply logic, evidence, and critical thinking to explore and interpret scientific phenomena, is essential in advancing knowled…
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
Yibo Yan, Shen Wang, Jiahao Huo +13
As the field of Multimodal Large Language Models (MLLMs) continues to evolve, their potential to revolutionize artificial intelligence is particularly promising, especially in addr…