From the 1 of 9 linked papers with an AI index.
9 papers
MemeMind: Reference-Guided Trace Construction for Offline Context Optimization
Run Yang, Weihang Wang, Boheng Sheng +7
Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollouts on an adaptation set, bu…
MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
Weihang Wang, Kainan Tu, Jielei Zhang +9
The paper presents MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes that evaluates how large vision‑language models handle cultural and background knowledge, an…
Large Language Model as Token Compressor and Decompressor
Wenbing Li, Yiran Wang, Zikai Song +4
In this paper, we study whether an off-the-shelf LLM can be adapted into a discrete, variable-length token compressor and decompressor for long-context processing. To this end, we…
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
Yu Xie, Jielei Zhang, Pengyu Chen +5
Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to…
Efficient Causal Structure Learning via Modular Subgraph Integration
Haixiang Sun, Pengchao Tian, Zihan Zhou +3
Learning causal structures from observational data remains a fundamental yet computationally intensive task, particularly in high-dimensional settings where existing methods face c…
Improving VQA Reliability: A Dual-Assessment Approach with Self-Reflection and Cross-Model Verification
Xixian Wu, Yang Ou, Pengchao Tian +4
Vision-language models (VLMs) have demonstrated significant potential in Visual Question Answering (VQA). However, the susceptibility of VLMs to hallucinations can lead to overconf…