From the 1 of 4 linked papers with an AI index.
4 papers
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
Yuqi Zhang, Cheng Chen, Yuyu Guo +6
Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions eit…
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval
Wenjie Yang, Hang Yu, Yuyu Guo +1
The paper introduces SOLAR, a self‑supervised two‑stage framework for symmetric multimodal‑to‑multimodal retrieval that learns intersection masks from large unlabeled image‑text pa…
OpAgent: Operator Agent for Web Navigation
Yuyu Guo, Wenjie Yang, Siyuan Yang +12
To fulfill user instructions, autonomous web agents must contend with the inherent complexity and volatile nature of real-world websites. Conventional paradigms predominantly rely…
LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection
Jian Wu, Hang Yu, Bingchang Liu +4
Adapting large language models (LLMs) to specific domains often faces a critical bottleneck: the scarcity of high-quality, human-curated data. While large volumes of unchecked data…