2 papers
q-bio.BM2026
STELLA: A Multimodal LLM for Protein Functional Annotation via Unified Sequence-Structure Encoding
Hongwang Xiao, Wenjun Lin, Xi Chen +7
Understanding the intricate interplay among sequence, structure, and function remains a fundamental challenge in proteomics. The sequence-structure-function paradigm posits that bi…
cs.CV2025
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
Chelsi Jain, Yiran Wu, Yifan Zeng +5
Document Visual Question Answering (DocVQA) is a practical yet challenging task, which is to ask questions based on documents while referring to multiple pages and different modali…