3 papers
cs.CV2026
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
Joongwon Chae, Zhenyu Wang, Peiwu Qin
Despite significant advances in vision-language understanding, implementing image segmentation within multimodal architectures remains a fundamental challenge in modern artificial…
q-bio.BM2025
pLDDT-Predictor: High-speed Protein Screening Using Transformer and ESM2
Joongwon Chae, Zhenyu Wang, Ijaz Gul +3
Recent advancements in protein structure prediction, particularly AlphaFold2, have revolutionized structural biology by achieving near-experimental accuracy ($\text{average RMSD} <…
cs.CV2024
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
Joongwon Chae, Zhenyu Wang, Lian Zhang +2
Recent advances in multimodal models have demonstrated impressive capabilities in object recognition and scene understanding. However, these models often struggle with precise spat…