9 papers
FARM: Enhancing Molecular Representations with Functional Group Awareness
Thao Nguyen, Kuan-Hao Huang, Ge Liu +3
We introduce Functional Group-Aware Representations for Small Molecules (FARM), a novel foundation model designed to bridge the gap between SMILES, natural language, and molecular…
From Papers to Panoramas: Building Hierarchies of Scientific Literature at Scale
Muhan Gao, Jash Shah, Weiqi Wang +2
Scientific knowledge is growing rapidly, making it difficult to track progress and high-level conceptual links across broad disciplines. While tools like citation networks and sear…
Visually Descriptive Language Model for Vector Graphics Reasoning
Zhenhailong Wang, Joy Hsu, Xingyao Wang +4
Despite significant advancements, large multimodal models (LMMs) still struggle to bridge the gap between low-level visual perception -- focusing on shapes, sizes, and layouts -- a…
Contrastive Visual Data Augmentation
Yu Zhou, Bingxuan Li, Mohan Tang +6
Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-s…
Eliminating Position Bias of Language Models: A Mechanistic Approach
Ziqi Wang, Hanlin Zhang, Xiner Li +6
Position bias has proven to be a prevalent issue of modern language models (LMs), where the models prioritize content based on its position within the given context. This bias ofte…
ARMADA: Attribute-Based Multimodal Data Augmentation
Xiaomeng Jin, Jeonghwan Kim, Yu Zhou +4
In Multimodal Language Models (MLMs), the cost of manually annotating high-quality image-text pair data for fine-tuning and alignment is extremely high. While existing multimodal d…