Publications (8)
Tread lightly interpreting group differences in genetic risk
Nicole Kleman, Meng Lin, Christopher R. Gignoux +1
Observed differences in mean phenotypic values across human groups have attracted renewed interest with the rise of large-scale genomic studies and polygenic risk prediction. Howev…
WildSVG: Towards Reliable SVG Generation Under Real-Word Conditions
Marco Terral, Haotian Zhang, Tianyang Zhang +8
We introduce the task of SVG extraction, which consists in translating specific visual inputs from an image into scalable vector graphics. Existing multimodal models achieve strong…
Text Smoothing: Enhance Various Data Augmentation Methods on Text Classification Tasks
Xing Wu, Chaochen Gao, Meng Lin +3
Before entering the neural network, a token is generally converted to the corresponding one-hot representation, which is a discrete distribution of the vocabulary. Smoothed represe…
ActiShade: Activating Overshadowed Knowledge to Guide Multi-Hop Reasoning in Large Language Models
Huipeng Ma, Luan Zhang, Dandan Song +10
In multi-hop reasoning, multi-round retrieval-augmented generation (RAG) methods typically rely on LLM-generated content as the retrieval query. However, these approaches are inher…
ConTextual Masked Auto-Encoder for Dense Passage Retrieval
Xing Wu, Guangyuan Ma, Meng Lin +3
Dense passage retrieval aims to retrieve the relevant passages of a query from a large corpus based on dense representations (i.e., vectors) of the query and the passages. Recent s…
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
Juan Rodriguez, Haotian Zhang, Abhay Puri +13
We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding.…
Imbalanced Sentiment Classification Enhanced with Discourse Marker
Tao Zhang, Xing Wu, Meng Lin +2
Imbalanced data commonly exists in real world, espacially in sentiment-related corpus, making it difficult to train a classifier to distinguish latent sentiment in text data. We ob…
CoT-MAE v2: Contextual Masked Auto-Encoder with Multi-view Modeling for Passage Retrieval
Xing Wu, Guangyuan Ma, Peng Wang +4
Growing techniques have been emerging to improve the performance of passage retrieval. As an effective representation bottleneck pretraining technique, the contextual masked auto-e…