3 citations · 6 across the 8 of their papers we have counts for
4 papers · 1 filter
DataProphet: Demystifying Supervision Data Generalization in Multimodal LLMs
Xuan Qi, Luxi He, Dan Roth +1
Conventional wisdom for selecting supervision data for multimodal large language models (MLLMs) is to prioritize datasets that appear similar to the target benchmark, such as text-…
Statutory Construction and Interpretation for Artificial Intelligence
Luxi He, Nimra Nadeem, Michel Liao +4
AI systems are increasingly governed by natural language principles, yet a key challenge arising from reliance on language remains underexplored: interpretive ambiguity. As in lega…
Metadata Conditioning Accelerates Language Model Pre-training
Tianyu Gao, Alexander Wettig, Luxi He +3
The vast diversity of styles, domains, and quality levels present in language model pre-training corpora is essential in developing general model capabilities, but efficiently lear…
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Zirui Wang, Mengzhou Xia, Luxi He +10
Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. Howeve…