7 papers
FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
Jiyoon Pyo, Yuankun Jiao, Dongwon Jung +11
Cartographic reasoning is the skill of interpreting geographic relationships by aligning legends, map scales, compass directions, map texts, and geometries across one or more map i…
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
James Y. Huang, Wenxuan Zhou, Nan Xu +5
The ability of Large Language Models (LLMs) to generate structured outputs that follow arbitrary schemas is crucial to a wide range of downstream tasks that require diverse structu…
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
Qin Liu, Wenxuan Zhou, Nan Xu +5
One critical challenge for large language models (LLMs) for making complex reasoning is their reliance on matching reasoning patterns from training data, instead of proactively sel…
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
Bangzheng Li, Fei Wang, Wenxuan Zhou +5
Vision-Language Models (VLMs) leverage aligned visual encoders to transform images into visual tokens, allowing them to be processed similarly to text by the backbone large languag…
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
Nan Xu, Fei Wang, Sheng Zhang +2
Motivated by in-context learning (ICL) capabilities of Large Language Models (LLMs), multimodal LLMs with additional visual modality are also exhibited with similar ICL abilities w…
Monotonic Paraphrasing Improves Generalization of Language Model Prompting
Qin Liu, Fei Wang, Nan Xu +3
Performance of large language models (LLMs) may vary with different prompts or instructions of even the same task. One commonly recognized factor for this phenomenon is the model's…