4 papers
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
Zheda Mai, Arpita Chowdhury, Zihe Wang +5
The rise of vision foundation models (VFMs) calls for systematic evaluation. A common approach pairs VFMs with large language models (LLMs) as general-purpose heads, followed by ev…
ExpressMind: A Multimodal Pretrained Large Language Model for Expressway Operation
Zihe Wang, Yihuan Wang, Haiyang Yu. Zhiyong Cui +4
The current expressway operation relies on rule-based and isolated models, which limits the ability to jointly analyze knowledge across different systems. Meanwhile, Large Language…
Static Segmentation by Tracking: A Label-Efficient Approach for Fine-Grained Specimen Image Segmentation
Zhenyang Feng, Zihe Wang, Jianyang Gu +22
We study image segmentation in the biological domain, particularly trait segmentation from specimen images (e.g., butterfly wing stripes, beetle elytra). This fine-grained task is…
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
Jihyung Kil, Zheda Mai, Justin Lee +6
The ability to compare objects, scenes, or situations is crucial for effective decision-making and problem-solving in everyday life. For instance, comparing the freshness of apples…