3 papers
cs.CV2025
Open Ad-hoc Categorization with Contextualized Feature Learning
Zilin Wang, Sangwoo Mo, Stella X. Yu +2
Adaptive categorization of visual scenes is essential for AI agents to handle changing tasks. Unlike fixed common categories for plants or animals, ad-hoc categories are created dy…
cs.LG2025
MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
Qinghua Liu, Sam Heshmati, Zheda Mai +3
Effective analysis of time series data presents significant challenges due to the complex temporal dependencies and cross-channel interactions in multivariate data. Inspired by the…
cs.CV2025
VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels
Xiwei Xuan, Xiaoqi Wang, Wenbin He +4
The advances in multi-modal foundation models (FMs) (e.g., CLIP and LLaVA) have facilitated the auto-labeling of large-scale datasets, enhancing model performance in challenging do…