8 papers
Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs
Xi Xiao, Chen Liu, Chih-Ting Liao +9
Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text. Despite inheriting strong reason…
Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson +9
Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing…
Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space
Yue Cao, Jianyang Gu, Vyacheslav Kungurtsev +4
Dataset distillation (DD) has proven to reduce training cost while preserving accuracy. While promising, the factors that make one distilled dataset more effective than another rem…
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes
Zihang Fu, Fanxiao Li, Jianyang Gu +5
Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social platforms. However, they stil…
AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network
Yu Hu, Jianyang Gu, Hao Liu +4
Adapting vision-language models to remote sensing imagery remains challenging due to two key factors: limited semantic coverage in textual representations and insufficient adaptabi…
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
Lin Zhao, Xinru Jiang, Xi Xiao +7
Dataset distillation often prioritizes global semantic proximity when creating small surrogate datasets for original large-scale ones. However, object semantics are inherently hier…