8 papers
Artemis: Structured Visual Reasoning for Perception Policy Learning
Wei Tang, Yanpeng Sun, Shan Zhang +6
Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indica…
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
Junming Liu, Yuqi Li, Yifei Sun +4
Vision-Language Models (VLMs) have made striking progress, yet their spatial reasoning remains fragile: models that answer an original input correctly can still fail under paired t…
PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularization
Yao Ni, Shan Zhang, Piotr Koniusz
Parameter-Efficient Fine-Tuning (PEFT) effectively adapts pre-trained transformers to downstream tasks. However, the optimization of tasks performance often comes at the cost of ge…
Amortized Active Generation of Pareto Sets
Daniel M. Steinberg, Asiri Wijesinghe, Rafael Oliveira +3
We introduce active generation of Pareto sets (A-GPS), a new framework for online discrete black-box multi-objective optimization (MOO). A-GPS learns a generative model of the Pare…
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
Yanpeng Sun, Shan Zhang, Wei Tang +5
Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they…
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
Yao Ni, Song Wen, Piotr Koniusz +1
Fine-tuning Stable Diffusion enables subject-driven image synthesis by adapting the model to generate images containing specific subjects. However, existing fine-tuning methods suf…