Publications (15)
IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing
Tao Hu, Jiaxin Ai, Licheng Wen +12
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
Mengqi He, Xinyu Tian, Xin Shen +6
SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
Shu Zou, Xinyu Tian, Qinyu Zhao +2
IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD
Nianchen Deng, Jiaxin Ai, Tao Hu +10
The paper introduces IndustryForge-27B, a multimodal foundation model fine‑tuned on diverse industrial CAD data to understand drawings, generate parametric modeling scripts, and co…
Asymptotic order of the geometric mean error for self-affine measures on Bedford-McMullen carpets
Sanguo Zhu, Shu Zou
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
Jiaxin Ai, Tao Hu, Xuemeng Yang +11
ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
Xinyu Tian, Shu Zou, Zhaoyuan Yang +1
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
Xinyu Tian, Shu Zou, Zhaoyuan Yang +1
Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition
Xinyu Tian, Shu Zou, Zhaoyuan Yang +2
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
Xinyu Tian, Shu Zou, Zhaoyuan Yang +5
MemHarness: Memory Is Reconstructed, Not Replayed
Rong Wu, Daocheng Fu, Licheng Wen +10
The paper introduces MemHarness, a framework that lets large language model agents reconstruct and adapt retrieved past experiences to the current context instead of replaying them…
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
Mengqi He, Xinyu Tian, Xin Shen +4
ASSEMCAD: Production-Ready CAD Assembly Generation from Natural Language
Yurui Dong, Shu Zou, Siqi Li +7
All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
Xinyu Tian, Shu Zou, Zhaoyuan Yang +3
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
Shu Zou, Xinyu Tian, Lukas Wesemann +3