1 paper
Mingfei Gao, Rui Tian, Haiming Gang +11
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and…