17 papers
Predicting Signed Distance Functions for Visual Instance Segmentation
Emil Brissman, Joakim Johnander, Michael Felsberg
Visual instance segmentation is a challenging problem and becomes even more difficult if objects of interest varies unconstrained in shape. Some objects are well described by a rec…
One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context
Yingzi Ma, Zichen Zhu, Ming Jiang +1
On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts. Existing teachers…
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
Zijian Wang, Hanqi Li, Ziyue Yang +17
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside…
SSR: Can Simulated Patients Learn to Stigmatize Themselves? Modeling Self-Stigma through Internal Monologue
Kunyao Lan, Bingrui Jin, Zichen Zhu +1
Simulating patients with large language models (LLMs) is a promising tool for mental health training, but existing approaches fail to capture a key clinical reality: self-stigma. P…
IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment
Zichen Zhu, Yuheng Sun, Mingxuan Zhu +11
Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes. Creations by generative models may contain…
Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents
Zhijie Ding, Weinan Hong, Zicheng Zhu +6
Multimodal large language models (MLLMs) have substantially advanced mobile agents, yet proactive mobile assistance remains challenging because agents must decide \emph{when} to in…