collaborators

6 papers

cs.CV2026

Rethinking the Adaptation of Vision Foundation Models for Efficient Cell Segmentation

Qing Xu, Xiangjian He, Wenting Duan +2

Cell segmentation is critical for computational pathology and biomedical discovery. While recent Vision Foundation Models (VFMs) have demonstrated remarkable universal feature repr…

cs.CL2026

Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs

Shanshan Wang, Fengying Ye, Hanjia Lyu +6

Previous detection studies have shown that LLMs cannot be effectively used as detectors, but these studies have not addressed modern Chinese poetry. Moreover, no relevant research…

cs.CV2026

Aurora: Unified Video Editing with a Tool-Using Agent

Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao +4

Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and reference images, and one set o…

cs.CV2026

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents

Ziyun Zeng, Hang Hua, Bocheng Zou +3

Recent GUI agents have made substantial progress in visual grounding and action prediction, yet they remain brittle in long-horizon tasks that require maintaining task state across…

cs.LG2026

scHelix: Asymmetric Dual-Stream Integration via Explicit Gene-Level Disentanglement

Xichen Yan, Zelin Zang, Changxi Chi +8

A critical challenge in single-cell RNA sequencing (scRNA-seq) integration is resolving the tension between eliminating batch effects and maintaining biological fidelity. While rec…

cs.CV2025

How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment

Zhen Chen, Qing Xu, Jinlin Wu +7

Foundation models in video generation are demonstrating remarkable capabilities as potential world models for simulating the physical world. However, their application in high-stak…