papers

Publications (16)

cs.CV2026

Distilling Physical Priors into Streaming World Models

Liangliang Zhao, Junying Wang, Danni Yang +5

Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical c…

cs.CV2025

IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation

Qi Chen, Changli Wu, Jiayi Ji +3

3D Referring Expression Segmentation (3D-RES) aims to segment point cloud scenes based on a given expression. However, existing 3D-RES approaches face two major challenges: feature…

cs.CV2026

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework

Guanzhou Chen, Erfei Cui, Changyao Tian +6

Instruction-based image editing has emerged as a key capability for unified multimodal models (UMMs), yet constructing large-scale, diverse, and high-quality editing datasets witho…

cs.CV2025

TokensGen: Harnessing Condensed Tokens for Long Video Generation

Wenqi Ouyang, Zeqi Xiao, Danni Yang +5

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often…

cond-mat.mtrl-sci2020

Selection of strain and fitting schemes for calculating higher-order elastic constants

Mingqing Liao, Yong Liu, Fei Zhou +6

Criteria of selecting strain and fitting schemes are proposed for the calculation of higher-order elastic constants more efficiently, robustly and accurately. As demonstrated by th…

cs.CV2025

RealDPO: Real or Not Real, that is the Preference

Guo Cheng, Danni Yang, Ziqi Huang +3

Video generative models have recently achieved notable advancements in synthesis quality. However, generating complex motions remains a critical challenge, as existing models often…

cs.LG2026

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

Yicheng Zou, Dongsheng Zhu, Lin Zhu +174

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…

cs.LG2023

Asynchronous Federated Learning with Incentive Mechanism Based on Contract Theory

Danni Yang, Yun Ji, Zhoubin Kou +2

To address the challenges posed by the heterogeneity inherent in federated learning (FL) and to attract high-quality clients, various incentive mechanisms have been employed. Howev…

cs.CV2025

MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites

Zhenxin Lei, Zhangwei Gao, Changyao Tian +12

Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains.…

cond-mat.mtrl-sci2021

Revisiting the third-order elastic constants of diamond: the higher-order effect

Mingqing Liao, Yong Liu, Yi Wang +7

In this letter, we study the higher-order effect on the third-order elastic constants (TOECs) of diamond using longitudinal stress-uniaxial strain (LSUS) approach based on density…

cs.CV2024

Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model

Danni Yang, Ruohan Dong, Jiayi Ji +4

Recently, diffusion models have increasingly demonstrated their capabilities in vision understanding. By leveraging prompt-based learning to construct sentences, these models have…

cond-mat.mes-hall2017

Nanoscale Bandgap Tuning across an Inhomogeneous Ferroelectric Interface

Jing Wang, Houbing Huang, Wangqiang He +10

We report nanoscale bandgap engineering via a local strain across the inhomogeneous ferroelectric interface, which is controlled by the visible-light-excited probe voltage. Switcha…

cs.LG2025

Decentralized Dynamic Cooperation of Personalized Models for Federated Continual Learning

Danni Yang, Zhikang Chen, Sen Cui +6

Federated continual learning (FCL) has garnered increasing attention for its ability to support distributed computation in environments with evolving data distributions. However, t…

cs.CV2024

SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation

Danni Yang, Jiayi Ji, Yiwei Ma +4

In this paper, we introduce SemiRES, a semi-supervised framework that effectively leverages a combination of labeled and unlabeled data to perform RES. A significant hurdle in appl…

cs.CV2023

Semi-Supervised Panoptic Narrative Grounding

Danni Yang, Jiayi Ji, Xiaoshuai Sun +4

Despite considerable progress, the advancement of Panoptic Narrative Grounding (PNG) remains hindered by costly annotations. In this paper, we introduce a novel Semi-Supervised Pan…

cs.CV2026

InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing

Changyao Tian, Danni Yang, Guanzhou Chen +26

Unified multimodal models (UMMs) that integrate understanding, reasoning, generation, and editing face inherent trade-offs between maintaining strong semantic comprehension and acq…