Publications (16)
Distilling Physical Priors into Streaming World Models
Liangliang Zhao, Junying Wang, Danni Yang +5
Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical c…
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
Qi Chen, Changli Wu, Jiayi Ji +3
3D Referring Expression Segmentation (3D-RES) aims to segment point cloud scenes based on a given expression. However, existing 3D-RES approaches face two major challenges: feature…
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
Guanzhou Chen, Erfei Cui, Changyao Tian +6
Instruction-based image editing has emerged as a key capability for unified multimodal models (UMMs), yet constructing large-scale, diverse, and high-quality editing datasets witho…
TokensGen: Harnessing Condensed Tokens for Long Video Generation
Wenqi Ouyang, Zeqi Xiao, Danni Yang +5
Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often…
Selection of strain and fitting schemes for calculating higher-order elastic constants
Mingqing Liao, Yong Liu, Fei Zhou +6
Criteria of selecting strain and fitting schemes are proposed for the calculation of higher-order elastic constants more efficiently, robustly and accurately. As demonstrated by th…
RealDPO: Real or Not Real, that is the Preference
Guo Cheng, Danni Yang, Ziqi Huang +3
Video generative models have recently achieved notable advancements in synthesis quality. However, generating complex motions remains a critical challenge, as existing models often…
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Yicheng Zou, Dongsheng Zhu, Lin Zhu +174
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…
Asynchronous Federated Learning with Incentive Mechanism Based on Contract Theory
Danni Yang, Yun Ji, Zhoubin Kou +2
To address the challenges posed by the heterogeneity inherent in federated learning (FL) and to attract high-quality clients, various incentive mechanisms have been employed. Howev…
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
Zhenxin Lei, Zhangwei Gao, Changyao Tian +12
Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains.…
Revisiting the third-order elastic constants of diamond: the higher-order effect
Mingqing Liao, Yong Liu, Yi Wang +7
In this letter, we study the higher-order effect on the third-order elastic constants (TOECs) of diamond using longitudinal stress-uniaxial strain (LSUS) approach based on density…
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
Danni Yang, Ruohan Dong, Jiayi Ji +4
Recently, diffusion models have increasingly demonstrated their capabilities in vision understanding. By leveraging prompt-based learning to construct sentences, these models have…
Nanoscale Bandgap Tuning across an Inhomogeneous Ferroelectric Interface
Jing Wang, Houbing Huang, Wangqiang He +10
We report nanoscale bandgap engineering via a local strain across the inhomogeneous ferroelectric interface, which is controlled by the visible-light-excited probe voltage. Switcha…
Decentralized Dynamic Cooperation of Personalized Models for Federated Continual Learning
Danni Yang, Zhikang Chen, Sen Cui +6
Federated continual learning (FCL) has garnered increasing attention for its ability to support distributed computation in environments with evolving data distributions. However, t…
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
Danni Yang, Jiayi Ji, Yiwei Ma +4
In this paper, we introduce SemiRES, a semi-supervised framework that effectively leverages a combination of labeled and unlabeled data to perform RES. A significant hurdle in appl…
Semi-Supervised Panoptic Narrative Grounding
Danni Yang, Jiayi Ji, Xiaoshuai Sun +4
Despite considerable progress, the advancement of Panoptic Narrative Grounding (PNG) remains hindered by costly annotations. In this paper, we introduce a novel Semi-Supervised Pan…
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
Changyao Tian, Danni Yang, Guanzhou Chen +26
Unified multimodal models (UMMs) that integrate understanding, reasoning, generation, and editing face inherent trade-offs between maintaining strong semantic comprehension and acq…