papers

Publications (35)

cs.CL2024

Synergy-of-Thoughts: Eliciting Efficient Reasoning in Hybrid Language Models

Yu Shang, Yu Li, Fengli Xu +1

Large language models (LLMs) have shown impressive emergent abilities in a wide range of tasks, but the associated expensive API cost greatly limits the real application. Previous…

cs.CV2026

WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

Yu Shang, Zhuohang Li, Yiding Ma +18

While world models have emerged as a cornerstone of embodied intelligence by enabling agents to reason about environmental dynamics through action-conditioned prediction, their eva…

quant-ph2025

Distributed Quantum Neural Networks on Distributed Photonic Quantum Computing

Kuan-Cheng Chen, Chen-Yu Liu, Yu Shang +2

We introduce a distributed quantum-classical framework that synergizes photonic quantum neural networks (QNNs) with matrix-product-state (MPS) mapping to achieve parameter-efficien…

cs.NE2024

Towards Biologically Plausible Computing: A Comprehensive Comparison

Changze Lv, Yufei Gu, Zhengkang Guo +16

Backpropagation is a cornerstone algorithm in training neural networks for supervised learning, which uses a gradient descent method to update network weights by minimizing the dis…

cs.CV2024

UrbanWorld: An Urban World Model for 3D City Generation

Yu Shang, Yuming Lin, Yu Zheng +6

Cities, as the essential environment of human life, encompass diverse physical elements such as buildings, roads and vegetation, which continuously interact with dynamic entities l…

gr-qc2006

On Newman-Penrose constants of stationary space-times

Xiaoning Wu, Yu Shang

We consider the general asymptotic expression of stationary space-time. Using Killing equation, we reduce the dynamical freedom of Einstein equation to the in-going gravitational w…

cs.RO2025

KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models

Sibo Li, Qianyue Hao, Yu Shang +1

Robotic world models are a promising paradigm for forecasting future environment states, yet their inference speed and the physical plausibility of generated trajectories remain cr…

cs.CL2024

RNG: Reducing Multi-level Noise and Multi-grained Semantic Gap for Joint Multimodal Aspect-Sentiment Analysis

Yaxin Liu, Yan Zhou, Ziming Li +4

As an important multimodal sentiment analysis task, Joint Multimodal Aspect-Sentiment Analysis (JMASA), aiming to jointly extract aspect terms and their associated sentiment polari…

cs.IR2021

Genetic Meta-Structure Search for Recommendation on Heterogeneous Information Network

Zhenyu Han, Fengli Xu, Jinghan Shi +4

In the past decade, the heterogeneous information network (HIN) has become an important methodology for modern recommender systems. To fully leverage its power, manually designed n…

cs.CL2025

Understanding World or Predicting Future? A Comprehensive Survey of World Models

Jingtao Ding, Yunke Zhang, Yu Shang +12

The concept of world models has garnered significant attention due to advancements in multimodal large language models such as GPT-4 and video generation models such as Sora, which…

cond-mat.mtrl-sci2026

Stoichiometric cluster learning for few-shot property prediction of multi-ionic integrated energetic materials

Ming-Yu Guo, Wei-Jia Zou, Yu Shang +1

Multi-ionic materials pose a distinct representational challenge in machine learning-driven materials design. Different from single-molecule or composition-based materials, their p…

cs.RO2025

AirScape: An Aerial Generative World Model with Motion Controllability

Baining Zhao, Rongze Tang, Mingyuan Jia +9

How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general s…

cs.CL2025

AgentSquare: Automatic LLM Agent Search in Modular Design Space

Yu Shang, Yu Li, Keyu Zhao +4

Recent advancements in Large Language Models (LLMs) have led to a rapid growth of agentic systems capable of handling a wide range of complex tasks. However, current research large…

gr-qc2009

The search for black hole binaries using a genetic algorithm

Antoine Petiteau, Yu Shang, Stanislav Babak

In this work we use genetic algorithm to search for the gravitational wave signal from the inspiralling massive Black Hole binaries in the simulated LISA data. We consider a single…

gr-qc2012

EMRI data analysis with a phenomenological waveform

Yan Wang, Yu Shang, Stanislav Babak

Extreme mass ratio inspirals (EMRIs) (capture and inspiral of a compact stellar mass object into a Massive Black Hole (MBH)) are among the most interesting objects for the gravitat…

cs.CV2025

Kaleidoscopic Background Attack: Disrupting Pose Estimation with Multi-Fold Radial Symmetry Textures

Xinlong Ding, Hongwei Yu, Jiawei Li +5

Camera pose estimation is a fundamental computer vision task that is essential for applications like visual localization and multi-view stereo reconstruction. In the object-centric…

cs.RO2026

Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space

Weichen Zhang, Peizhi Tang, Xin Zeng +12

Unmanned aerial vehicles (UAVs) have emerged as powerful embodied agents. One of the core abilities is autonomous navigation in large-scale three-dimensional environments. Existing…

cs.RO2026

Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control

Jianjie Fang, Yongyan Xu, Ziyou Wang +13

World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, fo…

cs.CV2025

RoboScape: Physics-informed Embodied World Model

Yu Shang, Xin Zhang, Yinzhou Tang +4

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data…

gr-qc2007

Light Cone Structure near Null Infinity of the Kerr Metric

Shan Bai, Zhoujian Cao, Xuefei Gong +3

Motivated by our attempt to understand the question of angular momentum of a relativistic rotating source carried away by gravitational waves, in the asymptotic regime near future…

cs.CV2025

RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale

Shengyuan Wang, Zhiheng Zheng, Yu Shang +6

City-scale 3D generation is of great importance for the development of embodied intelligence and world models. Existing methods, however, face significant challenges regarding qual…

cs.IR2025

AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems

Yu Shang, Peijie Liu, Yuwei Yan +9

The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasonin…

cond-mat.mtrl-sci2025

Nitrogen-Vacancy Engineering for Controlled Phase Transitions in CrN(111) Epitaxial Films

XiaoXu Zhang, Yang Li, Yu Shang +6

The phase transition in CrN epitaxial films is substantially suppressed by epitaxial constraint. Here, we propose that nitrogen (N) vacancies can be taken as a knob to regulate the…

cs.CV2025

LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE

Yu Shang, Lei Jin, Yiding Ma +4

Video-based world models hold significant potential for generating high-quality embodied manipulation data. However, current video generation methods struggle to achieve stable lon…

gr-qc2010

The search for spinning black hole binaries in mock LISA data using a genetic algorithm

Antoine Petiteau, Yu Shang, Stanislav Babak +1

Coalescing massive Black Hole binaries are the strongest and probably the most important gravitational wave sources in the LISA band. The spin and orbital precessions bring complex…

cs.IR2025

AgentSociety Challenge: Designing LLM Agents for User Modeling and Recommendation on Web Platforms

Yuwei Yan, Yu Shang, Qingbin Zeng +9

The AgentSociety Challenge is the first competition in the Web Conference that aims to explore the potential of Large Language Model (LLM) agents in modeling user behavior and enha…

cs.RO2026

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

Yu Shang, Yinzhou Tang, Yiding Ma +22

World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…

gr-qc2010

The Mock LISA Data Challenges: from Challenge 3 to Challenge 4

Stanislav Babak, John G. Baker, Matthew J. Benacquista +27

The Mock LISA Data Challenges are a program to demonstrate LISA data-analysis capabilities and to encourage their development. Each round of challenges consists of one or more data…

cs.MM2025

A Large-scale Dataset with Behavior, Attributes, and Content of Mobile Short-video Platform

Yu Shang, Chen Gao, Nian Li +1

Short-video platforms show an increasing impact on people's daily lives nowadays, with billions of active users spending plenty of time each day. The interactions between users and…

gr-qc2009

The Mock LISA Data Challenges: from Challenge 1B to Challenge 3

Stanislav Babak, John G. Baker, Matthew J. Benacquista +27

The Mock LISA Data Challenges are a programme to demonstrate and encourage the development of LISA data-analysis capabilities, tools and techniques. At the time of this workshop, t…

physics.soc-ph2026

Perceiving exposure segregation with open urban imagery

Yunke Zhang, Ruolong Ma, Xin Zhang +4

Socioeconomic exposure segregation -- the lack of daily interaction between income groups -- erodes social capital and entrenches inequality, yet the specific physical features tha…

cs.RO2026

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

Baining Zhao, Jiacheng Xu, Weicheng Feng +13

Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial V…

cs.RO2026

Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning

Yinzhou Tang, Jingbo Xu, Yu Shang +4

World Action Models (WAMs) offer a promising approach to embodied intelligence, yet existing methods rely heavily on video prediction as action priors and lack adaptive multimodal…

cs.RO2025

RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL

Yinzhou Tang, Yu Shang, Yinuo Chen +8

Achieving generalizable embodied policies remains a key challenge. Traditional policy learning paradigms, including both Imitation Learning (IL) and Reinforcement Learning (RL), st…

cs.CV2026

MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation

Yangcheng Yu, Xin Jin, Yu Shang +4

Embodied action planning is a core challenge in robotics, requiring models to generate precise actions from visual observations and language instructions. While video generation wo…