Publications (31)
OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
Xinyu Zhang, Boxuan Zhang, Yuchen Wan +5
While Large Language Models (LLMs) demonstrate remarkable reasoning, complex optimization tasks remain challenging, requiring domain knowledge and robust implementation. However, e…
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
Zhan'ao Yao, Liang Yin, Zhihao Gao +9
Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model…
Personalized Attraction Enhanced Sponsored Search with Multi-task Learning
Wei Zhao, Boxuan Zhang, Beidou Wang +6
We study a novel problem of sponsored search (SS) for E-Commerce platforms: how we can attract query users to click product advertisements (ads) by presenting them features of prod…
Ternary Spiking Neural Networks Enhanced by Complemented Neurons and Membrane Potential Aggregation
Boxuan Zhang, Jiaxin Wang, Zhen Xu +1
Spiking Neural Networks (SNNs) are promising energy-efficient models and powerful framworks of modeling neuron dynamics. However, existing binary spiking neurons exhibit limited bi…
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5
Dongrui Liu, Yi Yu, Jie Zhang +18
To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, Frontier AI Risk Management Framework in Practice presents a comp…
MatMind: A Structure-Activity Knowledge-Driven Generative Foundation Model for Materials Science
Zhan'ao Yao, Boxuan Zhang, Jingyuan Shu +10
Progress in AI-driven crystal materials science has so far been carried by narrow architectures purpose-built for individual tasks -- graph neural networks for property prediction,…
ACE-BERT: Adversarial Cross-modal Enhanced BERT for E-commerce Retrieval
Boxuan Zhang, Chao Wei, Yan Jin +1
Nowadays on E-commerce platforms, products are presented to the customers with multiple modalities. These multiple modalities are significant for a retrieval system while providing…
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
Boxuan Zhang, Jianing Zhu, Zeru Shi +2
LLM-based multi-agent systems are increasingly deployed on long-horizon tasks, but a single decisive error is often accepted by downstream agents and cascades into trajectory-level…
Shifting Uncertainty to Critical Moments: Towards Reliable Uncertainty Quantification for VLA Model
Yanchuan Tang, Taowen Wang, Yuefei Chen +3
Vision-Language-Action (VLA) models enable general-purpose robotic policies by mapping visual observations and language instructions to low-level actions, but they often lack relia…
Approximating full conformal prediction: distribution free guarantees via the tournament correction
Aabesh Bhattacharyya, Boxuan Zhang, Rina Foygel Barber
Conformal prediction is a framework for providing prediction intervals with distribution-free validity, guaranteeing predictive coverage for data drawn from any distribution. Its t…
Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers
Zicong He, Boxuan Zhang, Lu Cheng
Large language models (LLMs) are known to hallucinate, a phenomenon often linked to creativity. While previous research has primarily explored this connection through theoretical o…
What Shapes a Creative Machine Mind? Comprehensively Benchmarking Creativity in Foundation Models
Zicong He, Boxuan Zhang, Weihao Liu +2
The meteoric rise of foundation models (FMs) has expanded their capabilities far beyond conventional tasks. Creativity, long regarded as a hallmark of human intelligence and a driv…
Rendering Stable Features Improves Sampling-Based Localisation with Neural Radiance Fields
Boxuan Zhang, Lindsay Kleeman, Michael Burke
Neural radiance fields (NeRFs) are a powerful tool for implicit scene representations, allowing for differentiable rendering and the ability to make predictions about unseen viewpo…
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
Jiaqian Li, Yanshu Li, Boxuan Zhang +2
LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long before they surface in the fin…
Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving
Xinyu Zhang, Yuchen Wan, Boxuan Zhang +4
Large Language Models (LLMs) often struggle with structural ambiguity in optimization problems, where a single problem admits multiple related but conflicting modeling paradigms, h…
DyMoDreamer: World Modeling with Dynamic Modulation
Boxuan Zhang, Runqing Wang, Wei Xiao +5
A critical bottleneck in deep reinforcement learning (DRL) is sample inefficiency, as training high-performance agents often demands extensive environmental interactions. Model-bas…
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
Boxuan Zhang, Yi Yu, Jiaxuan Guo +1
The prevalent deployment of Large Language Model agents such as OpenClaw unlocks potential in real-world applications, while amplifying safety concerns. Among these concerns, the s…
SPECTRA: Context-Conditioned Spectral Movement Primitives for Robot Skill Generalization
Boxuan Zhang, Sheng Liu, Chenlin Ming +1
The paper introduces Spectral Movement Primitives, a frequency‑domain approach that learns robot manipulation skills from demonstrations using low‑frequency Fourier coefficients an…
Boosting Semi-Supervised Object Detection in Remote Sensing Images With Active Teaching
Boxuan Zhang, Zengmao Wang, Bo Du
The lack of object-level annotations poses a significant challenge for object detection in remote sensing images (RSIs). To address this issue, active learning (AL) and semi-superv…
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Fei Tang, Zhiqiong Lu, Boxuan Zhang +4
GUI agents drive applications through their visual interfaces instead of programmatic APIs, interacting with arbitrary software via taps, swipes, and keystrokes, reaching a long ta…
What If the Input is Expanded in OOD Detection?
Boxuan Zhang, Jianing Zhu, Zengmao Wang +3
Out-of-distribution (OOD) detection aims to identify OOD inputs from unknown classes, which is important for the reliable deployment of machine learning models in the open world. V…
Data Augmentation for High-Fidelity Generation of CAR-T/NK Immunological Synapse Images
Xiang Zhang, Boxuan Zhang, Alireza Naghizadeh +4
Chimeric antigen receptor (CAR)-T and NK cell immunotherapies have transformed cancer treatment, and recent studies suggest that the quality of the CAR-T/NK cell immunological syna…
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
Minghao Guo, Qingyue Jiao, Zeru Shi +14
Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later reasoning. In prior work, many…
Differentiable Geometric Indexing for End-to-End Generative Retrieval
Xujing Wang, Yufeng Chen, Boxuan Zhang +7
Generative Retrieval (GR) has emerged as a promising paradigm to unify indexing and search within a single probabilistic framework. However, existing approaches suffer from two int…
Mixture-of-World Models: Scaling Multi-Task Reinforcement Learning with Modular Latent Dynamics
Boxuan Zhang, Weipu Zhang, Zhaohan Feng +4
A fundamental challenge in multi-task reinforcement learning (MTRL) is achieving sample efficiency in visual domains where tasks exhibit substantial heterogeneity in both observati…
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
Boxuan Zhang, Ruqi Zhang
Large language models (LLMs) excel in many tasks but struggle to accurately quantify uncertainty in their generated responses. This limitation makes it challenging to detect misinf…
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
Shanghai AI Lab, :, Xiaoyang Chen +35
To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier…
Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts
Boxuan Zhang, Jianing Zhu, Qifan Wang +2
Recent generative models can produce images that appear highly realistic, raising challenges in distinguishing real and AI-generated images. Yet existing detectors based on pre-tra…
Stochastic Trajectory Optimization for Robotic Skill Acquisition From a Suboptimal Demonstration
Chenlin Ming, Zitong Wang, Boxuan Zhang +3
Learning from Demonstration (LfD) has emerged as a crucial method for robots to acquire new skills. However, when given suboptimal task trajectory demonstrations with shape charact…
Temporal Regularization Training: Unleashing the Potential of Spiking Neural Networks
Boxuan Zhang, Zhen Xu, Kuan Tao
Spiking Neural Networks (SNNs) have received widespread attention due to their event-driven and low-power characteristics, making them particularly effective for processing neuromo…
S1-MatAgent: A planner driven multi-agent system for material discovery
Xinrui Wang, Chengbo Li, Boxuan Zhang +5
The discovery of high-performance materials is crucial for technological advancement. Inverse design using multi-agent systems (MAS) shows great potential for new material discover…