papers

Publications (17)

cs.CV2026

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

Jiaxin Li, Yuxiang Wu, Zhenkai Zhang +11

Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs. Ho…

cs.RO2026

STAR-Filter: Efficient Convex Free-Space Approximation via Starshaped Set Filtering in Noisy Environments

Yuwei Wu, Yichen Zhao, Dexter Ong +1

Approximating collision-free space is fundamental to robot planning in complex environments. Convex geometric representations, such as polytopes and ellipsoids, are widely employed…

cs.CV2025

MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models

Yu Huang, Zelin Peng, Yichen Zhao +3

Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabil…

cs.AI2026

KISS - Knowledge Infrastructure for Scientific Simulation: A Scaffolding for Agentic Earth Science

Ziwei Li, Liujun Zhu, Yuchen Liu +5

Process-based simulation models encode decades of scientific understanding across the Earth sciences, yet the communities most exposed to climate risk and resource scarcity are the…

cs.RO2025

Super4DR: 4D Radar-centric Self-supervised Odometry and Gaussian-based Map Optimization

Zhiheng Li, Weihua Wang, Qiang Shen +2

Conventional SLAM systems using visual or LiDAR data often struggle in poor lighting and severe weather. Although 4D radar is suited for such environments, its sparse and noisy poi…

math.OC2025

Asynchronous Nonlinear Sheaf Diffusion for Multi-Agent Coordination

Yichen Zhao, Tyler Hanks, Hans Riess +3

Cellular sheaves and sheaf Laplacians provide a far-reaching generalization of graphs and graph Laplacians, resulting in a wide array of applications ranging from machine learning…

cs.AR2026

VHDLSuite: Unified Pipeline for LLM VHDL Generation with Data Synthesis and Evaluation

Yijun Shen, Minghao Shao, Yichen Zhao +4

Large Language Models (LLM) have shown impressive capabilities in Register Transfer Level (RTL) code generation, particularly for Verilog. However, evaluating their performance wit…

cs.LG2024

Predicting Chaotic System Behavior using Machine Learning Techniques

Huaiyuan Rao, Yichen Zhao, Qiang Lai

Recently, machine learning techniques, particularly deep learning, have demonstrated superior performance over traditional time series forecasting methods across various applicatio…

cs.CV2025

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang +125

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…

math.OC2025

A Distributed Asynchronous Generalized Momentum Algorithm Without Delay Bounds

Ellie Pond, Yichen Zhao, Matthew Hale

Asynchronous optimization algorithms often require delay bounds to prove their convergence, though these bounds can be difficult to obtain in practice. Therefore, we introduce a no…

cs.CV2026

NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding

Zelin Peng, Yichen Zhao, Yu Huang +5

Computer-aided medical image analysis is crucial for disease diagnosis and treatment planning. While vision-language models (VLMs) such as CLIP exhibit strong generalization abilit…

cs.CV2024

Visual Grounding with Multi-modal Conditional Adaptation

Ruilin Yao, Shengwu Xiong, Yichen Zhao +1

Visual grounding is the task of locating objects specified by natural language expressions. Existing methods extend generic object detection frameworks to tackle this task. They ty…

cs.AI2025

An overview of domain-specific foundation model: key technologies, applications and challenges

Haolong Chen, Hanzhi Chen, Zijian Zhao +6

The impressive performance of ChatGPT and other foundation-model-based products in human language understanding has prompted both academia and industry to explore how these models…

cs.AR2023

Parallel in-memory wireless computing

Cong Wang, Gong-Jie Ruan, Zai-Zheng Yang +13

Parallel wireless digital communication with ultralow power consumption is critical for emerging edge technologies such as 5G and Internet of Things. However, the physical separati…

cs.CV2026

MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation

Yichen Zhao, Zelin Peng, Fenghe Tang +3

Chest X-ray (CXR) reporting follows a region-based clinical workflow in which radiologists inspect anatomical regions and integrate localized findings into a final report. However,…

cs.AI2025

Integrating Natural Language Processing and Exercise Monitoring for Early Diagnosis of Metabolic Syndrome: A Deep Learning Approach

Yichen Zhao, Yuhua Wang, Xi Cheng +2

Metabolic syndrome (MetS) is a medication condition characterized by abdominal obesity, insulin resistance, hypertension and hyperlipidemia. It increases the risk of majority of ch…

cond-mat.mtrl-sci2021

Scalable massively parallel computing using continuous-time data representation in nanoscale crossbar array

Cong Wang, Shi-Jun Liang, Chen-Yu Wang +10

The growth of connected intelligent devices in the Internet of Things has created a pressing need for real-time processing and understanding of large volumes of analogue data. The…