Publications (17)
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
Jiaxin Li, Yuxiang Wu, Zhenkai Zhang +11
Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs. Ho…
STAR-Filter: Efficient Convex Free-Space Approximation via Starshaped Set Filtering in Noisy Environments
Yuwei Wu, Yichen Zhao, Dexter Ong +1
Approximating collision-free space is fundamental to robot planning in complex environments. Convex geometric representations, such as polytopes and ellipsoids, are widely employed…
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
Yu Huang, Zelin Peng, Yichen Zhao +3
Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabil…
KISS - Knowledge Infrastructure for Scientific Simulation: A Scaffolding for Agentic Earth Science
Ziwei Li, Liujun Zhu, Yuchen Liu +5
Process-based simulation models encode decades of scientific understanding across the Earth sciences, yet the communities most exposed to climate risk and resource scarcity are the…
Super4DR: 4D Radar-centric Self-supervised Odometry and Gaussian-based Map Optimization
Zhiheng Li, Weihua Wang, Qiang Shen +2
Conventional SLAM systems using visual or LiDAR data often struggle in poor lighting and severe weather. Although 4D radar is suited for such environments, its sparse and noisy poi…
Asynchronous Nonlinear Sheaf Diffusion for Multi-Agent Coordination
Yichen Zhao, Tyler Hanks, Hans Riess +3
Cellular sheaves and sheaf Laplacians provide a far-reaching generalization of graphs and graph Laplacians, resulting in a wide array of applications ranging from machine learning…
VHDLSuite: Unified Pipeline for LLM VHDL Generation with Data Synthesis and Evaluation
Yijun Shen, Minghao Shao, Yichen Zhao +4
Large Language Models (LLM) have shown impressive capabilities in Register Transfer Level (RTL) code generation, particularly for Verilog. However, evaluating their performance wit…
Predicting Chaotic System Behavior using Machine Learning Techniques
Huaiyuan Rao, Yichen Zhao, Qiang Lai
Recently, machine learning techniques, particularly deep learning, have demonstrated superior performance over traditional time series forecasting methods across various applicatio…
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…
A Distributed Asynchronous Generalized Momentum Algorithm Without Delay Bounds
Ellie Pond, Yichen Zhao, Matthew Hale
Asynchronous optimization algorithms often require delay bounds to prove their convergence, though these bounds can be difficult to obtain in practice. Therefore, we introduce a no…
NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
Zelin Peng, Yichen Zhao, Yu Huang +5
Computer-aided medical image analysis is crucial for disease diagnosis and treatment planning. While vision-language models (VLMs) such as CLIP exhibit strong generalization abilit…
Visual Grounding with Multi-modal Conditional Adaptation
Ruilin Yao, Shengwu Xiong, Yichen Zhao +1
Visual grounding is the task of locating objects specified by natural language expressions. Existing methods extend generic object detection frameworks to tackle this task. They ty…
An overview of domain-specific foundation model: key technologies, applications and challenges
Haolong Chen, Hanzhi Chen, Zijian Zhao +6
The impressive performance of ChatGPT and other foundation-model-based products in human language understanding has prompted both academia and industry to explore how these models…
Parallel in-memory wireless computing
Cong Wang, Gong-Jie Ruan, Zai-Zheng Yang +13
Parallel wireless digital communication with ultralow power consumption is critical for emerging edge technologies such as 5G and Internet of Things. However, the physical separati…
MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation
Yichen Zhao, Zelin Peng, Fenghe Tang +3
Chest X-ray (CXR) reporting follows a region-based clinical workflow in which radiologists inspect anatomical regions and integrate localized findings into a final report. However,…
Integrating Natural Language Processing and Exercise Monitoring for Early Diagnosis of Metabolic Syndrome: A Deep Learning Approach
Yichen Zhao, Yuhua Wang, Xi Cheng +2
Metabolic syndrome (MetS) is a medication condition characterized by abdominal obesity, insulin resistance, hypertension and hyperlipidemia. It increases the risk of majority of ch…
Scalable massively parallel computing using continuous-time data representation in nanoscale crossbar array
Cong Wang, Shi-Jun Liang, Chen-Yu Wang +10
The growth of connected intelligent devices in the Internet of Things has created a pressing need for real-time processing and understanding of large volumes of analogue data. The…