Publications (35)
Crocs: Cross-Technology Clock Synchronization for WiFi and ZigBee
Zihao Yu, Chengkun Jiang, Yuan He +2
Clock synchronization is a key function in embedded wireless systems and networks. This issue is equally important and more challenging in IoT systems nowadays, which often include…
ARCADE: A Real-Time Data System for Hybrid and Continuous Query Processing across Diverse Data Modalities
Jingyi Yang, Songsong Mo, Jiachen Shi +4
The explosive growth of multimodal data - spanning text, image, video, spatial, and relational modalities, coupled with the need for real-time semantic search and retrieval over th…
NTIRE 2024 Quality Assessment of AI-Generated Content Challenge
Xiaohong Liu, Xiongkuo Min, Guangtao Zhai +111
This paper reports on the NTIRE 2024 Quality Assessment of AI-Generated Content Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancemen…
Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation
Diandian Gu, Jing Lin, Gaohong Liu +25
We present Seed3D 2.0, an advanced 3D content generation system built on Seed3D 1.0, with substantial improvements across generation fidelity, simulation-ready capabilities, and ap…
NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results
Xiaohong Liu, Xiongkuo Min, Qiang Hu +92
This paper reports on the NTIRE 2025 XGC Quality Assessment Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) a…
ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection
Zihao Yu, Xiang Li, Jing Zhang
The rapid dissemination of rumors on social media highlights the urgent need for automatic detection methods to safeguard societal trust and stability. While existing multimodal ru…
CAMAL: Optimizing LSM-trees via Active Learning
Weiping Yu, Siqiang Luo, Zihao Yu +1
We use machine learning to optimize LSM-tree structure, aiming to reduce the cost of processing various read/write operations. We introduce a new approach Camal, which boasts the f…
MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
Zihao Yu, Xiu Yuan, Chongjie Zhang
The paper introduces MEMORA, a system that builds and uses a persistent memory of egocentric video experiences to help robots plan long‑term tasks by storing environment, entity, a…
Protected Transverse Electric Waves in Topological Dielectric Waveguides
Rui Zhou, Minglin L. N. Chen, Xingtong Shi +5
Waveguides are fundamental components in communication systems. However, they suffer from reflection and scattering losses at sharp routes or defects. The breakthrough in developin…
GSIM: Accelerating RTL Simulation for Large-Scale Designs
Lu Chen, Dingyi Zhao, Zihao Yu +2
Register Transfer Level (RTL) simulation is widely used in design space exploration, verification, debugging, and preliminary performance evaluation for hardware design. Among vari…
iEDA: An Open-Source Intelligent Physical Implementation Toolkit and Library
Xingquan Li, Simin Tao, Zengrong Huang +53
Open-source EDA shows promising potential in unleashing EDA innovation and lowering the cost of chip design. This paper presents an open-source EDA project, iEDA, aiming for buildi…
Design and optimization of DBSCAN Algorithm based on CUDA
Bingchen Wang, Chenglong Zhang, Lei Song +3
DBSCAN is a very classic algorithm for data clus- tering, which is widely used in many fields. However, with the data scale growing much more bigger than before, the traditional se…
Dynamics of thin film flows on a vertical fibre with vapor absorption
Souradip Chattopadhyay, Zihao Yu, Y. Sungtaek Ju +1
Water vapor capture through free surface flows plays a crucial role in various industrial applications, such as liquid desiccant air conditioning systems, water harvesting, and dew…
Accelerating Text-to-Image Editing via Cache-Enabled Sparse Diffusion Inference
Zihao Yu, Haoyang Li, Fangcheng Fu +2
Due to the recent success of diffusion models, text-to-image generation is becoming increasingly popular and achieves a wide range of applications. Among them, text-to-image editin…
Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
Feng Wang, Zihao Yu
Reinforcement Learning (RL) has recently emerged as a powerful technique for improving image and video generation in Diffusion and Flow Matching models, specifically for enhancing…
Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
Jiashi Feng, Xiu Li, Jing Lin +25
Developing embodied AI agents requires scalable training environments that balance content diversity with physics accuracy. World simulators provide such environments but face dist…
RF-Transformer: A Unified Backscatter Radio Hardware Abstraction
Xiuzhen Guo, Yuan He, Zihao Yu +3
This paper presents RF-Transformer, a unified backscatter radio hardware abstraction that allows a low-power IoT device to directly communicate with heterogeneous wireless receiver…
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
Dasong Li, Sizhuo Ma, Hang Hua +40
This paper presents an overview of the VQualA 2025 Challenge on Engagement Prediction for Short Videos, held in conjunction with ICCV 2025. The challenge focuses on understanding a…
AdaSfM: From Coarse Global to Fine Incremental Adaptive Structure from Motion
Yu Chen, Zihao Yu, Shu Song +3
Despite the impressive results achieved by many existing Structure from Motion (SfM) approaches, there is still a need to improve the robustness, accuracy, and efficiency on large-…
Leggiero: Analog WiFi Backscatter with Payload Transparency
Xin Na, Xiuzhen Guo, Zihao Yu +3
Backscatter is an enabling technology for battery-free sensing in today's Artificial Intelligence of Things (AIOT). Building a backscatter-based sensing system, however, is a daunt…
Video Quality Assessment Based on Swin TransformerV2 and Coarse to Fine Strategy
Zihao Yu, Fengbin Guan, Yiting Lu +2
The objective of non-reference video quality assessment is to evaluate the quality of distorted video without access to reference high-definition references. In this study, we intr…
Linear Codes from Simplicial Complexes over
Hongwei Liu, Zihao Yu
In this article we mainly study linear codes over and their binary subfield codes. We construct linear codes over whose defining sets are the…
DHIL-GT: Scalable Graph Transformer with Decoupled Hierarchy Labeling
Ningyi Liao, Zihao Yu, Siqiang Luo
Graph Transformer (GT) has recently emerged as a promising neural network architecture for learning graph-structured data. However, its global attention mechanism with quadratic co…
Unifews: You Need Fewer Operations for Efficient Graph Neural Networks
Ningyi Liao, Zihao Yu, Ruixiao Zeng +1
Graph Neural Networks (GNNs) have shown promising performance, but at the cost of resource-intensive operations on graph-scale matrices. To reduce computational overhead, previous…
QMamba: On First Exploration of Vision Mamba for Image Quality Assessment
Fengbin Guan, Xin Li, Zihao Yu +2
In this work, we take the first exploration of the recently popular foundation model, i.e., State Space Model/Mamba, in image quality assessment (IQA), aiming at observing and exca…
4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
Yiting Lu, Wei Luo, Peiyan Tu +8
World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct rea…
Cross-Technology Communication for the Internet of Things: A Survey
Yuan He, Xiuzhen Guo, Xiaolong Zheng +5
The ever-developing Internet of Things (IoT) brings the prosperity of wireless sensing and control applications. In many scenarios, different wireless technologies coexist in the s…
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
Zhongkai Yu, Yue Guan, Zihao Yu +6
Large-scale Mixture of Experts (MoE) Large Language Models (LLMs) have recently become the frontier open-weight models, achieving remarkable model capability similar to proprietary…
COSMIC: Generalized Refusal Direction Identification in LLM Activations
Vincent Siu, Nicholas Crispino, Zihao Yu +5
Large Language Models (LLMs) encode behaviors such as refusal within their activation space, yet identifying these behaviors remains a significant challenge. Existing methods often…
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model
Fengbin Guan, Zihao Yu, Yiting Lu +2
Video quality assessment tasks rely heavily on the rich features required for video understanding, such as semantic information, texture, and temporal motion. The existing video fo…
Super-robust telecommunications enabled by topological half-supermodes
Rui Zhou, Xintong Shi, Hai Lin +6
Topological photonics offers transformative potential for robust integrated waveguide devices due to their backscattering-immune properties. However, their integration faces two fu…
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
Shuhao Han, Haotian Fan, Fangyuan Kong +112
This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoratio…
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
Jianhong Tu, Zhuohao Ni, Nicholas Crispino +8
We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base.…
Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation
Jiahao Xu, Peiyuan Wang, Hanzhuo Zhang +7
In robotic manipulation, the tight coupling between grasping and motion planning often obscures the true source of failure, leading to inefficient trial-and-error. To enable effici…
NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results
Xin Li, Kun Yuan, Bingchen Li +110
This paper presents a review for the NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement. The challenge comprises two tracks: (i) Efficient Video Qualit…