6 papers
EZBlender: Efficient 3D Editing with Plan-and-ReAct Agent
Hao Wang, Wenhui Zhu, Shao Tang +8
As a cornerstone of the modern digital economy, 3D modeling and rendering demand substantial resources and manual effort when scene editing is performed in the traditional manner.…
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
Yanxi Chen, Wenhui Zhu, Xiwen Chen +9
Although Large Audio-Language Models (LALMs) deliver state-of-the-art (SOTA) performance, they frequently suffer from hallucinations, e.g. generating text not grounded in the audio…
nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis
Xin Li, Wenhui Zhu, Xuanzhao Dong +4
Retinal imaging is a critical, non-invasive modality for the early detection and monitoring of ocular and systemic diseases. Deep learning, particularly convolutional neural networ…
VAOT: Vessel-Aware Optimal Transport for Retinal Fundus Enhancement
Xuanzhao Dong, Wenhui Zhu, Yujian Xiong +8
Color fundus photography (CFP) is central to diagnosing and monitoring retinal disease, yet its acquisition variability (e.g., illumination changes) often degrades image quality, w…
Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis
Xiwen Chen, Peijie Qiu, Wenhui Zhu +8
While multiple instance learning (MIL) has shown to be a promising approach for histopathological whole slide image (WSI) analysis, its reliance on permutation invariance significa…
Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences
Xiwen Chen, Peijie Qiu, Wenhui Zhu +5
Since its introduction, the transformer has shifted the development trajectory away from traditional models (e.g., RNN, MLP) in time series forecasting, which is attributed to its…