7 papers
PresentAgent: Multimodal Agent for Presentation Video Generation
Jingwei Shi, Zeyu Zhang, Biao Wu +4
We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides…
DC-Scene: Data-Centric Learning for 3D Scene Understanding
Ting Huang, Zeyu Zhang, Ruicheng Zhang +1
3D scene understanding plays a fundamental role in vision applications such as robotics, autonomous driving, and augmented reality. However, advancing learning-based 3D scene under…
MediAug: Exploring Visual Augmentation in Medical Imaging
Xuyin Qi, Zeyu Zhang, Canxuan Gang +4
Data augmentation is essential in medical imaging for improving classification accuracy, lesion detection, and organ segmentation under limited data conditions. However, two signif…
MedConv: Convolutions Beat Transformers on Long-Tailed Bone Density Prediction
Xuyin Qi, Zeyu Zhang, Huazhan Zheng +19
Bone density prediction via CT scans to estimate T-scores is crucial, providing a more precise assessment of bone health compared to traditional methods like X-ray bone density tes…
PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images
Yang Luo, Shiru Wang, Jun Liu +7
Breast cancer survival prediction in computational pathology presents a remarkable challenge due to tumor heterogeneity. For instance, different regions of the same tumor in the pa…
PedDet: Adaptive Spectral Optimization for Multimodal Pedestrian Detection
Rui Zhao, Zeyu Zhang, Yi Xu +6
Pedestrian detection in intelligent transportation systems has made significant progress but faces two critical challenges: (1) insufficient fusion of complementary information bet…