From the 2 of 5 linked papers with an AI index.
5 papers
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
Zihan Deng, Chuanzhi Xu, Huiqi Liang +3
The paper introduces SciFigQual-Bench, a benchmark dataset that evaluates the quality of scientific figures within full manuscript context across five dimensions, and presents a cr…
SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence
Chuanzhi Xu, Zihan Deng, Huiqi Liang +4
The paper introduces SciFigAlign, a multimodal model that scores scientific figures by aligning visual content with manuscript evidence, using fine‑tuned CLIP and SciBERT to predic…
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
Zelin Zhang, Kedi Li, Huiqi Liang +3
Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, existing approaches often rely on car…
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
Chuanzhi Xu, Huiqi Liang, Bang Shi +7
Long video generation requires high-fidelity synthesis, coherent narrative structure, and user control over extended time spans. Existing text-to-video methods often rely on a sing…
RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature
Hanzheng Li, Xi Fang, Yixuan Li +9
The integration of Multimodal Large Language Models (MLLMs) into chemistry promises to revolutionize scientific discovery, yet their ability to comprehend the dense, graphical lang…