6 papers
3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
Zhenyu Zhao, Nanshan Jia, Jihyeon Je +7
Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automated judge depends on the…
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff +4
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizin…
Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Wanlong Fang, Tianle Zhang, Wen Tao +1
Understanding how multimodal large language models use different modalities is important for reliable reasoning. We employ Partial Information Decomposition (PID) as a decision-lev…
How Creative Are Large Language Models in Generating Molecules?
Wen Tao, Yiwei Wang, Peng Zhou +6
Molecule generation requires satisfying multiple chemical and biological constraints while searching a large and structured chemical space. This makes it a non-binary problem, wher…
Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
Tianle Zhang, Wanlong Fang, Jonathan Woo +3
The remarkable performance of Large Language Models (LLMs) can be enhanced with test-time computation, which relies on external tools and even other deep learning models. However,…
To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
Wanlong Fang, Tianle Zhang, Alvin Chan
Multimodal learning often relies on aligning representations across modalities to enable effective information integration, an approach traditionally assumed to be universally bene…