3 papers
cs.AI2025
ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
Pengze Li, Jiaqi Liu, Junchi Yu +5
Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these output…
cs.AI2025
Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
Jiaqi Liu, Songning Lai, Pengze Li +12
Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to u…
cs.CV2024
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
Peng Xia, Siwei Han, Shi Qiu +9
Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal…