6 papers
E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory
Lin Huang, Chengxiang Huang, Ziang Wang +7
Equivariant Graph Neural Networks (EGNNs) have become a widely used approach for modeling 3D atomistic systems. However, mainstream architectures face critical scalability bottlene…
Hard Negative Sample-Augmented DPO Post-Training for Small Language Models
Haocheng Lu, Minjun Zhu, Henry Yu
Large language models (LLMs) continue to struggle with mathematical reasoning, and common post-training pipelines often reduce each generated solution to a binary outcome: correct…
UBio-MolFM: A Universal Molecular Foundation Model for Bio-Systems
Lin Huang, Arthur Jiang, XiaoLi Liu +8
All-atom molecular simulation serves as a quintessential ``computational microscope'' for understanding the machinery of life, yet it remains fundamentally limited by the trade-off…
Uni-Parser Technical Report
Xi Fang, Haoyi Tao, Shuwen Yang +8
This technical report introduces Uni-Parser, an industrial-grade document parsing engine tailored for scientific literature and patents, delivering high throughput, robust accuracy…
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
Haocheng Lu, Nan Zhang, Wei Tao +4
Streaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as video frames arrive sequentially and user queries…
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
Wei Tao, Haocheng Lu, Xiaoyang Qu +4
One of the primary challenges in optimizing large language models (LLMs) for long-context inference lies in the high memory consumption of the Key-Value (KV) cache. Existing approa…