18 papers
Fighting Numerical Hallucinations via Data-centric Compilation for Online Financial QA
Hao Chen, Xing Tang, Qirui Liu +6
Large Language Models (LLMs) have significantly advanced online data services, particularly in the domain of financial question answering (FinQA). However, such systems remain susc…
Less Is More: Elevating RAG via Performance-Driven Context Compression
Ziqiang Cui, Yunpeng Weng, Xing Tang +7
Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for improving the timeliness of knowledge updates and the factual accuracy of large language models. Howeve…
Looking Farther with Confidence: Uncertainty-Guided Future Learning for Sequential Recommendation
Ziqiang Cui, Xing Tang, Peiyang Liu +4
Sequential recommendation effectively models dynamic user interests but continues to face challenges related to data sparsity. While self-supervised learning has alleviated this is…
Nonlinearity as Rank: Generative Low-Rank Adapter with Radial Basis Functions
Yihao Ouyang, Shiwei Li, Haozhao Wang +6
Low-rank adaptation (LoRA) approximates the update of a pretrained weight matrix using the product of two low-rank matrices. However, standard LoRA follows an explicit-rank paradig…
Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA
Xing Tang, Hao Chen, Shiwei Li +7
Large language models (LLMs) have been incorporated into numerous industrial applications. Meanwhile, a vast array of API assets is scattered across various functions in the financ…
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
Ailin Huang, Ang Li, Aobo Kong +213
We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most wh…