3 citations · 7 across the 3 of their papers we have counts for
3 papers
Tool Calling: Enhancing Medication Consultation via Retrieval-Augmented Large Language Models
Zhongzhen Huang, Kui Xue, Yongqi Fan +5
Large-scale language models (LLMs) have achieved remarkable success across various language tasks but suffer from hallucinations and temporal misalignment. To mitigate these shortc…
Augmenting Hessians with Inter-Layer Dependencies for Mixed-Precision Post-Training Quantization
Clemens JS Schaefer, Navid Lambert-Shirzad, Xiaofan Zhang +7
Efficiently serving neural network models with low latency is becoming more challenging due to increasing model complexity and parameter count. Model quantization offers a solution…
Mixed Precision Post Training Quantization of Neural Networks with Sensitivity Guided Search
Clemens JS Schaefer, Elfie Guo, Caitlin Stanton +7
Serving large-scale machine learning (ML) models efficiently and with low latency has become challenging owing to increasing model size and complexity. Quantizing models can simult…