1 paper · 1 filter
Abu Hanif Muhammad Syarubany, Chang Dong Yoo
Enterprise deployments of large-language model (LLM) demand continuously changing document collections with sub-second latency and predictable GPU cost requirements that classical…