Discrete-event simulation of multi-LoRA adapter serving strategies on a single GPU — comparing naive swap, hot-set preloading, and batch-by-adapter under variable VRAM pressure and arrival rates.
-
Updated
Jul 18, 2026 - Python
Discrete-event simulation of multi-LoRA adapter serving strategies on a single GPU — comparing naive swap, hot-set preloading, and batch-by-adapter under variable VRAM pressure and arrival rates.
Discrete-event simulation of LLM request routing across multiple serving instances: round-robin, least-load, prefix-aware, and hybrid cache-aware routing with queue-depth threshold sweep.
Add a description, image, and links to the s-lora topic page so that developers can more easily learn about it.
To associate your repository with the s-lora topic, visit your repo's landing page and select "manage topics."