Development of a Large Language Model Inference Platform with Request Routing on Consumer-Grade Hardware

D.K. Sviridenko, E.V. Bobrova, P.V. Anni

Abstract


This paper presents the design, development, and testing of a prototype platform for large language model (LLM) inference built on a master–worker architecture using consumer-grade hardware running Windows + WSL2 + Docker with a fully open-source software stack. The platform provides a single point of entry via a web interface and an OpenAI-compatible API, routes requests to compute nodes by model name, and supports authentication and rate limiting. New nodes can be added without any changes to the client side. All nodes are hosted on MEPhI servers, so all traffic remains within the organization's perimeter. The nodes serve models of different classes, including a mixture-of-experts (MoE) model distributed across two GPUs. An analytical model of node serving capacity is proposed, based on the VRAM budget and KV-cache size. Practical significance: the platform provides access to LLMs on existing hardware, eliminating dependence on external cloud services and ensuring that data never leaves the organization's perimeter.


Full Text:

PDF (Russian)

References


Zhao W. X., Zhou K., Li J. et al. A Survey of Large Lan-guage Models // arXiv preprint arXiv:2303.18223. 2023.

Zhou Z., Ning X., Hong K. et al. A Survey on Efficient Inference for Large Language Models // arXiv preprint arXiv:2404.14294. 2024.

Yuan Z., Shang Y., Zhou Y. et al. LLM Inference Un-veiled: Survey and Roofline Model Insights // arXiv preprint arXiv:2402.16363. 2024.

Zhen R., Li J., Ji Y. et al. Taming the Titans: A Survey of Efficient LLM Inference Serving // Proceedings of the 18th Interna-tional Natural Language Generation Conference (INLG). Hanoi, 2025. P. 522–541.

Miao X. et al. Towards Efficient Generative Large Lan-guage Model Serving: A Survey from Algorithms to Systems // ACM Computing Surveys. 2025.

Lee J. et al. A Survey on Inference Engines for Large Language Models: Perspectives on Optimization and Efficiency // arXiv preprint arXiv:2505.01658. 2025.

Girija S. S. et al. Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Tech-niques // 2025 IEEE 49th Annual Computers, Software, and Applica-tions Conference (COMPSAC). 2025.

Liu J., Tang P., Wang W. et al. A Survey on Inference Optimization Techniques for Mixture of Experts Models // ACM Computing Surveys. 2026. DOI: 10.1145/3794845.

G.G. Hajilov, A.V. Khatunov, T.A. Voloshin, M.G. Zhab-itsky MEPhI Higher Engineering School Digital polygon for educa-tional and practical projects infrastructure support International Jour-nal of Open Information Technologies ISSN: 2307-8162 vol. 13, no. 8, 2025

Ollama Project. Ollama: Get up and running with large language models locally. URL: https://ollama.com (дата обращения: 07.07.2026).

Yang A., Li A., Yang B. et al. Qwen3 Technical Report // arXiv preprint arXiv:2505.09388. 2025.

Gemma Team (Google Deep-Mind). Gemma 4 Model Card. URL: https://ai.google.dev/gemma/docs/core/model_card_4 (дата обращения: 07.07.2026).

Frantar E., Ashkboos S., Hoefler T., Alistarh D. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers // ICLR. 2023.

Dao T., Fu D. Y., Ermon S. et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness // NeurIPS. 2022.

Microsoft. GPU acceleration in WSL 2. URL: https://learn.microsoft.com/windows/wsl/gpu-acceleration (дата обращения: 07.07.2026).

Kwon W., Li Z., Zhuang S. et al. Efficient Memory Man-agement for Large Language Model Serving with PagedAttention // SOSP. 2023. P. 611–626.

Ainslie J., Lee-Thorp J., de Jong M. et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints // Proceedings of EMNLP. 2023.


Refbacks

  • There are currently no refbacks.


Abava  Кибербезопасность Monetec 2026 СНЭ

ISSN: 2307-8162