InferMatrix
Popular repositories Loading
-
-
vllm
vllm PublicForked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
-
light-llm-simulator
light-llm-simulator PublicA python-based LLM performance simulator with vLLM, notable for its lightweight design, easy scalability.
-
vllm-ascend
vllm-ascend PublicForked from vllm-project/vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
-
-
Repositories
- SparseGQA Public
- vllm Public Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- vllm-ascend Public Forked from vllm-project/vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
- vllm-omni Public Forked from vllm-project/vllm-omni
A framework for efficient model inference with omni-modality models
- light-llm-simulator Public
A python-based LLM performance simulator with vLLM, notable for its lightweight design, easy scalability.
- agentic-stack Public Forked from vllm-project/agentic-api
Stateful API logic for agentic applications using vLLM
- LM-service Public
- Mooncake Public Forked from kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Top languages
Loading…
Most used topics
Loading…