#
sm70
Here are 4 public repositories matching this topic...
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
-
Updated
Jul 10, 2026 - Python
Serve Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB (DGX-1, TP8) for agentic coding & ops — a downstream fork of 1Cat-vLLM.
-
Updated
Jul 3, 2026 - Python
VastLLM: a production-oriented FastLLM fork for native C++ inference, V100/SM70, long context, and Qwen3.5/3.6; upstream: ztxz16/fastllm
-
Updated
Aug 2, 2026 - C++
Improve this page
Add a description, image, and links to the sm70 topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the sm70 topic, visit your repo's landing page and select "manage topics."