Heterogeneous GPU Sharing on Kubernetes
-
Updated
Oct 8, 2026 - Go
Heterogeneous GPU Sharing on Kubernetes
K100AI 模型部署实践 —— 海光 K100-AI 上大模型的成功配置线拉起包(当前:Qwen3.8-27B 六条精选线;持续更新)
海光 K100 LC(gfx926)上从零自研的 Qwen3.8-27B 推理运行时:RT4 权重格式、int4 算子、MTP3 投机解码、本地视觉塔与 OpenAI 兼容服务
MiniMax H3 INT8 ConvRot optimization and dual-K100AI QKV/Attention inference for Hygon K100AI (gfx928)
dsh plugin for scnet.cn
Open-source model serving and heterogeneous GPU management platform for enterprise private AI. NVIDIA / Ascend / Cambricon / Hygon, unified through OpenAI-compatible APIs.
A ~2,000-line miniature of vLLM v1 with a vLLM-style Platform/op-family hardware abstraction layer. One engine, four backends: NVIDIA CUDA, Hygon DCU (DTK/HIP), Ascend NPU, Moore Threads MUSA. Real continuous batching, paged KV cache, TP, CUDA graph — runs Qwen3 end-to-end
To associate your repository with the hygon topic, visit your repo's landing page and select "manage topics."