Skip to content
#

cuda-mps

Here are 3 public repositories matching this topic...

Language: All
Filter by language

Predictive preemption engine monitoring active UI and audio telemetry at 10Hz to forecast LLM prompts. Speculatively streams KV cache snapshots to local inference servers while dynamically throttling background workloads via CUDA MPS or OS level process priority.

  • Updated Sep 20, 2026
  • Python

Add this topic to your repo

To associate your repository with the cuda-mps topic, visit your repo's landing page and select "manage topics."

Learn more