Skip to content

docs(ai): add garak LLM security scanning on Workbench - #848

Merged
luohua13 merged 11 commits into
mainfrom
docs/garak-llm-security-scanning
Sep 24, 2026
Merged

luohua13 merged 11 commits into
mainfrom
docs/garak-llm-security-scanning

Conversation

@luohua13

Copy link
Copy Markdown
Contributor

Adds a Solution article for scanning a published inference service with garak from Alauda AI Workbench.

docs/en/solutions/AI/LLM_Security_Scanning_with_garak_on_Workbench.md

Why

Installing garak on site is impractical in an isolated environment: it pulls packages from PyPI and detector models and datasets from Hugging Face. The article uses a prebuilt Workbench image with garak 0.17.0 and all of its offline assets, so the scan runs with no internet access.

Contents

  • Obtaining the published image and pushing it to the Private Registry, plus the offline tar path
  • Importing the WorkspaceKind the console needs in order to offer the image
  • Preparing scan.yaml, smoke test, running a scan, reading the reports
  • Adapting the chat template for non-Qwen models (with the steps to derive it from the model files), using the chat endpoint instead, selecting probes, custom probes, LLM judge detector
  • Image upgrade path

Verification

Everything in the article was executed end to end against qwen3-5-0-8b (Qwen3.5-0.8B on vLLM):

  • 21 probes in 775s; report, HTML summary and hitlog written to the Workspace volume
  • Detector models and datasets load from the image with the container fully offline (--net none)
  • Image pulled on a node other than the build host, confirming it is self-contained

Open item

ProductsVersion is not set in the frontmatter. The verification cluster runs aml-server v2.8.0-beta.4 with Workbench chart 2.0.0, which I could not map to a released AML version; the versions are stated in the Environment section instead. Happy to add the field if a reviewer can name the right version.

Prebuilt Workbench image with garak 0.17.0 and its offline detector
models, datasets and NLTK corpora, so an isolated cluster can scan a
published inference service without reaching PyPI or Hugging Face.

Covers obtaining the image, importing the WorkspaceKind the console
needs to offer it, preparing the scan configuration, running a scan and
reading the reports, plus adapting the chat template for non-Qwen models
and using an LLM judge detector.

Verified against Qwen3.5-0.8B on vLLM: 21 probes in 775s, reports
written to the Workspace volume, detectors loading from the image with
the container offline.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- drop the verified-versions line from Environment
- any registry the cluster can pull from, not specifically the platform
  Private Registry
- remove the Image contents section
- condense image retrieval to the address plus the push commands, and
  keep section anchors only where another section links to them

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The shipped sample posts to /v1/completions because the chat endpoint of
the verification service did not respond, but chat is the endpoint most
services expose and it lets the server apply the model's own chat
template. Present both plugins sections, chat first, and keep the
template-adaptation section for the completions fallback only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The chat endpoint is what services expose for instruct models and what
applications actually call, and the server applies the model's own chat
template. Scanning through /v1/completions needed a hand-built template
per model, and the rest generator drops all but the last turn, which
silently degrades the multi-turn probes; drop that path and the template
adaptation section with it.

Fix the thinking-mode setting: chat_template_kwargs has to be nested
under extra_body, otherwise the OpenAI client rejects it as an unexpected
keyword argument.

Add a section on what to scan: the model endpoint measures the model,
while the risk that matters lives in the application in front of it, so
cover passing the production system prompt and pointing the rest
generator at the application's own API.

Results table re-measured through the chat endpoint (21 probes, 701s),
plus a note on the report corruption that can break HTML generation on
long parallel runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The target does not have to be the in-cluster predictor Service: a
gateway, an ingress or any other endpoint serving /v1/chat/completions
works, so describe it as a base URL and mention the API key case.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- reports go to ~/garak/report via an absolute reporting.report_dir; a
  relative path lands under ~/.local/share/garak, which is awkward to
  open from the file browser
- remove the verification-results section and the report-corruption note
- refresh the smoke test output for the chat generator

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@luohua13
luohua13 merged commit 28362d2 into main Sep 24, 2026
1 check passed
@luohua13
luohua13 deleted the docs/garak-llm-security-scanning branch September 24, 2026 08:01

This branch was successfully deployed

1 active deployment
translate — 13e3783d Deployed Sep 24, 2026 by luohua13 via build #2715
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant