Skip to content

feat(asap-tools): validate recommended sketch configs end to end on cluster traces - #752

Closed
zzylol wants to merge 5 commits into
feat/trace-label-skewfrom
feat/e2e-recommended-sketch-configs
Closed

zzylol wants to merge 5 commits into
feat/trace-label-skewfrom
feat/e2e-recommended-sketch-configs

Conversation

@zzylol

@zzylol zzylol commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #746. Running these configs end to end with the local provider needs the runner fixes in #751.

Summary

This PR validates, end to end in ASAPQuery, the sketch configs that ProjectASAP/sketch-bench#130's recommend_config.py picks from #746's worst-case data parameters. It changes neither the planner nor the optimizer.

  • recommended_sketch_configs.py reads recommendations.csv and writes, for each query group, an experiment YAML with sketch_parameters set to the recommended CMS / KLL config, plus a twin YAML with the default config. It also maps feat(asap-tools): label query sets and skew fits for cluster traces #746 names to exporter names (cpu_rate → google_mean_cpu_usage_rate_0, msname → ms_name). CountSketch, DDSketch and top-k recommendations are skipped, because the planner has no matching sketch.

  • Ten generated configs:

    • Google sum by (job_id), instant and 5m;
    • Google quantile(0.99, cpu);
    • Alibaba MSMetrics sum by (ms_name), instant and 5m.

    Each comes in a recommended and a default version. The Alibaba configs scrape every 10 s, because one scrape is about 470k series.

  • Per-run summary with an all-keys error column and the query engine's memory; kll_rank_error.py for the p99 rank error; results in recommended_sketch_configs/README.md and results_summary.csv.

How the experiment runs (existing infrastructure)

The measurement path is ASAPQuery's existing e2e infrastructure. This PR only generates configs for it.

  • Runner: asap-tools/experiments/experiment_run_e2e.py. Each generated YAML uses the existing sketchdb mode with query_prometheus_too: true, so experiment_utils/config.py gives the querier both servers: sketchdb (:8088) and prometheus (:9090).
  • Querier: asap-tools/queriers/prometheus-client/main_prometheus_client.py computes one query_unix_time per repetition and sends every query in the group to every configured server at that time. Both engines therefore answer the same PromQL instant query evaluated at the same t, and Prometheus is the exact baseline.
  • Data: the existing cluster_data_exporter replays the traces, and both Prometheus and ASAP ingest from it.
  • Comparison: the existing post_experiment/single_experiment/calculate_fidelity.py and compare_latencies.py.

New in this PR, on top of that infrastructure:

  • Configs: the generator and the 10 configs. They set only sketch_parameters and the queries; the planner is unchanged.
  • Error metric for the sum by queries: average relative error over the 100 largest keys, the metric sketch-bench uses for CMS, plus query-engine memory, both in the run summary. calculate_fidelity.py returns NaN/inf for these queries because some keys sum to 0.
  • p99 rank error: kll_rank_error.py computes it offline against the replayed trace values. calculate_fidelity.py reports value error, not rank error.
  • Single-machine runs: the fixes in fix(asap-tools): run cluster-trace e2e experiments with the local provider #751, needed to run this infrastructure with the local provider.

Results

Each config ran once on an idle 56-core machine with 20 query repetitions; values are medians. ASAP and Prometheus are queried at the same timestamp, and Prometheus is the exact baseline. Error is the average relative error over the 100 largest keys for CMS, and rank error for p99. Latency is ASAP / Prometheus.

Query Config Predicted error Measured error Target met? Query engine memory p50 latency (ms)
Google sum by (job_id), instant Recommended CMS 3×4096 0.042 0.00005 ✅ 99 MB 124 / 1639
Default 3×1024 0.0031 ✅ 79 MB 116 / 1847
Google, same query, 5m range Recommended 3×4096 0.043 0.00005 ✅ 90 MB 100 / 3804
Default 0.0031 ✅ 74 MB 92 / 4262
Google quantile(0.99) Recommended KLL k=200 0.0020 0.0019 ✅ 44 MB 3.5 / 1860
Runner default K=20 0.0153 ❌ 43 MB 3.7 / 1857
Planner default K=500 0.0004 ✅ 46 MB 3.3 / 2070
Alibaba sum by (ms_name), instant Recommended CMS 3×16384 0.021 0.0013 ✅ 317 MB 629 / 3235
Default 3×1024 0.20 ❌ 275 MB 620 / 3130
Alibaba, same query, 5m range Recommended 3×16384 0.021 0.0013 ✅ 282 MB 678 / 4350
Default 0.20 ❌ 258 MB 640 / 4494
  • Every recommended config met its target. Two defaults did not: Alibaba sum by (ms_name) at 0.20, four times the 0.05 target, and KLL with the runner's default K=20.
  • Predictions are conservative upper bounds for CMS: measured error is 16–900× below the prediction. The prediction uses the worst case over the whole trace, while each run replays only its first minutes, and CMS's bound is loose. KLL's prediction is within 5%.
  • Larger sketches cost memory, not latency: 16–42 MB more query-engine memory, and ASAP latency is unchanged.
  • ASAPQuery has two KLL defaults: the e2e runner's config.yaml uses K=20, the planner uses K=500.

Limits

  • One run per config, covering the first minutes of one trace file per dataset.
  • Memory is the whole query engine's footprint, not per-sketch bytes.
  • CountSketch, DDSketch and top-k recommendations can't run until the planner has those sketches.

Validation

python -m unittest discover -s tests passes in asap-tools/experiments; black, isort and flake8 are clean. Running kll_rank_error.py on the stored outputs reproduces the p99 medians.

🤖 Generated with Claude Code

zzylol and others added 5 commits October 2, 2026 11:46
recommended_sketch_configs.py turns sketch-bench's recommendations.csv
into experiment_type configs (recommended sketch_parameters plus a
default twin) for the dataset-analysis queries the cluster_data_exporter
can replay, and summarizes finished runs against the predicted error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…onfigs

One scrape holds about 470k series and takes about 5 s, so the 1 s
default never completes; queries repeat no faster than the scrape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@milindsrivastava1997
milindsrivastava1997 deleted the branch feat/trace-label-skew October 3, 2026 22:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants