Conversation
recommended_sketch_configs.py turns sketch-bench's recommendations.csv into experiment_type configs (recommended sketch_parameters plus a default twin) for the dataset-analysis queries the cluster_data_exporter can replay, and summarizes finished runs against the predicted error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…onfigs One scrape holds about 470k series and takes about 5 s, so the 1 s default never completes; queries repeat no faster than the scrape. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR validates, end to end in ASAPQuery, the sketch configs that ProjectASAP/sketch-bench#130's
recommend_config.pypicks from #746's worst-case data parameters. It changes neither the planner nor the optimizer.recommended_sketch_configs.pyreadsrecommendations.csvand writes, for each query group, an experiment YAML withsketch_parametersset to the recommended CMS / KLL config, plus a twin YAML with the default config. It also maps feat(asap-tools): label query sets and skew fits for cluster traces #746 names to exporter names (cpu_rate→google_mean_cpu_usage_rate_0,msname→ms_name). CountSketch, DDSketch and top-k recommendations are skipped, because the planner has no matching sketch.Ten generated configs:
sum by (job_id), instant and 5m;quantile(0.99, cpu);sum by (ms_name), instant and 5m.Each comes in a recommended and a default version. The Alibaba configs scrape every 10 s, because one scrape is about 470k series.
Per-run summary with an all-keys error column and the query engine's memory;
kll_rank_error.pyfor the p99 rank error; results inrecommended_sketch_configs/README.mdandresults_summary.csv.How the experiment runs (existing infrastructure)
The measurement path is ASAPQuery's existing e2e infrastructure. This PR only generates configs for it.
asap-tools/experiments/experiment_run_e2e.py. Each generated YAML uses the existingsketchdbmode withquery_prometheus_too: true, soexperiment_utils/config.pygives the querier both servers:sketchdb(:8088) andprometheus(:9090).asap-tools/queriers/prometheus-client/main_prometheus_client.pycomputes onequery_unix_timeper repetition and sends every query in the group to every configured server at that time. Both engines therefore answer the same PromQL instant query evaluated at the samet, and Prometheus is the exact baseline.cluster_data_exporterreplays the traces, and both Prometheus and ASAP ingest from it.post_experiment/single_experiment/calculate_fidelity.pyandcompare_latencies.py.New in this PR, on top of that infrastructure:
sketch_parametersand the queries; the planner is unchanged.sum byqueries: average relative error over the 100 largest keys, the metric sketch-bench uses for CMS, plus query-engine memory, both in the run summary.calculate_fidelity.pyreturns NaN/inf for these queries because some keys sum to 0.kll_rank_error.pycomputes it offline against the replayed trace values.calculate_fidelity.pyreports value error, not rank error.Results
Each config ran once on an idle 56-core machine with 20 query repetitions; values are medians. ASAP and Prometheus are queried at the same timestamp, and Prometheus is the exact baseline. Error is the average relative error over the 100 largest keys for CMS, and rank error for p99. Latency is ASAP / Prometheus.
sum by (job_id), instantquantile(0.99)sum by (ms_name), instantsum by (ms_name)at 0.20, four times the 0.05 target, and KLL with the runner's default K=20.config.yamluses K=20, the planner uses K=500.Limits
Validation
python -m unittest discover -s testspasses inasap-tools/experiments; black, isort and flake8 are clean. Runningkll_rank_error.pyon the stored outputs reproduces the p99 medians.🤖 Generated with Claude Code