Skip to content

test(example4): B1 vs B2 as model-relative invariants (Q49) - #612

Draft
zzylol wants to merge 1 commit into
stack/509-y2-deployment-docfrom
stack/509-q49-example4-invariants
Draft

zzylol wants to merge 1 commit into
stack/509-y2-deployment-docfrom
stack/509-q49-example4-invariants

Conversation

@zzylol

@zzylol zzylol commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Stack: Wave 3 chain: #605 → #607 → #608 → #609 → #610 → this PR

Rebased on main d4869a7 (DF 54).

Problem

Example 4's spec says B1 (tumbling KLL panes built at ingestion time) is never costlier than B2 (rebuild every pane from raw samples at each evaluation). The built-in model contradicts that for Pattern B, so both tests were #[ignore]d:

  • B1 keeps 6 panes of 1M per-series KLLs, about 6.1 GB, which costs 768.58 cost/s.
  • B2 costs 5.17 cost/s, or 45.17 when raw retention is charged. It is over the 200 ms latency limit anyway.

Q49 (a): write the tests as invariants that follow from the model, and add a workload where B1 actually wins.

What the model implies

B1 keeps one KLL per series per pane for as long as the window lasts. B2 keeps no state, but it reads the window's raw samples at every evaluation. It pays to retain those samples only when the deployment doesn't keep raw data (#609). So B1 wins only when its panes are smaller than the raw samples they cover, and keeping those samples would otherwise be charged to the plan. When raw data is kept, B1 wins only if B2 breaks the latency limit.

Workload (p99 quantile_over_time) Raw data kept B1 cost/s B2 cost/s Test
Pattern B: 1M series, 15 s sampling, 5 m window, every 1 m yes 768.58 invalid (310 ms > 200 ms) stage3_b_rebuilding_a_million_series_breaks_the_latency_bound
10k series, 15 s, 5 m, every 1 m yes 7.69 0.05 stage3_b_kept_raw_data_makes_rebuilding_cheaper
10k series, 15 s, 5 m, every 1 m no 7.69 0.45 stage3_b_panes_larger_than_their_raw_data_lose
1k series, 1 s, 1 h window, every 10 m no 0.90 (selected) 7.30 stage3_b_panes_smaller_than_their_raw_data_win

The tests check which option is cheaper, not the numbers. The numbers above are from the illustrative v2 calibration.

Changes

  • planner_layering_example4.rs: removes the two ignored tests stage3_b_rebuilding_every_window_costs_most and stage3_b_prefers_ingestion_time_tumbling_windows, and adds the four tests above, a pattern_b_variant workload builder, and run_with_raw_data.
  • planner_layering_common: adds run_stages_with, which plans with the given deployment inputs. run_stages calls it with the executor's inputs.

Test plan

  • cargo fmt --check and cargo clippy -p asap-integration-tests --all-targets -D warnings are clean.
  • cargo test -p asap-integration-tests: 209 passed, 0 failed.

🤖 Generated with Claude Code

Replace the two ignored Pattern B cost tests, which asserted B1 <= B2 for
any workload, with invariants the cost model actually implies:

- at 1M series, rebuilding the panes (B2) breaks the 200 ms latency bound;
- when the deployment keeps raw data, B2 is cheaper whenever it is valid;
- panes larger than the raw samples they cover (15 s sampling, 1-min panes)
  lose even when the deployment does not keep raw data;
- panes smaller than their raw samples (1 s sampling, 10-min panes over 1 h)
  win when raw data is not kept, and the planner selects B1.

run_stages_with plans with given deployment inputs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the stack/509-y2-deployment-doc branch from 73ded2f to c0f007f Compare October 5, 2026 06:21
@zzylol
zzylol force-pushed the stack/509-q49-example4-invariants branch from d28b006 to a23ab1a Compare October 5, 2026 06:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant