Skip to content

Stage 3: satisfies checks the error metric - #627

Draft
zzylol wants to merge 1 commit into
stack/univmon-l2-accuracyfrom
stack/satisfies-error-metric
Draft

zzylol wants to merge 1 commit into
stack/univmon-l2-accuracyfrom
stack/satisfies-error-metric

Conversation

@zzylol

@zzylol zzylol commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Stack: #574 → #620 → #618 → #621 → #627 → #625 → #628 → #632 → #634 → #616 → #617 → #622 → #624 → #629 → #630 → #631 → #633 → #635 → #636 → #637

Problem

DefaultAccuracyModel::satisfies compares only a guarantee's bound and δ with the target. Stage 3's build_violation therefore accepts a bound stated in the wrong metric. For example, a Count-Min L1 frequency bound (Frequency) passes a target on a distinct count (Cardinality) whenever its number is small enough.

Changes

  • AccuracyModel::answers(statistic, guarantee): a new trait method with a default implementation. It is a separate method because satisfies(guarantee, target) never sees the statistic. Its callers in Pass 1 check composed root guarantees, where no SketchStatistic exists.
    • An exact guarantee answers every statistic.

    • Otherwise the guarantee's metric must be one the statistic's ε is stated in:

      Statistic Accepted metrics
      Quantile Rank, RelativeValue
      Cardinality Cardinality, RelativeValue
      FrequencyL2, FrequencyEntropy RelativeValue
      PointCount Frequency, L2Frequency
      TopK Frequency, L2Frequency, TopKMembership (the score bound Pass 1 already checks top-k targets against)
    • A deployment with a registered cross-metric conversion can override answers. EstimatorAccuracy delegates it to its base model.

  • Stage 3: build_violation checks answers before satisfies. On a mismatch it rejects with "… guarantees a bound, which does not bound ".
  • Tests:
    • Unit test answers_requires_the_statistics_metric.
    • plan-selection test a_bound_in_another_metric_misses_the_target: a CountSketch+heap build passes for its TopK readout, but is rejected for Cardinality even at ε = δ = 0.5.

No example's selection changes. Examples planner-layering-1, 2, 3a, 3b, 4a and 4b select the same plans as before, and none of their rejections has the new reason.

Test plan

  • cargo fmt --all --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace: 1679 passed, 0 failed, 21 ignored (re-run after the rebase on main d4869a7)

Stacked on #621.

🤖 Generated with Claude Code

@zzylol
zzylol force-pushed the stack/univmon-l2-accuracy branch from 18cb590 to ba716f2 Compare October 5, 2026 03:07
@zzylol
zzylol force-pushed the stack/satisfies-error-metric branch from ba0c9e2 to 64ebe46 Compare October 5, 2026 03:07
@zzylol
zzylol force-pushed the stack/univmon-l2-accuracy branch from ba716f2 to e34c85c Compare October 5, 2026 06:21
@zzylol
zzylol force-pushed the stack/satisfies-error-metric branch from 64ebe46 to 4f65ff1 Compare October 5, 2026 06:21
Stage 3 compared only a summary guarantee's bound and failure
probability with the target, so a bound in one metric (Count-Min's L1
frequency error, say) could satisfy a target on another statistic (a
distinct count). `AccuracyModel::answers` maps each statistic to the
metrics its epsilon is stated in, and `build_violation` checks it
before `satisfies`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the stack/satisfies-error-metric branch from 4f65ff1 to e30d21a Compare October 8, 2026 16:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant