From 86edc7d4425face621e62f35609003aa678cb1fb Mon Sep 17 00:00:00 2001 From: Selvomega Date: Mon, 28 Sep 2026 02:07:47 +0000 Subject: [PATCH 01/46] docs: propose operator flattening (#468) Add a design proposal for merging the pre- and post-ASAP operator sets into one Operator { Basic(OriginalOp), Ext(ASAPOp) } tree, and list it in the proposals index. Co-Authored-By: Claude Opus 5.5 --- docs/design_docs/proposals/README.md | 1 + .../proposals/operator-flattening.md | 500 ++++++++++++++++++ 2 files changed, 501 insertions(+) create mode 100644 docs/design_docs/proposals/operator-flattening.md diff --git a/docs/design_docs/proposals/README.md b/docs/design_docs/proposals/README.md index 76cc23402..78084caa4 100644 --- a/docs/design_docs/proposals/README.md +++ b/docs/design_docs/proposals/README.md @@ -8,3 +8,4 @@ extensions. A design document is not a promise of downstream runtime support. - [ASAPQuery rule coverage](asapquery-rule-coverage.md) - [UnivMon frequency summary](univmon-frequency-summary.md) - [ASAP-aware mapping proposals](asap-aware-mapping/README.md) +- [Operator flattening](operator-flattening.md) diff --git a/docs/design_docs/proposals/operator-flattening.md b/docs/design_docs/proposals/operator-flattening.md new file mode 100644 index 000000000..577ba45fb --- /dev/null +++ b/docs/design_docs/proposals/operator-flattening.md @@ -0,0 +1,500 @@ +# Operator flattening + +> Status: proposed, not implemented. Problem statement: +> [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). Code +> references and counts are against `main` at `8acb472`. + +**In one sentence**: split `QueryExpr` into original operators (`OriginalOp`) and +scalar expressions (`ScalarExpr`); introduce `Operator { Basic(OriginalOp), Ext(ASAPOp) }` +so that every operator's children are `Rc`; then delete the parallel +post-ASAP IR (`SummaryExpr`, `SummaryNode`, `SummarySchema`, and the relational +variants of `ValueOperation`). + +**Roadmap**: + +| Part | Sections | Question answered | +|---|---|---| +| I. New IR | §1 Types, §2 Per-node information | Which node kinds exist, and what each carries | +| II. Upstream and downstream changes | §3 Frontends and entry → §4 Planner → §5 Validation → §6 Export → §7 Other consumers | How each stage changes, in data-flow order | +| III. Implementation plan | §8 Stages and tests, §9 Out of scope | How to land it, and what is deliberately left out | + +--- + +# I. New IR + +## 1. Types + +### 1.1 Overview + +``` +Operator operator: one node of the DAG; every child is Rc> +├─ Basic(OriginalOp) original operator: today's relational QueryExpr variants (Scan / Filter / Join / Aggregate / SetOp …) +└─ Ext(ASAPOp) ASAP operator: summary operators (SummaryAgg / SummaryEstimate …) + +ScalarExpr scalar expression: only appears inside operator fields (predicates, SELECT lists); never a node +``` + +- **Naming**: the suffix says what a type is — `*Op` is an operator, `*Expr` is an + expression. `OriginalOp` holds the operators pre-ASAP already has, as opposed to + the extension `ASAPOp`. The name `QueryExpr` goes away. +- **The generic `C`**: the existing column-reference stage of `QueryExpr` — + `ColumnRef` (by name) out of the frontends, `ColumnId` (by position) after + `resolve_root`. A parent and its children must be in the same stage, so all + three types carry the same `C`. + +### 1.2 `OriginalOp` and `ScalarExpr`: splitting `QueryExpr` + +Today one `QueryExpr` enum holds two kinds of things: + +```sql +SELECT l_quantity * 2 AS q2 FROM lineitem WHERE l_quantity > 10 +``` +``` +Project { cols: [ProjectItem { expr: Arithmetic(Column(4) * Literal(2)) }], ← scalar expression: computes one value per row + child: Filter { pred: Predicate(Compare(Column(4) > Literal(10))), ← scalar expression + child: Scan lineitem } } ← relational operator: rows in, rows out +``` + +`Filter.child` (rows) and the `Compare` inside `Filter.pred` (a value) are both +`QueryExpr`; only the field position tells them apart. The code already knows +which variants are scalars: `output_schema` returns `ScalarHasNoRowSchema` for all +of them (`types/src/pre_asap/query_expr.rs:1447`). Split them into two types: + +```rust +pub enum OriginalOp { + Scan { .. }, Filter { pred: Predicate, child: Rc> }, Project { cols: Vec>, child }, + Aggregate { .. }, Join { .. }, SetOp { .. }, Concat { .. }, Dedup { .. }, Sort { .. }, Limit { .. }, BinaryOp { .. }, + SQLWindowFunc { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, + ScalarBridge(Rc>), // formerly PromqlScalarBridge: a scalar in operator position (the `2` in PromQL `v * 2`, a bare scalar query) + EvalTimestamp, // PromQL time(): operator position, one row, one column +} +pub enum ScalarExpr { // children are ScalarExpr only + Column(C), Literal(ScalarValue), Compare { .. }, BoolAnd(..), BoolOr(..), Not(..), IsNull(..), IsNotNull(..), + Cast { .. }, InList { .. }, FunctionCall { .. }, Arithmetic { .. }, Case { .. }, CurrentTimestamp, +} +pub struct Predicate(pub Rc>); +pub struct ProjectItem { pub alias: Option, pub expr: ScalarExpr } +``` + +- **Ambiguous variants**: `PromqlScalarBridge`, `EvalTimestamp`, + `PromqlScalarFromVector` and `PromqlVectorFromScalar` appear in operator position + and have a row schema, so they belong to `OriginalOp`. `CurrentTimestamp` (SQL + `NOW()`) is used both ways: it has a row schema (`query_expr.rs:1388`) and takes + part in scalar type inference (`:1674`, e.g. `WHERE ts > NOW()`). It goes into + `ScalarExpr` and is written `ScalarBridge(CurrentTimestamp)` in operator + position. **Open**: confirm how the frontends place `CurrentTimestamp` in + operator position. +- **Benefits**: a scalar in operator position (`Basic(Column(3))`) is no longer + expressible; the `ScalarHasNoRowSchema` runtime error disappears; `canonicalize` + and `pre_asap/cse` already walk only the relational skeleton and treat scalars as + opaque data (see the `pre_asap/cse.rs` module docs) — the types now say so. + +### 1.3 `ASAPOp` + +```rust +pub enum ASAPOp { + SummaryAgg { child: Rc>, family: ASAPType, input, reduction, grouping, // family was SummaryFamilyType, "never Plain" by convention; now by type + guarantee: Option }, // Some only for the ExactAggregate family + SummaryEstimate { child, query: SketchQuery, + guarantee: Option }, // None = the AccuracyModel has no error model for this family + SummaryMerge { children: Vec>>, timing: ExecutionTiming }, + SummarySubtract { left, right }, + SummaryDelete { child, key: C }, // was ColumnRef; now C like everything else + SummaryJoin { outer, inner, key: C, family: ASAPType }, + // The four summary-specific ValueOperation variants, lifted to the top level instead of { child, operation, timing } + FinalizeExactAccumulator { child, timing: ExecutionTiming }, + MaintainPopulation { child, population }, // always ingestion time + ReadPopulation { child, readout }, // always query time + Extension { child, name: String }, +} +impl ASAPOp { pub fn guarantee(&self) -> Option<&ResultGuarantee>; } // only the two variants above store one; see §2.2 +``` + +**Deleted post-ASAP types** — each is expressed by an existing `OriginalOp` variant: + +| Deleted | Expressed as | +|---|---| +| `SummaryExpr::KeepPreAsap(q)` | no wrapper: the subtree `q` is itself `Basic(..)` | +| `ValueOperation::{Project, Filter, Sort, Limit}` | `OriginalOp::{Project, Filter, Sort, Limit}` | +| `SummaryExpr::BinaryOp`, `SummaryExpr::RelationalJoin` | `OriginalOp::BinaryOp`, `OriginalOp::Join` | +| `ValueOperation::Exact(Aggregate)`, `ExactOperation` | `OriginalOp::Aggregate` | +| `SummaryNode` (the node wrapper) | not needed; per-node information is covered in §2 | + +### 1.4 Child slots + +```rust +// Today // After +Filter { Filter { + pred: Predicate(Rc), // scalar pred: Predicate(Rc), + child: Rc, // rows child: Rc, // either Basic(..) or Ext(SummaryEstimate ..) +} } +``` + +**Rule**: every `OriginalOp` field that holds input rows has type `Rc>`. +There are about 25 such fields — `Filter.child`, `Join.left` / `Join.right`, both +sides of `SetOp`, `Concat.children`, `BinaryOp.lhs` / `BinaryOp.rhs`, the `child` of +each `Promql*` operator, and so on. + +**Effect**: an ASAP operator can sit directly under any original operator. Both +sides of a `SetOp`, for example, can be `SummaryEstimate`s. + +**Exception: `Concat.children`**. Today it is `Vec`: branches are stored +by value and have no `Rc` identity. The planner's search and assembly (§4) +identify nodes by `Rc` pointer — targets are registered by pointer and holes are +recognized with `Rc::ptr_eq` — so a `Concat` branch can be neither a target nor a +hole. This affects, for example, the branches SQL `ROLLUP` and PromQL +`histogram_quantiles` are lowered into. As `Vec>` it is handled like +any other child slot and §4 needs no special case (§8 stage 0). + +## 2. Per-node information + +A post-ASAP node carries three pieces of information today. After the change: + +| Field | Meaning | Today | After | +|---|---|---|---| +| `schema` | which columns the node outputs, and their types | stored on every `SummaryNode` | not stored; `output_schema()` computes it on demand (§2.1) | +| `guarantee` | how far the node's output value can be off (metric, bound, failure probability, provenance) | stored on every `SummaryNode` | stored on `SummaryAgg` / `SummaryEstimate`; derived by `GuaranteeIndex` for everything else (§2.2) | +| `timing` | whether the node runs at ingestion time or at query time | stored on the `BinaryOp` / `ValueOperation` / `SummaryMerge` variants | not stored on `OriginalOp` (§5); kept on `SummaryMerge` and `FinalizeExactAccumulator` | + +### 2.1 Schema: two types become one + +A schema is a node's list of output columns: name, type, nullability. There are two today: + +| | pre-ASAP | post-ASAP | +|---|---|---| +| Type | `Schema { columns: Vec, time_index, unique_keys, closed }` | `SummarySchema { fields: Vec, time_index }` | +| Column type | `Column.dtype: DataType` — plain values only (`Int64`, `Float64`, `Utf8`, …) | `SummaryField.dtype: SummaryFamilyType` — a plain value `Plain(DataType)` or summary state (`Sketch(Kll, …)`, …) | +| Where it lives | not stored; `QueryExpr::output_schema()` computes it | stored on every `SummaryNode` | + +In the new IR both kinds of operator share one tree, so `Operator::output_schema()` +must describe any node, and a column type must be able to hold either a plain +value or summary state: + +``` +SummaryEstimate(Quantile .99) → [p99: DataType(Float64)] + SummaryAgg(Kll) → [state: ASAPType(Sketch(Kll, k=269))] + Scan lineitem → [l_orderkey: DataType(Int64), l_quantity: DataType(Float64), …] +``` + +So keep `Schema`, delete `SummarySchema` / `SummaryField`, and give `Column.dtype` a new type, `ColumnType`: + +```rust +pub enum ColumnType { + DataType(DataType), // a plain value; the existing DataType, unchanged (Int64, Float64, Utf8, …) + ASAPType(ASAPType), // ASAP state +} +pub enum ASAPType { // SummaryFamilyType without its Plain variant + ExactAggregate(ExactKind, ExactParams), Sketch(SketchKind, GroupingStrategy), + Sample(SamplingKind, SamplingParams), Wavelet(WaveletKind, WaveletParams), StatModel(StatModelKind, StatModelParams), +} + +pub struct Column { pub name, pub dtype: ColumnType, pub nullable, pub table: Option } +impl Column { + pub fn plain(name, DataType) -> Self; // build a plain column (frontends, catalogs) + pub fn plain_dtype(&self) -> Option<&DataType>; // None for an ASAP-state column + pub fn expect_plain_dtype(&self) -> &DataType; // frontends and scalar type inference: a state column is a bug, panic +} +``` + +- **Naming**: as in §1, `DataType` / `ASAPType` mirror `OriginalOp` / `ASAPOp`. + The name `SummaryFamilyType` goes away: once every column carries it, a plain + column reading `SummaryFamilyType::Plain(Float64)` is a misnomer. With two + levels, "is this column state?" is a check on the outer variant only — exactly + the split `check_plain_operands` needs. +- **`DataType` itself is unchanged.** It remains part of the scalar type system + (the type of a `Literal`, the result type of `Arithmetic`, catalog column + declarations); it now sits inside `ColumnType::DataType(..)`. Code that reads + `Column.dtype` expecting a `DataType` unwraps one more level: code that only + ever sees plain columns (frontends, scalar type inference) uses + `expect_plain_dtype()`; code that may see state (planner, validation, export) + uses `plain_dtype()` or matches. A scalar expression never legitimately reads a + state column — `check_plain_operands` (§5) rejects such a DAG first. +- **Elements of nested types**: `DataType::List { element: Box }` and + `DataType::Struct { fields: Vec }` have `Column` elements. With + `Column.dtype: ColumnType` an element could become ASAP state (a "list of KLL + sketches"), which is not intended. These two use a plain-only field instead: + + ```rust + pub struct PlainField { pub name: String, pub dtype: DataType, pub nullable: bool } // nested fields are already unqualified (List's doc comment), so no table + DataType::List { element: Box } + DataType::Struct { fields: Vec } + ``` + + `DataType::Map { key: Box, value: Box, .. }` already uses + `DataType` only and is unchanged. There are 36 construction or match sites of + `DataType::{List, Struct, Map}`. +- **Why keep `Schema`**: `ColumnType::DataType(..)` represents every plain type, so + nothing is lost, and `Schema` additionally has `unique_keys` / `closed` / + `Column.table`, which merging the other way would drop. +- **Computed on demand**: schemas are no longer stored on nodes; `Operator::output_schema()` + computes them. `OriginalOp` keeps today's `QueryExpr::output_schema()` logic, and + `ASAPOp` follows the table below. These rules are currently spread over + `asap-aware-mapping/src/replacement.rs` and `types/src/post_asap/execution_data_state.rs`; + they move into `asap-types`. +- **Deleted conversions**: `lift()` (`replacement.rs:3352`) and `lift_plain` + (`execution_data_state.rs:725`), which turn a `Schema` into a `SummarySchema`, and + the reverse `plain_schema` (`execution_data_state.rs:707`). With one schema type + there is nothing to convert. +- **Size of the change**: 7 direct `Column { .. }` literals; the deleted + `SummarySchema {}` (36) and `SummaryField {}` (19) literals. In non-test code + `SummaryFamilyType` appears 197 times (42 of them with `Plain(`) and `.dtype` is + read at 67 sites; all are rewritten against the new types. + +| `ASAPOp` | Output schema | +|---|---| +| `SummaryAgg` | the grouping columns as-is + one `ASAPType(family)` column | +| `SummaryEstimate` | the row schema determined by the `SketchQuery` | +| `SummaryMerge` / `Subtract` / `Delete` / `Join` | one `ASAPType(family)` column | +| `FinalizeExactAccumulator` | the child's schema, with each `ASAPType(ExactAggregate ..)` column turned into the matching `DataType(..)` | +| `MaintainPopulation` / `ReadPopulation` | existing implementation | + +### 2.2 Guarantee: some stored, some derived + +| Node | Where its guarantee comes from | +|---|---| +| `SummaryAgg`, `SummaryEstimate` | **stored in a field**. Computed at binding time by the `AccuracyModel` from the summary's parameters and the evidence; it cannot be recovered afterwards, and selection needs it during search to decide whether a candidate meets the accuracy target (`replacement.rs:5831`) | +| `OriginalOp` (`Project`, `Join`, …), `FinalizeExactAccumulator` | composed from the children by rule (`relational_join_guarantee`; `exact_operation_rule` in `accuracy/composition.rs`). A subtree with no `Ext` descendant is exact | +| `MaintainPopulation`, `ReadPopulation` | always exact | +| `SummaryMerge` / `Subtract` / `Delete` / `Join` | output state, so no guarantee of their own — but they change the error of a later readout (see below) | + +The derived part goes into one table: + +```rust +/// One traversal, keyed by node pointer — the same approach as the ExecutionDataStateAssignment +/// returned by validate_execution_data_states. +pub struct GuaranteeIndex(HashMap<*const Operator, ResultGuarantee>); +pub fn guarantee_index(root: &Rc) -> GuaranteeIndex; +``` + +- The composition rules also depend on the `AccuracyModel`, so the planner must + compute `GuaranteeIndex` with the same model and hand it over together with the + assembled root (`assemble_selected_dag`). `compile_executable_dag` uses it to fill + `ExecutableDagNode.guarantee`. +- Not chosen: wrapping every node to store a guarantee — frontends and `resolve` + would then have to build the wrapper too. +- **Gap for state nodes**: `SummaryMerge` / `Subtract` / `Delete` / `Join` change the + error of a later readout. For a CMS, `a − b` has error bound `ε·(‖a‖₁ + ‖b‖₁)`, so + the relative error is large when `a` and `b` are close. A readout's guarantee is + computed today as if its input were a sketch built directly by a `SummaryAgg`, + ignoring any intermediate operation. The planner on `main` never emits these four + nodes (they are built only in tests, and `SummaryMerge` is documented as inserted by + a deployment), so the gap is latent. Not addressed here (§9). + +--- + +# II. Upstream and downstream changes + +## 3. Frontends and the optimizer entry + +Trees built by the frontends and by `resolve` contain only `Basic`. They access +children through `expect_basic()`; before a tree reaches the optimizer, the entry +checks it once and rejects any `Ext`: + +```rust +impl Operator { + pub fn contains_ext(&self) -> bool; + pub fn expect_basic(&self) -> &OriginalOp; // an Ext here is a bug: panic +} +``` + +**Where the check goes**: on `main` the optimizer entries are `search_workload` / +`search_workload_with` / `search_workload_with_targets`. They do not return +`Result`, so the check starts as `assert!(!root.contains_ext())` at each entry. If +the single `ParsedWorkload` entry point (#429) lands on `main` first, the check +moves into `ParsedWorkload::new` and returns an error (it already checks the entry +count). + +**Why not a compile-time guarantee** (an associated type `Ext` on `ColState`, with +`ColumnRef::Ext = Never`): + +| | Compile time | Run time (this proposal) | +|---|---|---| +| Types | one more associated type on `ColState`, plus an empty enum | the `Ext` variant holds `ASAPOp` directly | +| What it catches | only frontend code before `resolve`. A frontend already returns a `ColumnId` tree, where `Ext` is allowed | **every input to the optimizer**: frontend code after `resolve`, deserialized plans, third-party frontends, hand-written test IR | +| Frontends matching children | `basic()` is total | `expect_basic()`, which can in principle panic | + +The compile-time version protects only a small stretch inside the frontends; the +runtime check sits at the entry, covers more, and keeps the types simpler. So one +`Operator` type serves both pre- and post-ASAP, told apart by the entry check +rather than by a type parameter. + +## 4. Planner: search and assembly + +**Candidates**: three shapes become two. + +```rust +pub enum Replacement { + Subtree(Rc), // formerly Summary(Rc) and Rewrite(Rc); "introduces a summary" = contains_ext() + ExactComposition { .. }, // unchanged, except its plan becomes Rc instead of Rc +} +``` + +**Holes**: a subtree of a candidate that is `Rc::ptr_eq` to some target is a hole, +left for that target's own choice to fill. A candidate that needs a specific +implementation of a child (for example a `SummaryAgg` that needs an exact +accumulator value, as in `realize_temporal_average`) inlines the child as a new node; +the pointer differs, so it is not a hole. `realize_child_with` changes from +"realize the child recursively, else `keep_pre_asap`" to "else leave a hole". + +**Assembly**: one generic rule replaces `assemble_residual`. + +```rust +fn assemble(&self, t: &Rc) -> Rc { + memo by ptr; // a shared child stays one Rc — today's assembled_nodes + let body = match self.chosen(t) { + Some(Subtree(r)) => r, // Rewrites are assembled further down too (today replacement.rs:4571 wraps the whole rewrite in keep_pre_asap and ignores the choices below) + Some(ExactComposition{..}) => composition.plan, + None => t, // no candidate chosen: keep the node, recurse into its children — what assemble_residual does for only four operators + }; + body.map_children(|c| if is_target(c) { self.assemble(c) } else { c }) +} +``` + +After assembly, run the data-state validation (§5) once; if filling a hole is +illegal (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`), that hole falls +back to its original subtree. This generalizes the fallback `relink_summary` does +today for `Aggregate` only. Finally compute `GuaranteeIndex` (§2.2). + +- **`map_children`**: reuse `rebuild_children` from `pre_asap/cse.rs:576` as + `Operator::map_children`, dispatching `Basic` to `OriginalOp::map_children` and + `Ext` to `ASAPOp::map_children`. +- **Deleted**: `assemble_residual`, `relink_summary`, the `query_time_nested_sum` / + `contains_aggregate` special cases, `keep_pre_asap` / `keep_pre_asap_rc`, and the + branch of `finalize_exact_accumulator` that builds an empty schema for `KeepPreAsap`. + +How the three problems in #468 are resolved: + +| Problem | Resolution | +|---|---| +| 1. One operator, two spellings: a `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside it | one set of types; `assemble` no longer builds `ValueOperation`s | +| 2. `KeepPreAsap` is opaque: nothing outside can reference the `Scan` inside, so an exact aggregate and a sketch cannot share one scan | `Scan` is no longer wrapped; `SummaryAgg{child: scan}` and `Aggregate{child: scan}` can point to the same `Rc`. Splitting a multi-measure `Aggregate` into `avg` + KLL is a binding-rule change; this proposal only makes the split expressible | +| 3. `SetOp` and similar operators have no post-ASAP counterpart and must stay inside `KeepPreAsap`, so no subtree below them can use a summary | `SetOp` takes the `None => t` path and both children are assembled independently | + +## 5. Validation: execution data states + +Today's rule `produced_data_state(KeepPreAsap) = None` (decided by the edge that +reaches the node) extends to every `OriginalOp`: + +| Node | Produced state | +|---|---| +| `OriginalOp` | none of its own; assigned by the consuming edge (`QUERY_ROWS` at the root) and passed down unchanged | +| `SummaryAgg` | `{the child's timing, SummaryState}`; ingestion time when the child is an `OriginalOp` | +| other `ASAPOp` | unchanged | + +The `KeepPreAsap` / `BinaryOp` / `ValueOperation` / `RelationalJoin` arms of +`validate_execution_data_states` merge into one `Basic` arm: + +- Pass the node's state to each child; an `Ext` child is checked against the + `Ext` edge rules (e.g. `SummaryEstimate` only on the query side). +- `check_plain_operands` stays: every column an `OriginalOp` references must be + `ColumnType::DataType`; `Project` / `Filter` / `Sort` / `Limit` may pass + `ExactAggregate` columns through (today's exception). This is what rejects + `Project(Ext(SummaryAgg))`, i.e. projecting sketch state as if it were rows. +- The extra ingestion-side constraints on `BinaryOp` (arithmetic only, identical + schemas on both sides, exactly one `Float64`) move into this arm, per variant. +- `AmbiguousKeepPreAsap` is renamed `AmbiguousDataState`; its meaning is unchanged. + +## 6. Export: fragments in the executable DAG + +`ExecutableDag` has four payloads for original operators today — `fallback{expression: QueryExpr}`, +`binary`, `value` and `relational_join`. Afterwards there is one: + +```rust +ExecutableOperatorPayload::Relational { + /// An Operator with no Ext (checked at construction). Leaves are real Scans, or + /// Scan { source: Source::DagInput { role } } standing for an incoming edge whose + /// schema is that edge's intermediate_schema. + expression: Operator, +} +``` + +`Source` gains a variant `DagInput { role: EdgeRole }`. `compile_executable_dag` works as follows: + +1. Starting from the root, take the **largest connected subtree containing no + `Ext`** as one fragment node; wherever the fragment meets an `Ext`, cut an edge + and put a `DagInput` leaf in the fragment. +2. Each `Ext` node maps one-to-one onto the existing `summary_agg` / + `summary_estimate` / … payloads; `FinalizeExactAccumulator` / `MaintainPopulation` / + `ReadPopulation` / `Extension` become `value{operation}`. `ValueOperation` stays as a + wire type with only these four variants. +3. Today's `fallback` is a fragment with no `DagInput` leaf. A backend's existing + path that recursively lowers a `QueryExpr` handles every fragment after unwrapping + one `Basic` level and adding a `DagInput` arm (treat the incoming edge as a + materialized table); the dedicated lowerings for `binary` / `value::Project` / + `relational_join` can go. + +- **Wire version 5 → 6**: three fewer payloads; `fallback` is renamed `relational` + and allows `DagInput` leaves; `output_schema` / `intermediate_schema` change from + `SummarySchema` to `Schema` (adding `unique_keys` / `closed` / `table`). +- **Phase granularity**: ingestion / query phase is now one per fragment, and + `with_execution_phases` still assigns it per node. Switching phase inside a + fragment would need a materialization point, and materialization points are + exactly `Ext` nodes (`FinalizeExactAccumulator`, `MaintainPopulation`), so no real + granularity is lost. + +## 7. Other consumers + +| Location | Change | +|---|---| +| `post_asap/cse.rs` | delete. `ASAPOp` derives `PartialEq` + serde, so `share_common_subtrees` in `pre_asap/cse.rs` covers it | +| `dag_export.rs` | delete `build_summary` / `build_summary_hybrid` / `summary_kind_tag`; the one exporter gains an `Ext` arm. The viewer's `node-style.js` drops `KeepPreAsap` / `SummaryBinaryOp` / `ValueOperation` / `RelationalJoin` and adds the `ASAPOp` variant names; the pin test at `dag_export.rs:1843` follows | +| `summary_maintenance_cost/estimator.rs` (the largest, 80 sites) | the `KeepPreAsap` branches through `query_source_selections` / `retained_queries` switch to "largest subtree with no `Ext`" (exactly the §6 fragment); `BinaryOp \| RelationalJoin => "exact_binary"` and `ValueOperation => "value_operation"` fold into the fragment cost. This also fixes the missing `RelationalJoin` arm at `evidence.rs:292`, which can raise `InconsistentOperatorStatistics` | +| `physical_plan_cost_model.rs::estimate_candidate` | `Summary(KeepPreAsap(q))` and `Rewrite(q)` already share `lower_query_physical_dag`; every fragment will go through it | +| `summary_maintenance_lifecycle.rs` | `selected_raw_recompute = matches!(root, KeepPreAsap)` becomes `!contains_ext(root)`; `keep_pre_asap(target)` at `:551` becomes `target` itself | +| `maintained_population.rs` | `KeepPreAsap(source)` becomes `source` itself; `MaintainPopulation`'s `population.matches_input(child)` inspects a `Basic` child directly | +| `exact_composition.rs` | the `ExactOperation::Aggregate` shell becomes a `Basic(Aggregate)` node whose `child` is a hole | +| `RelationalJoin.pruning` | production code never sets `Some`; delete it. Candidate pruning can return later as an `ASAPOp` variant; the `CandidateCompleteness` type stays | + +--- + +# III. Implementation plan + +## 8. Stages and tests + +`main` builds and passes all tests after every stage. + +| Stage | Content | Main changes | +|---|---|---| +| 0 Preparation | `Rc` for `Concat.children`; `rebuild_children` becomes `map_children`; add `Column::plain` | `asap-types`, mechanical | +| 1 Split operators and expressions | §1.2: split `QueryExpr` into `OriginalOp` + `ScalarExpr`; `Predicate` / `ProjectItem` etc. hold `ScalarExpr`; relational children stay `Rc` for now. No semantic change | every place that builds or matches scalars: the three frontends' expression lowering, column resolution in `resolve`, scalar rewrites in `canonicalize`, `scalar_signature.rs`, `column_resolution.rs`. Mechanical | +| 2 Two levels | §1.1, §1.4: add `Operator`, `ASAPOp` (a placeholder with no producer yet), `contains_ext()` and `expect_basic()`; every relational child slot becomes `Rc>`. No semantic change | every place that builds or matches children: the three frontends, `resolve`, `canonicalize`, `pre_asap/cse`, `output_schema`, `dag_export`, all of `asap-aware-mapping`. Each change is the same shape; afterwards the remaining work is local | +| 3 One schema | §2.1: `SummarySchema` → `Schema`, `SummaryFamilyType` → `ColumnType` + `ASAPType`, `PlainField` for `List` / `Struct` elements; delete `lift*` | all of `asap-types` plus schema construction in every crate. The wire format changes only in schema fields; no version bump yet | +| 4 New types alongside old | fill in the `ASAPOp` variants; add the `Ext` arms of `output_schema` and data-state validation (§5), `GuaranteeIndex` (§2.2) and the optimizer-entry check (§3); write `flatten(&SummaryNode) -> Rc`; `compile_executable_dag` and `dag_export` consume the flat tree (planner output goes through `flatten` first). Wire → 6 | `asap-types`, `devtools`, viewer. The planner is untouched, but export and execution already run on the new types | +| 5 Planner switch | §4: candidates and assembly results become `Rc`, `assemble` is rewritten; the cost / lifecycle / estimator code in §7 moves to the new types; delete `flatten` | `asap-aware-mapping`; the largest stage | +| 6 Cleanup | delete `SummaryExpr` / `SummaryNode` / the redundant `ValueOperation` variants / `ExactOperation` / `post_asap/cse.rs`; update `post-asap-ir.md`, `physical-plan-integration.md`, the developer guide and the viewer docs | mostly docs | + +An alternative transition would first make `ValueOperation` wrap a pre-ASAP operator +without changing the structure. Its benefit — export and execution see a flat +structure early — is already delivered by stage 4, which is built on the final types +and so is not throwaway code. That transition is therefore not planned. + +**Tests**: + +- One integration test per problem in §4, asserting the shape of the assembled plan: + 1. `WITH metric AS (SELECT avg(CASE WHEN l_quantity BETWEEN 1 AND 50 THEN 1.0 ELSE 0.0 END) AS in_range FROM lineitem) SELECT in_range, in_range = 1.0 AS ok FROM metric`: + no post-ASAP-only node other than `Ext`, and all three `Project`s are the same variant; + 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem`: + the `child` of the `avg` `Aggregate` and of the KLL `SummaryAgg` is the same `Scan` by `Rc::ptr_eq` (enabled once a binding rule splits measures); + 3. `SELECT approx_distinct(l_partkey) FROM lineitem UNION ALL SELECT approx_distinct(l_suppkey) FROM lineitem`: + each side of the `SetOp` has a `SummaryEstimate`. +- The optimizer entry rejects a tree containing `Ext`. +- Rewrite the 117 `SummaryExpr::` assertions in `sql_to_post_asap.rs` / + `promql_to_post_asap.rs` / `exact_composition.rs` against the new shape. +- Wire 6 round trip: a fragment with a `DagInput` leaf is equal after serialization + and deserialization; a version-5 document is rejected by `deny_unknown_fields`. +- Keep the shapes of the 52 existing tests in `execution_data_state.rs`; only their + construction changes. + +## 9. Out of scope + +- The binding rule that splits a multi-measure `Aggregate` into "exact + summary" + sharing one child (the second half of problem 2 in §4). +- Candidate pruning as an `ASAPOp` variant. +- **Accuracy of summary state** (§2.2): making a readout's guarantee account for + `SummaryMerge` / `Subtract` / `Delete` / `Join`. Either attach an accuracy + descriptor to state (e.g. "ε relative to the L1 norm"), or compose the error along + the state chain at readout. To be done once the planner starts emitting these nodes. +- Folding `ExactComposition` into `Subtree`. Semantically it equals an `Aggregate` + fragment (query time) or `Finalize(SummaryAgg{ExactAggregate})` (ingestion time), + both expressible afterwards; the cost machinery of `composition_plans` and + `CompositionDecision` stays as is until stage 5 is stable. From 755291e10ed035c2301c0fb9b24bd379ee1b3e6c Mon Sep 17 00:00:00 2001 From: Selvomega Date: Mon, 28 Sep 2026 02:32:27 +0000 Subject: [PATCH 02/46] proposal doc updated --- .../proposals/operator-flattening.md | 24 ++++++++++++++----- 1 file changed, 18 insertions(+), 6 deletions(-) diff --git a/docs/design_docs/proposals/operator-flattening.md b/docs/design_docs/proposals/operator-flattening.md index 577ba45fb..7d2eea9df 100644 --- a/docs/design_docs/proposals/operator-flattening.md +++ b/docs/design_docs/proposals/operator-flattening.md @@ -1,14 +1,26 @@ -# Operator flattening +# Sharing Operators Between Pre-ASAP IR and Post-ASAP IR > Status: proposed, not implemented. Problem statement: > [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). Code > references and counts are against `main` at `8acb472`. -**In one sentence**: split `QueryExpr` into original operators (`OriginalOp`) and -scalar expressions (`ScalarExpr`); introduce `Operator { Basic(OriginalOp), Ext(ASAPOp) }` -so that every operator's children are `Rc`; then delete the parallel -post-ASAP IR (`SummaryExpr`, `SummaryNode`, `SummarySchema`, and the relational -variants of `ValueOperation`). +**The idea.** Today the post-ASAP-plan is glued together by different operator types. +This proposal keeps one operator language and makes summary operators extra node kinds in it: any relational operator can sit above a summary, and a summary can read any relational subtree. +Nothing is wrapped and nothing is duplicated. + +``` +Today Proposed +ValueOperation(Project) ← a copy Project + SummaryEstimate Ext(SummaryEstimate) + SummaryAgg(Kll) Ext(SummaryAgg(Kll)) + KeepPreAsap(Scan lineitem) ← a black box Scan lineitem +``` + +Concretely: split `QueryExpr` into operators (`OriginalOp`) and scalar +expressions (`ScalarExpr`), make every operator's children `Rc` with +`Operator = Basic(OriginalOp) | Ext(ASAPOp)`, and delete the parallel post-ASAP +types (`SummaryExpr`, `SummaryNode`, `SummarySchema`, and the relational variants +of `ValueOperation`). **Roadmap**: From ebaba7d1c21b9691c6696d430a04669d34a094b6 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Tue, 29 Sep 2026 03:02:55 +0000 Subject: [PATCH 03/46] design doc updated --- docs/design_docs/proposals/README.md | 3 +- .../proposals/decoupling_op_and_expr.md | 114 ++++ .../proposals/operator-flattening.md | 512 ------------------ .../design_docs/proposals/operator-sharing.md | 428 +++++++++++++++ 4 files changed, 544 insertions(+), 513 deletions(-) create mode 100644 docs/design_docs/proposals/decoupling_op_and_expr.md delete mode 100644 docs/design_docs/proposals/operator-flattening.md create mode 100644 docs/design_docs/proposals/operator-sharing.md diff --git a/docs/design_docs/proposals/README.md b/docs/design_docs/proposals/README.md index 78084caa4..2cebb3b27 100644 --- a/docs/design_docs/proposals/README.md +++ b/docs/design_docs/proposals/README.md @@ -8,4 +8,5 @@ extensions. A design document is not a promise of downstream runtime support. - [ASAPQuery rule coverage](asapquery-rule-coverage.md) - [UnivMon frequency summary](univmon-frequency-summary.md) - [ASAP-aware mapping proposals](asap-aware-mapping/README.md) -- [Operator flattening](operator-flattening.md) +- [Operator sharing](operator-sharing.md) +- [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md new file mode 100644 index 000000000..9d34b4e4a --- /dev/null +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -0,0 +1,114 @@ +# Decoupling Operators From Scalar Expressions + +> Status: proposed, not implemented. Companion to [Operator sharing](operator-sharing.md) +> (same PR): this document splits `QueryExpr`; that one builds the shared operator +> language on the result. Code is referenced by file and function; counts are +> approximate, measured on `main` at `8acb472`. + +**The idea.** `QueryExpr` holds two different kinds of node in one enum. This proposal +splits it into `NonASAPOp` (operators) and `ScalarExpr` (scalar expressions), so the +field a node sits in decides its type. + +``` +Today Proposed +Filter { pred: Rc, Filter { pred: Predicate(Rc), + child: Rc } child: Rc } +``` + +## 1. Problem + +``` +SELECT l_quantity * 2 AS q2 FROM lineitem WHERE l_quantity > 10 + +Project ← operator: outputs a table + cols: [Column(4) * Literal(2)] ← scalar expression: outputs one value per input row + child: Filter ← operator + pred: Column(4) > Literal(10) ← scalar expression + child: Scan lineitem ← operator +``` + +- An **operator** outputs a table. It is a node of the plan: the planner can replace it, + share it, or put a summary under it. +- A **scalar expression** has no table of its own. `Column(4)` means "column 4 of the + input of the operator I sit in"; outside that operator it means nothing. + +Today both are `QueryExpr` variants, told apart only by field position. The code already +separates them, but only by convention: + +- `QueryExpr::output_schema` returns `ScalarHasNoRowSchema` for all 13 scalar variants + (`query_expr.rs`), so `Filter { child: Literal(2) }` compiles and fails at run time. +- `pre_asap/cse.rs` never descends into a scalar, and repeats a "scalar: nothing to do" + arm in each of its three traversals; `canonicalize` likewise never rewrites one. + +## 2. Types + +```rust +pub enum NonASAPOp { + Scan { .. }, Filter { pred: Predicate, child: Rc> }, Project { cols: Vec>, child }, + Aggregate { .. }, Join { .. }, SetOp { .. }, Concat { .. }, Dedup { .. }, Sort { .. }, Limit { .. }, BinaryOp { .. }, + SQLWindowFunc { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, + ScalarBridge(Rc>), // formerly PromqlScalarBridge (the `2` in PromQL `v * 2`) + EvalTimestamp, // PromQL time() +} +pub enum ScalarExpr { + Column(C), Literal(ScalarValue), Compare { .. }, BoolAnd(..), BoolOr(..), Not(..), IsNull(..), IsNotNull(..), + Cast { .. }, InList { .. }, FunctionCall { .. }, Arithmetic { .. }, Case { .. }, CurrentTimestamp, +} +pub struct Predicate(pub Rc>); +pub struct ProjectItem { pub alias: Option, pub expr: ScalarExpr } +``` + +```text +NonASAPOp +├─ children: Rc> (Rc> after operator sharing) +└─ scalar fields: Predicate / ProjectItem / SortKey / ScalarBridge / ... + └─ ScalarExpr + └─ children: ScalarExpr only, never an operator +``` + +- **Naming**: `NonASAPOp` is named for [Operator sharing](operator-sharing.md), where it + becomes the non-ASAP category of `Operator`. `QueryExpr` goes away. +- **Scalar fields**: `Filter.pred`, `Join.pred`, `Aggregate.having` (`Predicate`); + `Project.cols` (`ProjectItem`); `Sort` / `SQLWindowFunc` sort keys (`SortKey`); + `SQLWindowFunc.args`; `PromqlRelabel.value`. +- **Borderline variants** go by position, not by look. `PromqlScalarBridge`, + `EvalTimestamp`, `PromqlScalarFromVector` and `PromqlVectorFromScalar` sit in operator + position with a row schema (e.g. a `BinaryOp` operand) → `NonASAPOp`. `CurrentTimestamp` + (SQL `NOW()`) is produced only by scalar lowering (`df_expr_to_unresolved`); its only + operator-position use is in a unit test → `ScalarExpr`. +- **Gain**: neither a scalar in operator position nor an operator in scalar position is + expressible, and `ScalarHasNoRowSchema` is deleted. + +## 3. Changes + +| Location | Change | +|---|---| +| frontend expression lowering (`df_expr_to_unresolved`, PromQL `walk`) | scalar positions build `ScalarExpr`, operator positions `NonASAPOp` | +| `resolve`, `column_resolution.rs` | separate operator and scalar resolvers; a scalar resolves against its operator's input schema | +| `canonicalize`, `pre_asap/cse.rs` | the "scalar: nothing to do" arms go; scalars are hashed as plain data | +| `scalar_signature.rs`, `infer_expr_type` | take `ScalarExpr` | +| `QueryExpr::output_schema` | becomes `NonASAPOp::output_schema`; the scalar arms and `ScalarHasNoRowSchema` go | + +## 4. Implementation and tests + +This is stage 1 of the joint plan ([Operator sharing §8](operator-sharing.md#8-stages-and-tests)): +children stay `Rc`; operator sharing widens them to `Rc` in its +stage 2. + +**No wire change.** The `fallback` payload serializes a `QueryExpr`, externally tagged. +Variant names are kept, so a tree serializes the same; `ScalarBridge` keeps the name +`PromqlScalarBridge` with `#[serde(rename)]`. + +- Existing tests pass unchanged apart from construction syntax. +- Tests that place a scalar in operator position no longer compile and are rewritten or + deleted: the `CurrentTimestamp` unit test, the `ScalarHasNoRowSchema` tests, and the + `executable_dag.rs` tests using `QueryExpr::Literal` as a `fallback` expression. + +## 5. Limits + +The split relies on no scalar containing an operator. That holds today: SQL +`IN (SELECT …)` / `EXISTS` in a filter lower to a semi-join (`lower_filter`), and every +other subquery-valued expression is rejected (`frontend-sql/src/sql/expr.rs`). Supporting +a scalar subquery (`WHERE x > (SELECT avg(x) …)`) would add `ScalarExpr::Subquery(Rc<..>)`, +make the two types mutually recursive, and require CSE and the planner to look inside +scalars. diff --git a/docs/design_docs/proposals/operator-flattening.md b/docs/design_docs/proposals/operator-flattening.md deleted file mode 100644 index 7d2eea9df..000000000 --- a/docs/design_docs/proposals/operator-flattening.md +++ /dev/null @@ -1,512 +0,0 @@ -# Sharing Operators Between Pre-ASAP IR and Post-ASAP IR - -> Status: proposed, not implemented. Problem statement: -> [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). Code -> references and counts are against `main` at `8acb472`. - -**The idea.** Today the post-ASAP-plan is glued together by different operator types. -This proposal keeps one operator language and makes summary operators extra node kinds in it: any relational operator can sit above a summary, and a summary can read any relational subtree. -Nothing is wrapped and nothing is duplicated. - -``` -Today Proposed -ValueOperation(Project) ← a copy Project - SummaryEstimate Ext(SummaryEstimate) - SummaryAgg(Kll) Ext(SummaryAgg(Kll)) - KeepPreAsap(Scan lineitem) ← a black box Scan lineitem -``` - -Concretely: split `QueryExpr` into operators (`OriginalOp`) and scalar -expressions (`ScalarExpr`), make every operator's children `Rc` with -`Operator = Basic(OriginalOp) | Ext(ASAPOp)`, and delete the parallel post-ASAP -types (`SummaryExpr`, `SummaryNode`, `SummarySchema`, and the relational variants -of `ValueOperation`). - -**Roadmap**: - -| Part | Sections | Question answered | -|---|---|---| -| I. New IR | §1 Types, §2 Per-node information | Which node kinds exist, and what each carries | -| II. Upstream and downstream changes | §3 Frontends and entry → §4 Planner → §5 Validation → §6 Export → §7 Other consumers | How each stage changes, in data-flow order | -| III. Implementation plan | §8 Stages and tests, §9 Out of scope | How to land it, and what is deliberately left out | - ---- - -# I. New IR - -## 1. Types - -### 1.1 Overview - -``` -Operator operator: one node of the DAG; every child is Rc> -├─ Basic(OriginalOp) original operator: today's relational QueryExpr variants (Scan / Filter / Join / Aggregate / SetOp …) -└─ Ext(ASAPOp) ASAP operator: summary operators (SummaryAgg / SummaryEstimate …) - -ScalarExpr scalar expression: only appears inside operator fields (predicates, SELECT lists); never a node -``` - -- **Naming**: the suffix says what a type is — `*Op` is an operator, `*Expr` is an - expression. `OriginalOp` holds the operators pre-ASAP already has, as opposed to - the extension `ASAPOp`. The name `QueryExpr` goes away. -- **The generic `C`**: the existing column-reference stage of `QueryExpr` — - `ColumnRef` (by name) out of the frontends, `ColumnId` (by position) after - `resolve_root`. A parent and its children must be in the same stage, so all - three types carry the same `C`. - -### 1.2 `OriginalOp` and `ScalarExpr`: splitting `QueryExpr` - -Today one `QueryExpr` enum holds two kinds of things: - -```sql -SELECT l_quantity * 2 AS q2 FROM lineitem WHERE l_quantity > 10 -``` -``` -Project { cols: [ProjectItem { expr: Arithmetic(Column(4) * Literal(2)) }], ← scalar expression: computes one value per row - child: Filter { pred: Predicate(Compare(Column(4) > Literal(10))), ← scalar expression - child: Scan lineitem } } ← relational operator: rows in, rows out -``` - -`Filter.child` (rows) and the `Compare` inside `Filter.pred` (a value) are both -`QueryExpr`; only the field position tells them apart. The code already knows -which variants are scalars: `output_schema` returns `ScalarHasNoRowSchema` for all -of them (`types/src/pre_asap/query_expr.rs:1447`). Split them into two types: - -```rust -pub enum OriginalOp { - Scan { .. }, Filter { pred: Predicate, child: Rc> }, Project { cols: Vec>, child }, - Aggregate { .. }, Join { .. }, SetOp { .. }, Concat { .. }, Dedup { .. }, Sort { .. }, Limit { .. }, BinaryOp { .. }, - SQLWindowFunc { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, - ScalarBridge(Rc>), // formerly PromqlScalarBridge: a scalar in operator position (the `2` in PromQL `v * 2`, a bare scalar query) - EvalTimestamp, // PromQL time(): operator position, one row, one column -} -pub enum ScalarExpr { // children are ScalarExpr only - Column(C), Literal(ScalarValue), Compare { .. }, BoolAnd(..), BoolOr(..), Not(..), IsNull(..), IsNotNull(..), - Cast { .. }, InList { .. }, FunctionCall { .. }, Arithmetic { .. }, Case { .. }, CurrentTimestamp, -} -pub struct Predicate(pub Rc>); -pub struct ProjectItem { pub alias: Option, pub expr: ScalarExpr } -``` - -- **Ambiguous variants**: `PromqlScalarBridge`, `EvalTimestamp`, - `PromqlScalarFromVector` and `PromqlVectorFromScalar` appear in operator position - and have a row schema, so they belong to `OriginalOp`. `CurrentTimestamp` (SQL - `NOW()`) is used both ways: it has a row schema (`query_expr.rs:1388`) and takes - part in scalar type inference (`:1674`, e.g. `WHERE ts > NOW()`). It goes into - `ScalarExpr` and is written `ScalarBridge(CurrentTimestamp)` in operator - position. **Open**: confirm how the frontends place `CurrentTimestamp` in - operator position. -- **Benefits**: a scalar in operator position (`Basic(Column(3))`) is no longer - expressible; the `ScalarHasNoRowSchema` runtime error disappears; `canonicalize` - and `pre_asap/cse` already walk only the relational skeleton and treat scalars as - opaque data (see the `pre_asap/cse.rs` module docs) — the types now say so. - -### 1.3 `ASAPOp` - -```rust -pub enum ASAPOp { - SummaryAgg { child: Rc>, family: ASAPType, input, reduction, grouping, // family was SummaryFamilyType, "never Plain" by convention; now by type - guarantee: Option }, // Some only for the ExactAggregate family - SummaryEstimate { child, query: SketchQuery, - guarantee: Option }, // None = the AccuracyModel has no error model for this family - SummaryMerge { children: Vec>>, timing: ExecutionTiming }, - SummarySubtract { left, right }, - SummaryDelete { child, key: C }, // was ColumnRef; now C like everything else - SummaryJoin { outer, inner, key: C, family: ASAPType }, - // The four summary-specific ValueOperation variants, lifted to the top level instead of { child, operation, timing } - FinalizeExactAccumulator { child, timing: ExecutionTiming }, - MaintainPopulation { child, population }, // always ingestion time - ReadPopulation { child, readout }, // always query time - Extension { child, name: String }, -} -impl ASAPOp { pub fn guarantee(&self) -> Option<&ResultGuarantee>; } // only the two variants above store one; see §2.2 -``` - -**Deleted post-ASAP types** — each is expressed by an existing `OriginalOp` variant: - -| Deleted | Expressed as | -|---|---| -| `SummaryExpr::KeepPreAsap(q)` | no wrapper: the subtree `q` is itself `Basic(..)` | -| `ValueOperation::{Project, Filter, Sort, Limit}` | `OriginalOp::{Project, Filter, Sort, Limit}` | -| `SummaryExpr::BinaryOp`, `SummaryExpr::RelationalJoin` | `OriginalOp::BinaryOp`, `OriginalOp::Join` | -| `ValueOperation::Exact(Aggregate)`, `ExactOperation` | `OriginalOp::Aggregate` | -| `SummaryNode` (the node wrapper) | not needed; per-node information is covered in §2 | - -### 1.4 Child slots - -```rust -// Today // After -Filter { Filter { - pred: Predicate(Rc), // scalar pred: Predicate(Rc), - child: Rc, // rows child: Rc, // either Basic(..) or Ext(SummaryEstimate ..) -} } -``` - -**Rule**: every `OriginalOp` field that holds input rows has type `Rc>`. -There are about 25 such fields — `Filter.child`, `Join.left` / `Join.right`, both -sides of `SetOp`, `Concat.children`, `BinaryOp.lhs` / `BinaryOp.rhs`, the `child` of -each `Promql*` operator, and so on. - -**Effect**: an ASAP operator can sit directly under any original operator. Both -sides of a `SetOp`, for example, can be `SummaryEstimate`s. - -**Exception: `Concat.children`**. Today it is `Vec`: branches are stored -by value and have no `Rc` identity. The planner's search and assembly (§4) -identify nodes by `Rc` pointer — targets are registered by pointer and holes are -recognized with `Rc::ptr_eq` — so a `Concat` branch can be neither a target nor a -hole. This affects, for example, the branches SQL `ROLLUP` and PromQL -`histogram_quantiles` are lowered into. As `Vec>` it is handled like -any other child slot and §4 needs no special case (§8 stage 0). - -## 2. Per-node information - -A post-ASAP node carries three pieces of information today. After the change: - -| Field | Meaning | Today | After | -|---|---|---|---| -| `schema` | which columns the node outputs, and their types | stored on every `SummaryNode` | not stored; `output_schema()` computes it on demand (§2.1) | -| `guarantee` | how far the node's output value can be off (metric, bound, failure probability, provenance) | stored on every `SummaryNode` | stored on `SummaryAgg` / `SummaryEstimate`; derived by `GuaranteeIndex` for everything else (§2.2) | -| `timing` | whether the node runs at ingestion time or at query time | stored on the `BinaryOp` / `ValueOperation` / `SummaryMerge` variants | not stored on `OriginalOp` (§5); kept on `SummaryMerge` and `FinalizeExactAccumulator` | - -### 2.1 Schema: two types become one - -A schema is a node's list of output columns: name, type, nullability. There are two today: - -| | pre-ASAP | post-ASAP | -|---|---|---| -| Type | `Schema { columns: Vec, time_index, unique_keys, closed }` | `SummarySchema { fields: Vec, time_index }` | -| Column type | `Column.dtype: DataType` — plain values only (`Int64`, `Float64`, `Utf8`, …) | `SummaryField.dtype: SummaryFamilyType` — a plain value `Plain(DataType)` or summary state (`Sketch(Kll, …)`, …) | -| Where it lives | not stored; `QueryExpr::output_schema()` computes it | stored on every `SummaryNode` | - -In the new IR both kinds of operator share one tree, so `Operator::output_schema()` -must describe any node, and a column type must be able to hold either a plain -value or summary state: - -``` -SummaryEstimate(Quantile .99) → [p99: DataType(Float64)] - SummaryAgg(Kll) → [state: ASAPType(Sketch(Kll, k=269))] - Scan lineitem → [l_orderkey: DataType(Int64), l_quantity: DataType(Float64), …] -``` - -So keep `Schema`, delete `SummarySchema` / `SummaryField`, and give `Column.dtype` a new type, `ColumnType`: - -```rust -pub enum ColumnType { - DataType(DataType), // a plain value; the existing DataType, unchanged (Int64, Float64, Utf8, …) - ASAPType(ASAPType), // ASAP state -} -pub enum ASAPType { // SummaryFamilyType without its Plain variant - ExactAggregate(ExactKind, ExactParams), Sketch(SketchKind, GroupingStrategy), - Sample(SamplingKind, SamplingParams), Wavelet(WaveletKind, WaveletParams), StatModel(StatModelKind, StatModelParams), -} - -pub struct Column { pub name, pub dtype: ColumnType, pub nullable, pub table: Option } -impl Column { - pub fn plain(name, DataType) -> Self; // build a plain column (frontends, catalogs) - pub fn plain_dtype(&self) -> Option<&DataType>; // None for an ASAP-state column - pub fn expect_plain_dtype(&self) -> &DataType; // frontends and scalar type inference: a state column is a bug, panic -} -``` - -- **Naming**: as in §1, `DataType` / `ASAPType` mirror `OriginalOp` / `ASAPOp`. - The name `SummaryFamilyType` goes away: once every column carries it, a plain - column reading `SummaryFamilyType::Plain(Float64)` is a misnomer. With two - levels, "is this column state?" is a check on the outer variant only — exactly - the split `check_plain_operands` needs. -- **`DataType` itself is unchanged.** It remains part of the scalar type system - (the type of a `Literal`, the result type of `Arithmetic`, catalog column - declarations); it now sits inside `ColumnType::DataType(..)`. Code that reads - `Column.dtype` expecting a `DataType` unwraps one more level: code that only - ever sees plain columns (frontends, scalar type inference) uses - `expect_plain_dtype()`; code that may see state (planner, validation, export) - uses `plain_dtype()` or matches. A scalar expression never legitimately reads a - state column — `check_plain_operands` (§5) rejects such a DAG first. -- **Elements of nested types**: `DataType::List { element: Box }` and - `DataType::Struct { fields: Vec }` have `Column` elements. With - `Column.dtype: ColumnType` an element could become ASAP state (a "list of KLL - sketches"), which is not intended. These two use a plain-only field instead: - - ```rust - pub struct PlainField { pub name: String, pub dtype: DataType, pub nullable: bool } // nested fields are already unqualified (List's doc comment), so no table - DataType::List { element: Box } - DataType::Struct { fields: Vec } - ``` - - `DataType::Map { key: Box, value: Box, .. }` already uses - `DataType` only and is unchanged. There are 36 construction or match sites of - `DataType::{List, Struct, Map}`. -- **Why keep `Schema`**: `ColumnType::DataType(..)` represents every plain type, so - nothing is lost, and `Schema` additionally has `unique_keys` / `closed` / - `Column.table`, which merging the other way would drop. -- **Computed on demand**: schemas are no longer stored on nodes; `Operator::output_schema()` - computes them. `OriginalOp` keeps today's `QueryExpr::output_schema()` logic, and - `ASAPOp` follows the table below. These rules are currently spread over - `asap-aware-mapping/src/replacement.rs` and `types/src/post_asap/execution_data_state.rs`; - they move into `asap-types`. -- **Deleted conversions**: `lift()` (`replacement.rs:3352`) and `lift_plain` - (`execution_data_state.rs:725`), which turn a `Schema` into a `SummarySchema`, and - the reverse `plain_schema` (`execution_data_state.rs:707`). With one schema type - there is nothing to convert. -- **Size of the change**: 7 direct `Column { .. }` literals; the deleted - `SummarySchema {}` (36) and `SummaryField {}` (19) literals. In non-test code - `SummaryFamilyType` appears 197 times (42 of them with `Plain(`) and `.dtype` is - read at 67 sites; all are rewritten against the new types. - -| `ASAPOp` | Output schema | -|---|---| -| `SummaryAgg` | the grouping columns as-is + one `ASAPType(family)` column | -| `SummaryEstimate` | the row schema determined by the `SketchQuery` | -| `SummaryMerge` / `Subtract` / `Delete` / `Join` | one `ASAPType(family)` column | -| `FinalizeExactAccumulator` | the child's schema, with each `ASAPType(ExactAggregate ..)` column turned into the matching `DataType(..)` | -| `MaintainPopulation` / `ReadPopulation` | existing implementation | - -### 2.2 Guarantee: some stored, some derived - -| Node | Where its guarantee comes from | -|---|---| -| `SummaryAgg`, `SummaryEstimate` | **stored in a field**. Computed at binding time by the `AccuracyModel` from the summary's parameters and the evidence; it cannot be recovered afterwards, and selection needs it during search to decide whether a candidate meets the accuracy target (`replacement.rs:5831`) | -| `OriginalOp` (`Project`, `Join`, …), `FinalizeExactAccumulator` | composed from the children by rule (`relational_join_guarantee`; `exact_operation_rule` in `accuracy/composition.rs`). A subtree with no `Ext` descendant is exact | -| `MaintainPopulation`, `ReadPopulation` | always exact | -| `SummaryMerge` / `Subtract` / `Delete` / `Join` | output state, so no guarantee of their own — but they change the error of a later readout (see below) | - -The derived part goes into one table: - -```rust -/// One traversal, keyed by node pointer — the same approach as the ExecutionDataStateAssignment -/// returned by validate_execution_data_states. -pub struct GuaranteeIndex(HashMap<*const Operator, ResultGuarantee>); -pub fn guarantee_index(root: &Rc) -> GuaranteeIndex; -``` - -- The composition rules also depend on the `AccuracyModel`, so the planner must - compute `GuaranteeIndex` with the same model and hand it over together with the - assembled root (`assemble_selected_dag`). `compile_executable_dag` uses it to fill - `ExecutableDagNode.guarantee`. -- Not chosen: wrapping every node to store a guarantee — frontends and `resolve` - would then have to build the wrapper too. -- **Gap for state nodes**: `SummaryMerge` / `Subtract` / `Delete` / `Join` change the - error of a later readout. For a CMS, `a − b` has error bound `ε·(‖a‖₁ + ‖b‖₁)`, so - the relative error is large when `a` and `b` are close. A readout's guarantee is - computed today as if its input were a sketch built directly by a `SummaryAgg`, - ignoring any intermediate operation. The planner on `main` never emits these four - nodes (they are built only in tests, and `SummaryMerge` is documented as inserted by - a deployment), so the gap is latent. Not addressed here (§9). - ---- - -# II. Upstream and downstream changes - -## 3. Frontends and the optimizer entry - -Trees built by the frontends and by `resolve` contain only `Basic`. They access -children through `expect_basic()`; before a tree reaches the optimizer, the entry -checks it once and rejects any `Ext`: - -```rust -impl Operator { - pub fn contains_ext(&self) -> bool; - pub fn expect_basic(&self) -> &OriginalOp; // an Ext here is a bug: panic -} -``` - -**Where the check goes**: on `main` the optimizer entries are `search_workload` / -`search_workload_with` / `search_workload_with_targets`. They do not return -`Result`, so the check starts as `assert!(!root.contains_ext())` at each entry. If -the single `ParsedWorkload` entry point (#429) lands on `main` first, the check -moves into `ParsedWorkload::new` and returns an error (it already checks the entry -count). - -**Why not a compile-time guarantee** (an associated type `Ext` on `ColState`, with -`ColumnRef::Ext = Never`): - -| | Compile time | Run time (this proposal) | -|---|---|---| -| Types | one more associated type on `ColState`, plus an empty enum | the `Ext` variant holds `ASAPOp` directly | -| What it catches | only frontend code before `resolve`. A frontend already returns a `ColumnId` tree, where `Ext` is allowed | **every input to the optimizer**: frontend code after `resolve`, deserialized plans, third-party frontends, hand-written test IR | -| Frontends matching children | `basic()` is total | `expect_basic()`, which can in principle panic | - -The compile-time version protects only a small stretch inside the frontends; the -runtime check sits at the entry, covers more, and keeps the types simpler. So one -`Operator` type serves both pre- and post-ASAP, told apart by the entry check -rather than by a type parameter. - -## 4. Planner: search and assembly - -**Candidates**: three shapes become two. - -```rust -pub enum Replacement { - Subtree(Rc), // formerly Summary(Rc) and Rewrite(Rc); "introduces a summary" = contains_ext() - ExactComposition { .. }, // unchanged, except its plan becomes Rc instead of Rc -} -``` - -**Holes**: a subtree of a candidate that is `Rc::ptr_eq` to some target is a hole, -left for that target's own choice to fill. A candidate that needs a specific -implementation of a child (for example a `SummaryAgg` that needs an exact -accumulator value, as in `realize_temporal_average`) inlines the child as a new node; -the pointer differs, so it is not a hole. `realize_child_with` changes from -"realize the child recursively, else `keep_pre_asap`" to "else leave a hole". - -**Assembly**: one generic rule replaces `assemble_residual`. - -```rust -fn assemble(&self, t: &Rc) -> Rc { - memo by ptr; // a shared child stays one Rc — today's assembled_nodes - let body = match self.chosen(t) { - Some(Subtree(r)) => r, // Rewrites are assembled further down too (today replacement.rs:4571 wraps the whole rewrite in keep_pre_asap and ignores the choices below) - Some(ExactComposition{..}) => composition.plan, - None => t, // no candidate chosen: keep the node, recurse into its children — what assemble_residual does for only four operators - }; - body.map_children(|c| if is_target(c) { self.assemble(c) } else { c }) -} -``` - -After assembly, run the data-state validation (§5) once; if filling a hole is -illegal (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`), that hole falls -back to its original subtree. This generalizes the fallback `relink_summary` does -today for `Aggregate` only. Finally compute `GuaranteeIndex` (§2.2). - -- **`map_children`**: reuse `rebuild_children` from `pre_asap/cse.rs:576` as - `Operator::map_children`, dispatching `Basic` to `OriginalOp::map_children` and - `Ext` to `ASAPOp::map_children`. -- **Deleted**: `assemble_residual`, `relink_summary`, the `query_time_nested_sum` / - `contains_aggregate` special cases, `keep_pre_asap` / `keep_pre_asap_rc`, and the - branch of `finalize_exact_accumulator` that builds an empty schema for `KeepPreAsap`. - -How the three problems in #468 are resolved: - -| Problem | Resolution | -|---|---| -| 1. One operator, two spellings: a `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside it | one set of types; `assemble` no longer builds `ValueOperation`s | -| 2. `KeepPreAsap` is opaque: nothing outside can reference the `Scan` inside, so an exact aggregate and a sketch cannot share one scan | `Scan` is no longer wrapped; `SummaryAgg{child: scan}` and `Aggregate{child: scan}` can point to the same `Rc`. Splitting a multi-measure `Aggregate` into `avg` + KLL is a binding-rule change; this proposal only makes the split expressible | -| 3. `SetOp` and similar operators have no post-ASAP counterpart and must stay inside `KeepPreAsap`, so no subtree below them can use a summary | `SetOp` takes the `None => t` path and both children are assembled independently | - -## 5. Validation: execution data states - -Today's rule `produced_data_state(KeepPreAsap) = None` (decided by the edge that -reaches the node) extends to every `OriginalOp`: - -| Node | Produced state | -|---|---| -| `OriginalOp` | none of its own; assigned by the consuming edge (`QUERY_ROWS` at the root) and passed down unchanged | -| `SummaryAgg` | `{the child's timing, SummaryState}`; ingestion time when the child is an `OriginalOp` | -| other `ASAPOp` | unchanged | - -The `KeepPreAsap` / `BinaryOp` / `ValueOperation` / `RelationalJoin` arms of -`validate_execution_data_states` merge into one `Basic` arm: - -- Pass the node's state to each child; an `Ext` child is checked against the - `Ext` edge rules (e.g. `SummaryEstimate` only on the query side). -- `check_plain_operands` stays: every column an `OriginalOp` references must be - `ColumnType::DataType`; `Project` / `Filter` / `Sort` / `Limit` may pass - `ExactAggregate` columns through (today's exception). This is what rejects - `Project(Ext(SummaryAgg))`, i.e. projecting sketch state as if it were rows. -- The extra ingestion-side constraints on `BinaryOp` (arithmetic only, identical - schemas on both sides, exactly one `Float64`) move into this arm, per variant. -- `AmbiguousKeepPreAsap` is renamed `AmbiguousDataState`; its meaning is unchanged. - -## 6. Export: fragments in the executable DAG - -`ExecutableDag` has four payloads for original operators today — `fallback{expression: QueryExpr}`, -`binary`, `value` and `relational_join`. Afterwards there is one: - -```rust -ExecutableOperatorPayload::Relational { - /// An Operator with no Ext (checked at construction). Leaves are real Scans, or - /// Scan { source: Source::DagInput { role } } standing for an incoming edge whose - /// schema is that edge's intermediate_schema. - expression: Operator, -} -``` - -`Source` gains a variant `DagInput { role: EdgeRole }`. `compile_executable_dag` works as follows: - -1. Starting from the root, take the **largest connected subtree containing no - `Ext`** as one fragment node; wherever the fragment meets an `Ext`, cut an edge - and put a `DagInput` leaf in the fragment. -2. Each `Ext` node maps one-to-one onto the existing `summary_agg` / - `summary_estimate` / … payloads; `FinalizeExactAccumulator` / `MaintainPopulation` / - `ReadPopulation` / `Extension` become `value{operation}`. `ValueOperation` stays as a - wire type with only these four variants. -3. Today's `fallback` is a fragment with no `DagInput` leaf. A backend's existing - path that recursively lowers a `QueryExpr` handles every fragment after unwrapping - one `Basic` level and adding a `DagInput` arm (treat the incoming edge as a - materialized table); the dedicated lowerings for `binary` / `value::Project` / - `relational_join` can go. - -- **Wire version 5 → 6**: three fewer payloads; `fallback` is renamed `relational` - and allows `DagInput` leaves; `output_schema` / `intermediate_schema` change from - `SummarySchema` to `Schema` (adding `unique_keys` / `closed` / `table`). -- **Phase granularity**: ingestion / query phase is now one per fragment, and - `with_execution_phases` still assigns it per node. Switching phase inside a - fragment would need a materialization point, and materialization points are - exactly `Ext` nodes (`FinalizeExactAccumulator`, `MaintainPopulation`), so no real - granularity is lost. - -## 7. Other consumers - -| Location | Change | -|---|---| -| `post_asap/cse.rs` | delete. `ASAPOp` derives `PartialEq` + serde, so `share_common_subtrees` in `pre_asap/cse.rs` covers it | -| `dag_export.rs` | delete `build_summary` / `build_summary_hybrid` / `summary_kind_tag`; the one exporter gains an `Ext` arm. The viewer's `node-style.js` drops `KeepPreAsap` / `SummaryBinaryOp` / `ValueOperation` / `RelationalJoin` and adds the `ASAPOp` variant names; the pin test at `dag_export.rs:1843` follows | -| `summary_maintenance_cost/estimator.rs` (the largest, 80 sites) | the `KeepPreAsap` branches through `query_source_selections` / `retained_queries` switch to "largest subtree with no `Ext`" (exactly the §6 fragment); `BinaryOp \| RelationalJoin => "exact_binary"` and `ValueOperation => "value_operation"` fold into the fragment cost. This also fixes the missing `RelationalJoin` arm at `evidence.rs:292`, which can raise `InconsistentOperatorStatistics` | -| `physical_plan_cost_model.rs::estimate_candidate` | `Summary(KeepPreAsap(q))` and `Rewrite(q)` already share `lower_query_physical_dag`; every fragment will go through it | -| `summary_maintenance_lifecycle.rs` | `selected_raw_recompute = matches!(root, KeepPreAsap)` becomes `!contains_ext(root)`; `keep_pre_asap(target)` at `:551` becomes `target` itself | -| `maintained_population.rs` | `KeepPreAsap(source)` becomes `source` itself; `MaintainPopulation`'s `population.matches_input(child)` inspects a `Basic` child directly | -| `exact_composition.rs` | the `ExactOperation::Aggregate` shell becomes a `Basic(Aggregate)` node whose `child` is a hole | -| `RelationalJoin.pruning` | production code never sets `Some`; delete it. Candidate pruning can return later as an `ASAPOp` variant; the `CandidateCompleteness` type stays | - ---- - -# III. Implementation plan - -## 8. Stages and tests - -`main` builds and passes all tests after every stage. - -| Stage | Content | Main changes | -|---|---|---| -| 0 Preparation | `Rc` for `Concat.children`; `rebuild_children` becomes `map_children`; add `Column::plain` | `asap-types`, mechanical | -| 1 Split operators and expressions | §1.2: split `QueryExpr` into `OriginalOp` + `ScalarExpr`; `Predicate` / `ProjectItem` etc. hold `ScalarExpr`; relational children stay `Rc` for now. No semantic change | every place that builds or matches scalars: the three frontends' expression lowering, column resolution in `resolve`, scalar rewrites in `canonicalize`, `scalar_signature.rs`, `column_resolution.rs`. Mechanical | -| 2 Two levels | §1.1, §1.4: add `Operator`, `ASAPOp` (a placeholder with no producer yet), `contains_ext()` and `expect_basic()`; every relational child slot becomes `Rc>`. No semantic change | every place that builds or matches children: the three frontends, `resolve`, `canonicalize`, `pre_asap/cse`, `output_schema`, `dag_export`, all of `asap-aware-mapping`. Each change is the same shape; afterwards the remaining work is local | -| 3 One schema | §2.1: `SummarySchema` → `Schema`, `SummaryFamilyType` → `ColumnType` + `ASAPType`, `PlainField` for `List` / `Struct` elements; delete `lift*` | all of `asap-types` plus schema construction in every crate. The wire format changes only in schema fields; no version bump yet | -| 4 New types alongside old | fill in the `ASAPOp` variants; add the `Ext` arms of `output_schema` and data-state validation (§5), `GuaranteeIndex` (§2.2) and the optimizer-entry check (§3); write `flatten(&SummaryNode) -> Rc`; `compile_executable_dag` and `dag_export` consume the flat tree (planner output goes through `flatten` first). Wire → 6 | `asap-types`, `devtools`, viewer. The planner is untouched, but export and execution already run on the new types | -| 5 Planner switch | §4: candidates and assembly results become `Rc`, `assemble` is rewritten; the cost / lifecycle / estimator code in §7 moves to the new types; delete `flatten` | `asap-aware-mapping`; the largest stage | -| 6 Cleanup | delete `SummaryExpr` / `SummaryNode` / the redundant `ValueOperation` variants / `ExactOperation` / `post_asap/cse.rs`; update `post-asap-ir.md`, `physical-plan-integration.md`, the developer guide and the viewer docs | mostly docs | - -An alternative transition would first make `ValueOperation` wrap a pre-ASAP operator -without changing the structure. Its benefit — export and execution see a flat -structure early — is already delivered by stage 4, which is built on the final types -and so is not throwaway code. That transition is therefore not planned. - -**Tests**: - -- One integration test per problem in §4, asserting the shape of the assembled plan: - 1. `WITH metric AS (SELECT avg(CASE WHEN l_quantity BETWEEN 1 AND 50 THEN 1.0 ELSE 0.0 END) AS in_range FROM lineitem) SELECT in_range, in_range = 1.0 AS ok FROM metric`: - no post-ASAP-only node other than `Ext`, and all three `Project`s are the same variant; - 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem`: - the `child` of the `avg` `Aggregate` and of the KLL `SummaryAgg` is the same `Scan` by `Rc::ptr_eq` (enabled once a binding rule splits measures); - 3. `SELECT approx_distinct(l_partkey) FROM lineitem UNION ALL SELECT approx_distinct(l_suppkey) FROM lineitem`: - each side of the `SetOp` has a `SummaryEstimate`. -- The optimizer entry rejects a tree containing `Ext`. -- Rewrite the 117 `SummaryExpr::` assertions in `sql_to_post_asap.rs` / - `promql_to_post_asap.rs` / `exact_composition.rs` against the new shape. -- Wire 6 round trip: a fragment with a `DagInput` leaf is equal after serialization - and deserialization; a version-5 document is rejected by `deny_unknown_fields`. -- Keep the shapes of the 52 existing tests in `execution_data_state.rs`; only their - construction changes. - -## 9. Out of scope - -- The binding rule that splits a multi-measure `Aggregate` into "exact + summary" - sharing one child (the second half of problem 2 in §4). -- Candidate pruning as an `ASAPOp` variant. -- **Accuracy of summary state** (§2.2): making a readout's guarantee account for - `SummaryMerge` / `Subtract` / `Delete` / `Join`. Either attach an accuracy - descriptor to state (e.g. "ε relative to the L1 norm"), or compose the error along - the state chain at readout. To be done once the planner starts emitting these nodes. -- Folding `ExactComposition` into `Subtree`. Semantically it equals an `Aggregate` - fragment (query time) or `Finalize(SummaryAgg{ExactAggregate})` (ingestion time), - both expressible afterwards; the cost machinery of `composition_plans` and - `CompositionDecision` stays as is until stage 5 is stable. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md new file mode 100644 index 000000000..4ee165901 --- /dev/null +++ b/docs/design_docs/proposals/operator-sharing.md @@ -0,0 +1,428 @@ +# Sharing Operators Between Pre-ASAP IR and Post-ASAP IR + +> Status: proposed, not implemented. Problem statement: +> [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). Implementation +> starts after the single entry point and the pluggable pass (#429, #430) land on +> `main`. Builds on [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) +> (same PR), which splits `QueryExpr` into `NonASAPOp` and `ScalarExpr`. Code is +> referenced by file and function; counts are approximate, measured on `main` at `8acb472`. + +**The idea.** Today a post-ASAP plan is glued together from two sets of operator types. +This proposal keeps one operator language and makes summary operators extra node kinds in it: any relational operator can sit above a summary, and a summary can read any relational subtree. +Nothing is wrapped and nothing is duplicated. + +``` +Today Proposed +ValueOperation(Project) ← a copy NonASAP(Project) + SummaryEstimate ASAP(SummaryEstimate) + SummaryAgg(Kll) ASAP(SummaryAgg(Kll)) + KeepPreAsap(Scan lineitem) ← a black box NonASAP(Scan lineitem) +``` + +| Part | Sections | +|---|---| +| I. New IR | §1 Types, §2 Per-node information | +| II. Changes, in data-flow order | §3 Entry → §4 Planner → §5 Timing → §6 Export → §7 Other consumers | +| III. Implementation | §8 Stages and tests, §9 Out of scope, §10 Open questions | + +--- + +# I. New IR + +## 1. Types + +### 1.1 Overview + +Operator attributes differ in how widely they apply. Each is defined at the level +that matches its breadth: + +| Applies to | Examples | Defined as | +|---|---|---| +| every operator | children, schema, timing, guarantee | methods of `Operator`; the values may be stored per variant | +| one category | for all `NonASAP` operators, timing is derived from the consuming edge, and guarantee from the children | a variant of `Operator` | +| one operator | `Aggregate.measures`, `SummaryAgg.family` | fields of that variant | + +```rust +pub enum Operator { // category + NonASAP(NonASAPOp), // today's relational QueryExpr variants (§1.2) + ASAP(ASAPOp), // summary operators (§1.3) +} + +impl Operator { // every operator + pub fn children(&self) -> Vec<&Rc>>; + pub fn map_children(&self, f: impl FnMut(&Rc>) -> Rc>) -> Self; + pub fn output_schema(&self) -> Result; // computed (§2.1) + pub fn timing(&self) -> &Slot; // §2.3 + pub fn guarantee(&self) -> &Slot>; // §2.2; Set(None): unknown, never read as exact + pub fn with_timing(&self, timing: ExecutionTiming) -> Self; + pub fn with_guarantee(&self, guarantee: Option) -> Self; +} +pub enum Slot { Unset, Set(T) } + +pub fn derive_guarantees(root: &Rc, model: &dyn AccuracyModel, + evidence: &dyn AccuracyEvidenceProvider) -> Result, AccuracyError>; // §2.2 +pub fn derive_timings(root: &Rc) -> Result, ExecutionDataStateError>; // §2.3, §5 +``` + +### 1.2 `NonASAPOp` + +`NonASAPOp` and `ScalarExpr` come from splitting `QueryExpr` +([decoupling doc](decoupling_op_and_expr.md#2-types)). Here `NonASAPOp` becomes the +non-ASAP category of `Operator`: its child slots widen from `Rc` to +`Rc` (§1.4), and every variant gets the `timing` / `guarantee` slots (§1.1). +Scalar fields are unchanged; only `NonASAPOp` holds them. + +```text +Operator +├─ NonASAP(NonASAPOp) +│ ├─ children: Rc> → back to Operator: NonASAP or ASAP +│ ├─ timing / guarantee slots +│ └─ scalar fields: Predicate / ProjectItem / SortKey / ScalarBridge / ... +│ └─ ScalarExpr: never contains an Operator +└─ ASAP(ASAPOp) + ├─ children: Rc> → back to Operator: NonASAP or ASAP + └─ timing / guarantee slots +``` + +### 1.3 `ASAPOp` + +```rust +pub enum ASAPOp { + SummaryAgg { child: Rc>, family: ASAPType, input, reduction, grouping, + exact_rule: Option }, // ExactAggregate only, §2.2 + SummaryEstimate { child, query: SketchQuery, + local_guarantee: Option }, // §2.2 + SummaryMerge { children: Vec>> }, + SummarySubtract { left, right }, + SummaryDelete { child, key: C }, // key was ColumnRef + SummaryJoin { outer, inner, key: C, family: ASAPType }, + // the summary-specific ValueOperation variants, lifted to the top level + FinalizeExactAccumulator { child }, + MaintainPopulation { child, population }, // always ingestion time + ReadPopulation { child, readout }, // always query time + Extension { child, name: String }, +} +``` + +`family` was the original `SummaryFamilyType`. + +Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. + +**Unexercised variants**: `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` +and `Extension` are built only in tests today. They are migrated, but all their methods +return `Unimplemented`. + +Following table shows how some legacy types get expressed in the new framework. + +| Legacy types | Expressed as | +|---|---| +| `SummaryExpr::KeepPreAsap(q)` | `q` itself, an `NonASAP(..)` subtree | +| `ValueOperation::{Project, Filter, Sort, Limit}` | `NonASAPOp::{Project, Filter, Sort, Limit}` | +| `SummaryExpr::{BinaryOp, RelationalJoin}` | `NonASAPOp::{BinaryOp, Join}` | +| `ValueOperation::Exact(Aggregate)`, `ExactOperation` | `NonASAPOp::Aggregate` | +| `SummaryNode` | `Operator` itself: `schema` is computed, `timing` / `guarantee` are slots on every variant (§2) | + +### 1.4 Child field + +Non-ASAP operators now sit on the same level as ASAP operators, so their children must +be `Rc` to allow free placement: + +```rust +// After the decoupling doc // After this proposal +Filter { pred: Predicate(Rc), Filter { pred: Predicate(Rc), + child: Rc } child: Rc } // NonASAP(..) or ASAP(SummaryEstimate ..) +``` + +Now an original operator can also sit on ASAP operators, e.g. a `SetOp` sitting on two `SummaryEstimate` operators. + +`Concat.children` is `Vec` today: branches are stored by value and have no `Rc` identity, so the planner (§4), which identifies targets and holes by pointer, +cannot replace a branch — e.g. the branches of SQL `ROLLUP` or PromQL `histogram_quantiles`. It becomes `Vec>` (§8 stage 0). + +## 2. Per-node information + +Besides children (§1.4), every operator has the three attributes below: the +every-operator level of §1.1. Each row says where the value lives and who supplies it. + +| Field | Meaning | Today | After | +|---|---|---|---| +| `schema` | output columns and their types | stored on every `SummaryNode` | obtained by `output_schema()` (§2.1) | +| `guarantee` | how far the output value can be off | stored on every `SummaryNode` | obtained by `guarantee()`: derived for every node; binding stores only a sketch's own error, as a `SummaryEstimate` field (§2.2) | +| `timing` | ingestion time or query time | on `BinaryOp` / `ValueOperation` / `SummaryMerge` | obtained by `timing()`: set by binding and lifecycle (`SummaryAgg`) or the planner (`FinalizeExactAccumulator`); derived for the rest (§2.3) | + +### 2.1 Schema: fused into one type + +Today pre-ASAP uses `Schema { columns: Vec, time_index, unique_keys, closed }` +with `Column.dtype: DataType` (plain values only), and post-ASAP stores a +`SummarySchema { fields: Vec, time_index }` on every node, with +`SummaryField.dtype: SummaryFamilyType` (`Plain(DataType)` or summary state). One +tree now needs one schema type whose columns can be either: + +``` +SummaryEstimate(Quantile .99) → [p99: DataType(Float64)] + SummaryAgg(Kll) → [state: ASAPType(Sketch(Kll, k=269))] + Scan lineitem → [l_orderkey: DataType(Int64), l_quantity: DataType(Float64), …] +``` + +We merge semantics of the above two types into one `Schema` type. +The struct of `Schema` stays, with two changes: + +- `Column` is renamed `Field`, and `Schema.columns` `Schema.fields`: the struct describes + a column and holds none of its data. Arrow and DataFusion use the same names. +- `Field.dtype` widens from `DataType` to an enum `FieldType`, so a field can describe either a + plain value or ASAP state. + +`SummarySchema` / `SummaryField` are then redundant and deleted: + +```rust +pub enum FieldType { DataType(DataType), ASAPType(ASAPType) } +pub enum ASAPType { // SummaryFamilyType without Plain + ExactAggregate(ExactKind, ExactParams), Sketch(SketchKind, GroupingStrategy), + Sample(SamplingKind, SamplingParams), Wavelet(WaveletKind, WaveletParams), StatModel(StatModelKind, StatModelParams), +} +pub struct Schema { pub fields: Vec, pub time_index, pub unique_keys, pub closed } +pub struct Field { pub name, pub dtype: FieldType, pub nullable, pub table: Option } +impl Field { + pub fn plain(name, DataType) -> Self; + pub fn plain_dtype(&self) -> Option<&DataType>; // None for a state column + pub fn expect_plain_dtype(&self) -> &DataType; // frontends, scalar type inference; panics on state +} +``` + +Today every `SummaryNode` stores its schema, built at construction. After, every node +computes it with `output_schema()`: + +| Node | Today | After | +|---|---|---| +| `NonASAPOp` | `QueryExpr::output_schema()` lifted to `SummarySchema` (`KeepPreAsap`), or stored on the `ValueOperation` / `BinaryOp` / `RelationalJoin` copy | today's `QueryExpr::output_schema()` logic | +| `SummaryAgg` | the replaced `Aggregate`'s output with the measure column retyped to `family` | grouping columns + one `ASAPType(family)` column | +| `SummaryEstimate` | the replaced operator's output schema | the child's grouping columns + the value columns of the `SketchQuery` | +| `FinalizeExactAccumulator` | the logical operator's output, lifted | the child's schema, `ASAPType(ExactAggregate ..)` columns turned into `DataType(..)` | +| `MaintainPopulation` / `ReadPopulation` | the source's schema / the replaced aggregate's output | the same rules, computed from the child and the `readout` | +| unexercised variants | one field typed `family` | unimplemented (§1.3) | + +### 2.2 Guarantee: always derived + +| Node | Today | After | +|---|---|---| +| `SummaryEstimate` | stored at binding: the sketch's own error composed with the child's (`compose_guarantee`) | **derived**: `local_guarantee` composed with the child's. `local_guarantee` is a field set at binding: the sketch's error over an exact input, `None` when the model has no error model for the family | +| `SummaryAgg` | stored: ExactAggregate family composed with the child's; sketch families `None` | **derived**: ExactAggregate family: exact, composed with the child's under `exact_rule`; sketch families `Set(None)`, state has no guarantee | +| `NonASAPOp` | `KeepPreAsap`: exact; the `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: composed at construction | **derived**: composed from the children (`relational_join_guarantee`, `exact_operation_rule`); exact if no `ASAP` descendant | +| `FinalizeExactAccumulator` | copies the child's | **derived**: the child's | +| `MaintainPopulation` / `ReadPopulation` | stored: exact | **derived**: exact | +| unexercised variants | `None`: state has no guarantee of its own | unimplemented (§1.3) | + +- **Only the local part is stored.** Today binding stores the composed value, built + from the child it sees. Assembly can fill that child's hole with a different plan, and + a stored composition would go stale (`relink_agg_child` copies it today, relying on the + new child being exact). +- `derive_guarantees` uses the same `AccuracyModel` as binding, and the evidence for + `propagate`'s `PropagationStats`. It reads only the subtree, so search runs it on a + candidate to check its accuracy target (the candidate filter in + `search_workload_with_targets`). +- The value travels with the node through cloning, CSE and serialization, as + `SummaryNode.guarantee` does today. + +### 2.3 Timing: set where position does not decide it + +| Node | Today | After | +|---|---|---| +| `NonASAPOp` | `KeepPreAsap`: from the consuming edge; the `ValueOperation` / `BinaryOp` copies: a stored field | **derived** from the consuming edge (§5) | +| `SummaryAgg` | from the child; ingestion time under `KeepPreAsap` | **set**: binding sets `QueryTime`; a lifecycle decision may change it to `IngestionTime` | +| `FinalizeExactAccumulator` | a stored field, set by the planner | **set** by the planner: the same position allows either time | +| `SummaryEstimate` | query time, fixed by the kind | **derived** from the kind: query time | +| `MaintainPopulation` / `ReadPopulation` | a stored field, always ingestion / query time | **derived** from the kind: ingestion / query time | +| unexercised variants | `SummaryMerge`: a stored field; `Join` / `Subtract` / `Delete`: ingestion time | unimplemented (§1.3) | + +Unlike a guarantee, a timing depends on the parents, so `derive_timings` needs the whole +DAG and runs only after assembly. + +--- + +# II. Changes, in data-flow order + +## 3. Optimizer entry + +Frontends and `resolve` build `NonASAP` trees only and access children with +`expect_non_asap()`. `ParsedWorkload::new` rejects a tree that `contains_asap()`, +next to its existing entry-count check: + +```rust +impl Operator { + pub fn contains_asap(&self) -> bool; + pub fn expect_non_asap(&self) -> &NonASAPOp; // an ASAP node here is a bug: panic +} +``` + +A compile-time alternative — an associated type on `ColState` with +`ColumnRef::ASAP = Never` — only protects frontend code before `resolve`: frontends +already return `ColumnId` trees, where `ASAP` is allowed. The entry check covers every +input (frontends after `resolve`, deserialized plans, third-party frontends, test IR) +with simpler types. + +## 4. Planner: search and assembly + +```rust +pub enum Replacement { + Subtree(Rc), // formerly Summary(Rc) and Rewrite(Rc) + ExactComposition { .. }, // its plan becomes Rc +} +``` + +**Holes**: a subtree of a candidate that is `Rc::ptr_eq` to a target is filled by that +target's own choice. A candidate needing a specific child implementation (e.g. +`realize_temporal_average`) inlines a new node, so it is not a hole. +`realize_child_with` falls back to leaving a hole instead of `keep_pre_asap`. + +**Assembly** — one rule replaces `assemble_residual`: + +```rust +fn assemble(&self, t: &Rc) -> Rc { + memo by ptr; // shared children stay one Rc + let body = match self.chosen(t) { + Some(Subtree(r)) => r, // rewrites are assembled further down too; today assemble_target wraps them in keep_pre_asap + Some(ExactComposition{..}) => composition.plan, + None => t, // keep the node, recurse — assemble_residual does this for four operators only + }; + body.map_children(|c| if is_target(c) { self.assemble(c) } else { c }) +} +``` + +Then run `derive_timings` (§5) once: + +- **Illegal fill** (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`): + `derive_timings` returns an error, and the hole falls back to its original subtree — + `relink_summary`'s fallback, for every operator. +- **A shared subtree read at two timings**, e.g. a query-time `Aggregate` and an + ingestion-time `SummaryAgg` reading one `Scan`: both choices are legal, only the sharing + is not. `derive_timings` memoizes by (pointer, timing), so it builds one copy per + timing; a subtree read at one timing stays one `Rc`. + +Finally run `derive_guarantees` (§2.2). `map_children` is `rebuild_children` from +`pre_asap/cse.rs`, dispatching to `NonASAPOp::map_children` / `ASAPOp::map_children`. +Deleted: `assemble_residual`, `relink_summary`, the `query_time_nested_sum` / +`contains_aggregate` special cases, `keep_pre_asap` / `keep_pre_asap_rc`, and the +`KeepPreAsap` branch of `finalize_exact_accumulator`. + +| #468 problem | Resolution | +|---|---| +| 1. A `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside | one set of types | +| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds when both run at the same time — always, without lifecycle decisions; otherwise the scan is duplicated (above). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | +| 3. `SetOp` and similar have no post-ASAP copy, so no summary below them | `SetOp` takes `None => t`; both children are assembled | + +## 5. Timing: execution data states + +`validate_execution_data_states` becomes `derive_timings`: instead of returning the +`ExecutionDataStateAssignment` side table (deleted), it writes each node's `timing` slot. +`produced_data_state(KeepPreAsap) = None` (set by the consuming edge) extends to all +`NonASAPOp`s: + +| Node | Produced state | +|---|---| +| `NonASAPOp` | set by the consuming edge (`QUERY_ROWS` at the root), passed to its children | +| `SummaryAgg` | `{timing, SummaryState}` from its `timing` slot (§2.3); a `NonASAPOp` child takes the same timing | +| other `ASAPOp` | unchanged | + +A data state is the `timing` slot plus a primitive (`Raw` / `SummaryState` / …) fixed by +the kind; only the timing is stored. + +The `KeepPreAsap` / `BinaryOp` / `ValueOperation` / `RelationalJoin` arms of today's +`validate_execution_data_states` merge into one `NonASAP` arm of `derive_timings`: + +- Pass the state to each child; an `ASAP` child is checked by the `ASAP` edge rules. +- `check_plain_operands` stays: referenced columns must be `FieldType::DataType` + (`Project` / `Filter` / `Sort` / `Limit` may pass `ExactAggregate` columns through). + This rejects `Project(ASAP(SummaryAgg))`. +- `BinaryOp`'s ingestion-side constraints move into this arm. +- `AmbiguousKeepPreAsap` is deleted: a subtree read at two timings is copied (§4). + +## 6. Export: fragments in the executable DAG + +The four original-operator payloads (`fallback{expression: QueryExpr}`, `binary`, +`value`, `relational_join`) become one: + +```rust +ExecutableOperatorPayload::Relational { + /// No ASAP node inside. Leaves are Scans, or Scan { source: Source::DagInput { role } } + /// for an incoming edge whose schema is the edge's intermediate_schema. + expression: Operator, +} +``` + +`compile_executable_dag` takes each **largest connected subtree without `ASAP`** as one +fragment, cutting an edge with a `DagInput` leaf wherever it meets an `ASAP` node. +`ASAP` nodes map one-to-one onto the existing summary payloads; +`FinalizeExactAccumulator` / `MaintainPopulation` / `ReadPopulation` stay +`value{operation}`. A backend lowers every fragment with its existing `QueryExpr` +lowering plus a `DagInput` arm (an incoming edge as a materialized table); the +`binary` / `value::Project` / `relational_join` lowerings go. + +- **Wire 5 → 6**: three fewer payloads; `fallback` becomes `relational` with `DagInput` + leaves; `output_schema` / `intermediate_schema` become `Schema`. One cutover (§8 + stage 4), together with the downstream readers. +- **Timing and guarantee** are read from the node slots; an `Unset` slot is rejected. + An edge's `data_state` is its producer's timing plus the primitive of its kind (§5). + `compile_executable_dag` no longer re-runs data-state validation. +- **Phases** become per fragment. Switching phase inside a fragment would need a + materialization point, and those are `ASAP` nodes, so nothing is lost. + +## 7. Other consumers + +| Location | Change | +|---|---| +| `post_asap/cse.rs` | delete; `share_common_subtrees` covers `ASAPOp` (derives `PartialEq` + serde) | +| `dag_export.rs` | delete `build_summary` / `build_summary_hybrid` / `summary_kind_tag`; one exporter with an `ASAP` arm; update the viewer's `node-style.js` and the pin test `viewer_categorizes_exactly_the_exported_node_kinds` | +| `summary_maintenance_cost/estimator.rs` (80 sites) | `KeepPreAsap` branches (`query_source_selections`, `retained_queries`) use the §6 fragment; `exact_binary` / `value_operation` costs fold into it. Also fixes the missing `RelationalJoin` arm in `summary_operation_evidence` | +| `physical_plan_cost_model.rs::estimate_candidate` | every fragment goes through `lower_query_physical_dag` | +| `summary_maintenance_lifecycle.rs` | `selected_raw_recompute` becomes `!contains_asap(root)`; the `keep_pre_asap(target)` fallback in `assemble_selected_dag_with_summary_maintenance_lifecycles` becomes `target`; the lifecycle decision sets a `SummaryAgg` to `IngestionTime` with `with_timing` (§2.3) | +| `maintained_population.rs` | `KeepPreAsap(source)` becomes `source`; `population.matches_input` reads an `NonASAP` child directly | +| `exact_composition.rs` | `ExactOperation::Aggregate` becomes an `NonASAP(Aggregate)` whose child is a hole | +| `RelationalJoin.pruning` | never set to `Some` in production; delete. Candidate pruning can return as an `ASAPOp` variant | + +--- + +# III. Implementation + +## 8. Stages and tests + +`main` builds and passes all tests after every stage. + +| Stage | Content | Touches | +|---|---|---| +| 0 Preparation | `Rc` for `Concat.children`; `rebuild_children` → `map_children`; `Column::plain` | `asap-types` | +| 1 Split | [decoupling doc](decoupling_op_and_expr.md): `NonASAPOp` + `ScalarExpr`; children stay `Rc` | scalar code ([decoupling doc §3](decoupling_op_and_expr.md#3-changes)) | +| 2 Two levels | §1.1, §1.4: `Operator`, an empty `ASAPOp`, `contains_asap()`, `expect_non_asap()`; child slots become `Rc>`; every variant gets `timing` / `guarantee` slots, left `Unset` | every crate; the same mechanical change everywhere | +| 3 One schema | §2.1: `Column` → `Field` and `Schema.columns` → `fields` (serde keeps the name `columns` until stage 4); `FieldType`, `ASAPType`, `PlainField`, `Schema` everywhere except the `executable_dag.rs` wire types, which keep `SummarySchema` until stage 4. **No wire change** | `asap-types` + schema construction in every crate | +| 4 New types | fill `ASAPOp`; `ASAP` arms of `output_schema`; `derive_guarantees` and `derive_timings` (§2.2, §5); the entry check (§3); `flatten(&SummaryNode) -> Rc` so export runs on the new types; wire types become `Schema`, and `Schema.fields` serializes as `fields`. Wire → 6. **The only wire-breaking stage**; merged together with ASAPQuery-backend and ASAPCollector | `asap-types`, `devtools`, viewer | +| 5 Planner | §4: candidates and assembly on `Rc`; §7 moves to the new types; delete `flatten` | `asap-aware-mapping` | +| 6 Cleanup | delete `SummaryExpr`, `SummaryNode`, extra `ValueOperation` variants, `ExactOperation`, `post_asap/cse.rs`; update `post-asap-ir.md`, `physical-plan-integration.md`, developer and viewer docs | docs | + +Wrapping pre-ASAP operators in `ValueOperation` first is not planned: stage 4 gives +the same early flat export, on the final types. + +**Tests**: + +- One integration test per #468 problem: + 1. `WITH metric AS (SELECT avg(CASE WHEN l_quantity BETWEEN 1 AND 50 THEN 1.0 ELSE 0.0 END) AS in_range FROM lineitem) SELECT in_range, in_range = 1.0 AS ok FROM metric` — no post-ASAP-only node besides `ASAP`; all `Project`s are one variant. + 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` (once a binding rule splits measures). + 3. `SELECT approx_distinct(l_partkey) FROM lineitem UNION ALL SELECT approx_distinct(l_suppkey) FROM lineitem` — each side of the `SetOp` has a `SummaryEstimate`. +- A shared `Scan` read at two timings is copied once per timing; read at one timing, it stays one `Rc`. +- After `derive_*`, no slot is `Unset`; export rejects a tree with one. +- A `SummaryEstimate` over a hole reports the error of the plan that fills the hole. +- Each unexercised variant (§1.3) returns `Unimplemented` from `output_schema`, `derive_*` and export. +- `ParsedWorkload::new` rejects a tree containing `ASAP`. +- Rewrite the 117 `SummaryExpr::` assertions in `sql_to_post_asap.rs` / `promql_to_post_asap.rs` / `exact_composition.rs`. +- Wire 6 round trip with a `DagInput` fragment; a version-5 document is rejected. +- The 52 `execution_data_state.rs` tests keep their shapes; assertions read the `timing` slot instead of `ExecutionDataStateAssignment`. + +## 9. Out of scope + +- The binding rule splitting a multi-measure `Aggregate` into exact + summary over one child. +- Candidate pruning as an `ASAPOp` variant. +- Accuracy through `SummaryMerge` / `Subtract` / `Delete` / `Join` (§2.2): an accuracy + descriptor on state, or composing error along the state chain at readout. +- Folding `ExactComposition` into `Subtree` (both of its forms become expressible); + deferred until stage 5 is stable. + +## 10. Open questions + +None at present. From 19536469269363c0652e4ce32e93c5caee38a1a2 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Tue, 29 Sep 2026 16:51:13 +0000 Subject: [PATCH 04/46] document updated --- .../proposals/decoupling_op_and_expr.md | 2 +- .../design_docs/proposals/operator-sharing.md | 358 +++++++++++------- 2 files changed, 229 insertions(+), 131 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index 9d34b4e4a..b28727b4b 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -84,7 +84,7 @@ NonASAPOp | Location | Change | |---|---| | frontend expression lowering (`df_expr_to_unresolved`, PromQL `walk`) | scalar positions build `ScalarExpr`, operator positions `NonASAPOp` | -| `resolve`, `column_resolution.rs` | separate operator and scalar resolvers; a scalar resolves against its operator's input schema | +| `resolve`, `column_resolution.rs` | already separate: `resolve` walks operators and calls `resolve_expr` for scalars. Each scalar resolves against one schema its operator picks (usually the child's output; the `Aggregate`'s output for `HAVING`, left + right for a `Join` predicate, the `Scan`'s own schema for `Scan` predicates). The split only changes their signatures: `resolve` takes `NonASAPOp`, `resolve_expr` takes `ScalarExpr`. Leaf schemas are still inferred from scalar column references across the whole tree | | `canonicalize`, `pre_asap/cse.rs` | the "scalar: nothing to do" arms go; scalars are hashed as plain data | | `scalar_signature.rs`, `infer_expr_type` | take `ScalarExpr` | | `QueryExpr::output_schema` | becomes `NonASAPOp::output_schema`; the scalar arms and `ScalarHasNoRowSchema` go | diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 4ee165901..44254c4cd 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -1,11 +1,8 @@ # Sharing Operators Between Pre-ASAP IR and Post-ASAP IR -> Status: proposed, not implemented. Problem statement: -> [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). Implementation -> starts after the single entry point and the pluggable pass (#429, #430) land on -> `main`. Builds on [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) -> (same PR), which splits `QueryExpr` into `NonASAPOp` and `ScalarExpr`. Code is -> referenced by file and function; counts are approximate, measured on `main` at `8acb472`. +> - Status: proposed, not implemented. +> - Problem statement: [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). +> - Builds on [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) (same PR), which splits `QueryExpr` into `NonASAPOp` and `ScalarExpr`. **The idea.** Today a post-ASAP plan is glued together from two sets of operator types. This proposal keeps one operator language and makes summary operators extra node kinds in it: any relational operator can sit above a summary, and a summary can read any relational subtree. @@ -21,7 +18,7 @@ ValueOperation(Project) ← a copy NonASAP(Project) | Part | Sections | |---|---| -| I. New IR | §1 Types, §2 Per-node information | +| I. New IR | §1 Types, §2 Schema, guarantee, and timing | | II. Changes, in data-flow order | §3 Entry → §4 Planner → §5 Timing → §6 Export → §7 Other consumers | | III. Implementation | §8 Stages and tests, §9 Out of scope, §10 Open questions | @@ -31,86 +28,125 @@ ValueOperation(Project) ← a copy NonASAP(Project) ## 1. Types -### 1.1 Overview +### 1.1 Unified `Operator` type -Operator attributes differ in how widely they apply. Each is defined at the level -that matches its breadth: +Operator attributes differ in how widely they apply. +We define the `Operator` type structure based on the breadth of its attributes. | Applies to | Examples | Defined as | |---|---|---| -| every operator | children, schema, timing, guarantee | methods of `Operator`; the values may be stored per variant | -| one category | for all `NonASAP` operators, timing is derived from the consuming edge, and guarantee from the children | a variant of `Operator` | +| every operator | children, schema, timing, guarantee | methods implemented for `Operator` | +| one category | for all `NonASAP` operators, timing is derived from the consuming edge, and guarantee from the children | implementation specified to one enum branch of `Operator` | | one operator | `Aggregate.measures`, `SummaryAgg.family` | fields of that variant | ```rust -pub enum Operator { // category - NonASAP(NonASAPOp), // today's relational QueryExpr variants (§1.2) - ASAP(ASAPOp), // summary operators (§1.3) +pub enum Operator { + NonASAP(NonASAPOp), // today's relational and timeseries operators in `QueryExpr` (§1.2) + ASAP(ASAPOp), // summary operators (§1.3) } -impl Operator { // every operator +impl Operator { // implemented for every operator pub fn children(&self) -> Vec<&Rc>>; pub fn map_children(&self, f: impl FnMut(&Rc>) -> Rc>) -> Self; - pub fn output_schema(&self) -> Result; // computed (§2.1) - pub fn timing(&self) -> &Slot; // §2.3 - pub fn guarantee(&self) -> &Slot>; // §2.2; Set(None): unknown, never read as exact - pub fn with_timing(&self, timing: ExecutionTiming) -> Self; - pub fn with_guarantee(&self, guarantee: Option) -> Self; + pub fn output_schema(&self) -> Result; // schema: §2.1 + pub fn guarantee(&self) -> &Slot>; // accuracy guarantee: §2.2 + pub fn timing(&self) -> &Slot; // execution timing: §2.3 + pub fn with_guarantee(&self, guarantee: Option) -> Self; // Setter of accuracy guarantee + pub fn with_timing(&self, timing: ExecutionTiming) -> Self; // Setter of execution timing } + +/// `Slot` represents a value that may be unset or set. +/// In the current design, it will be used to wrap the `timing` and `guarantee` values, +/// whose values will only be determined after derivation. pub enum Slot { Unset, Set(T) } -pub fn derive_guarantees(root: &Rc, model: &dyn AccuracyModel, - evidence: &dyn AccuracyEvidenceProvider) -> Result, AccuracyError>; // §2.2 -pub fn derive_timings(root: &Rc) -> Result, ExecutionDataStateError>; // §2.3, §5 +/// Memo of one derivation, keyed by (node pointer, incoming timing). +/// One is shared by every root of a workload, so a node shared by two roots stays one `Rc`. +pub struct DerivationMemo { .. } + +/// Build the accuracy guarantee of one DAG root by derivation +pub fn derive_guarantees( + root: &Rc, + model: &dyn AccuracyModel, // accuracy model used for derivation + evidence: &dyn AccuracyEvidenceProvider, // evidence provider used for derivation + memo: &mut DerivationMemo, +) -> Result, AccuracyError>; +/// Build the execution timing of one DAG root by derivation +pub fn derive_timings( + root: &Rc, + memo: &mut DerivationMemo, +) -> Result, ExecutionDataStateError>; ``` -### 1.2 `NonASAPOp` - -`NonASAPOp` and `ScalarExpr` come from splitting `QueryExpr` -([decoupling doc](decoupling_op_and_expr.md#2-types)). Here `NonASAPOp` becomes the -non-ASAP category of `Operator`: its child slots widen from `Rc` to -`Rc` (§1.4), and every variant gets the `timing` / `guarantee` slots (§1.1). -Scalar fields are unchanged; only `NonASAPOp` holds them. +- **Derivation recomputes**: `derive_*` keep the values set at construction (§2.2, §2.3) + and recompute every other slot, so calling them again after a rewrite is safe. +- **Equality**: both slots take part in `PartialEq` and hashing, so CSE never merges two + nodes that differ in timing or guarantee. +Following diagram conceptually displays the structure of `Operator`: ```text Operator ├─ NonASAP(NonASAPOp) │ ├─ children: Rc> → back to Operator: NonASAP or ASAP -│ ├─ timing / guarantee slots -│ └─ scalar fields: Predicate / ProjectItem / SortKey / ScalarBridge / ... -│ └─ ScalarExpr: never contains an Operator +│ ├─ timing / guarantee +│ └─ scalar expressions: Predicate / ProjectItem / SortKey / ScalarBridge / ... +│ └─ ScalarExpr: never contains an Operator └─ ASAP(ASAPOp) ├─ children: Rc> → back to Operator: NonASAP or ASAP - └─ timing / guarantee slots + └─ timing / guarantee ``` +### 1.2 `NonASAPOp` + +`NonASAPOp` is the non-ASAP category of `Operator`. +It comes from splitting `QueryExpr` into "operator" and "scalar expression" parts ([decoupling doc](decoupling_op_and_expr.md#2-types)). + +```rust +pub enum NonASAPOp { + Scan { .. }, + Filter { pred: Predicate, child: Rc> }, + Project { cols: Vec>, child: Rc> }, + Aggregate { reduction, measures, having: Option>, child: Rc> }, + Join { kind, pred: Predicate, left: Rc>, right: Rc> }, + SetOp { kind, all, left: Rc>, right: Rc> }, + Concat { children: Vec>> }, + Sort { keys: Vec>, child: Rc> }, + Limit { n, offset, child: Rc> }, + BinaryOp { op, lhs, rhs }, + SQLWindowFunc { args: Vec>, order_by: Vec>, child: Rc>, .. }, + Dedup { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, + ScalarBridge(Rc>), // the `2` in PromQL `v * 2` + EvalTimestamp, // PromQL time() +} +``` + +Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. + ### 1.3 `ASAPOp` +`ASAPOp` is the ASAP category of `Operator`. +`ASAPOp` comes from today's `SummaryExpr`: its summary variants, and the summary-specific `ValueOperation` variants. + ```rust pub enum ASAPOp { SummaryAgg { child: Rc>, family: ASAPType, input, reduction, grouping, - exact_rule: Option }, // ExactAggregate only, §2.2 - SummaryEstimate { child, query: SketchQuery, - local_guarantee: Option }, // §2.2 + exact_rule: Option }, + SummaryEstimate { child: Rc>, query: SketchQuery, + local_guarantee: Option }, SummaryMerge { children: Vec>> }, - SummarySubtract { left, right }, - SummaryDelete { child, key: C }, // key was ColumnRef - SummaryJoin { outer, inner, key: C, family: ASAPType }, - // the summary-specific ValueOperation variants, lifted to the top level - FinalizeExactAccumulator { child }, - MaintainPopulation { child, population }, // always ingestion time - ReadPopulation { child, readout }, // always query time - Extension { child, name: String }, + SummarySubtract { left: Rc>, right: Rc> }, + SummaryDelete { child: Rc>, key: C }, + SummaryJoin { outer: Rc>, inner: Rc>, key: C, family: ASAPType }, + FinalizeExactAccumulator { child: Rc> }, + MaintainPopulation { child: Rc>, population }, + ReadPopulation { child: Rc>, readout }, + Extension { child: Rc>, name: String }, } ``` -`family` was the original `SummaryFamilyType`. - Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. -**Unexercised variants**: `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` -and `Extension` are built only in tests today. They are migrated, but all their methods -return `Unimplemented`. +**Unused branches**: `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` are built only in tests today. They are migrated, but for safety, we have all their methods return `Unimplemented`. Following table shows how some legacy types get expressed in the new framework. @@ -135,44 +171,35 @@ Filter { pred: Predicate(Rc), Filter { pred: Predicate(Rc` today: branches are stored by value and have no `Rc` identity, so the planner (§4), which identifies targets and holes by pointer, +`Concat.children` is `Vec` today: branches are stored by value and have no `Rc` identity, so the planner (§4), which identifies targets by pointer, cannot replace a branch — e.g. the branches of SQL `ROLLUP` or PromQL `histogram_quantiles`. It becomes `Vec>` (§8 stage 0). -## 2. Per-node information +## 2. Schema, Guarantee, and Timing -Besides children (§1.4), every operator has the three attributes below: the -every-operator level of §1.1. Each row says where the value lives and who supplies it. +This section discusses three key per-node attributes, `schema`, `guarantee`, and `timing`, as well as how they are stored and derived in the new framework. | Field | Meaning | Today | After | |---|---|---|---| -| `schema` | output columns and their types | stored on every `SummaryNode` | obtained by `output_schema()` (§2.1) | -| `guarantee` | how far the output value can be off | stored on every `SummaryNode` | obtained by `guarantee()`: derived for every node; binding stores only a sketch's own error, as a `SummaryEstimate` field (§2.2) | -| `timing` | ingestion time or query time | on `BinaryOp` / `ValueOperation` / `SummaryMerge` | obtained by `timing()`: set by binding and lifecycle (`SummaryAgg`) or the planner (`FinalizeExactAccumulator`); derived for the rest (§2.3) | +| `schema` | output columns and their types | pre-ASAP: computed by `QueryExpr::output_schema()`
post-ASAP: a `SummarySchema` stored on every `SummaryNode` | can be obtained by `output_schema()` | +| `guarantee` | accuracy bound | pre-ASAP: none
post-ASAP: stored on every `SummaryNode` | can be obtained by `guarantee()`
binding stores only each operator's own error
complete error bound need to be derived by `derive_guarantees()` | +| `timing` | execution time | pre-ASAP: none
post-ASAP, stored: a field on `BinaryOp` / `ValueOperation` / `SummaryMerge`
post-ASAP, not stored: `KeepPreAsap` from the consuming edge, `SummaryAgg` from the child. | can be obtained by `timing()`
set by binding (`SummaryAgg`) or the planner (`FinalizeExactAccumulator`)
timing of the rest of operators need to be derived by `derive_timings()` | ### 2.1 Schema: fused into one type -Today pre-ASAP uses `Schema { columns: Vec, time_index, unique_keys, closed }` -with `Column.dtype: DataType` (plain values only), and post-ASAP stores a -`SummarySchema { fields: Vec, time_index }` on every node, with -`SummaryField.dtype: SummaryFamilyType` (`Plain(DataType)` or summary state). One -tree now needs one schema type whose columns can be either: - -``` -SummaryEstimate(Quantile .99) → [p99: DataType(Float64)] - SummaryAgg(Kll) → [state: ASAPType(Sketch(Kll, k=269))] - Scan lineitem → [l_orderkey: DataType(Int64), l_quantity: DataType(Float64), …] -``` +Today schemas of pre-ASAP operators and post-ASAP operators are different: +- pre-ASAP uses `Schema { columns: Vec, time_index, unique_keys, closed }` with `Column.dtype: DataType` (plain values only), +- post-ASAP stores a `SummarySchema { fields: Vec, time_index }` on every node, with `SummaryField.dtype: SummaryFamilyType` (`Plain(DataType)` or summary state). +Now since the two operators types are unified into one, we need a unified schema type as well. -We merge semantics of the above two types into one `Schema` type. -The struct of `Schema` stays, with two changes: +We implement the new schema type based on the original `Schema` type used in pre-ASAP operators, with two changes: - `Column` is renamed `Field`, and `Schema.columns` `Schema.fields`: the struct describes - a column and holds none of its data. Arrow and DataFusion use the same names. -- `Field.dtype` widens from `DataType` to an enum `FieldType`, so a field can describe either a - plain value or ASAP state. + a column and holds none of its data. (Arrow and DataFusion use the same names.) +- `Field.dtype` widens from `DataType` to an enum `FieldType`, which covers both plain data types and ASAP summary types. -`SummarySchema` / `SummaryField` are then redundant and deleted: +`SummarySchema` / `SummaryField` are then redundant and deleted. +Detailed code design is shown below. ```rust pub enum FieldType { DataType(DataType), ASAPType(ASAPType) } pub enum ASAPType { // SummaryFamilyType without Plain @@ -188,54 +215,106 @@ impl Field { } ``` -Today every `SummaryNode` stores its schema, built at construction. After, every node -computes it with `output_schema()`: - | Node | Today | After | |---|---|---| -| `NonASAPOp` | `QueryExpr::output_schema()` lifted to `SummarySchema` (`KeepPreAsap`), or stored on the `ValueOperation` / `BinaryOp` / `RelationalJoin` copy | today's `QueryExpr::output_schema()` logic | +| `NonASAPOp` | post-ASAP `KeepPreAsap`: `QueryExpr` schema lifted to `SummarySchema` and stored
post-ASAP `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: stored at construction | using the same logic as `QueryExpr::output_schema()` | | `SummaryAgg` | the replaced `Aggregate`'s output with the measure column retyped to `family` | grouping columns + one `ASAPType(family)` column | | `SummaryEstimate` | the replaced operator's output schema | the child's grouping columns + the value columns of the `SketchQuery` | -| `FinalizeExactAccumulator` | the logical operator's output, lifted | the child's schema, `ASAPType(ExactAggregate ..)` columns turned into `DataType(..)` | +| `FinalizeExactAccumulator` | the logical operator's output, lifted | the child's schema, `ASAPType(ExactAggregate ..)` columns changed into `DataType(..)` | | `MaintainPopulation` / `ReadPopulation` | the source's schema / the replaced aggregate's output | the same rules, computed from the child and the `readout` | -| unexercised variants | one field typed `family` | unimplemented (§1.3) | +| unused variants | one field typed `family` | unimplemented | ### 2.2 Guarantee: always derived +A guarantee is filled in two steps: + +1. **Binding** records local accuracy guarantee: a `SummaryEstimate`'s `local_guarantee` (the sketch's error over an exact input) and an exact `SummaryAgg`'s `exact_rule`. No `guarantee` slot is set yet. +2. **`derive_guarantees`** fills every slot bottom-up: a node without an `ASAP` descendant is exact, and every other node composes its children's guarantees by its own rule. + +``` +Project p99 ±1% ← the child's + SummaryEstimate p99 ±1% ← local ±1%, composed with the child's + SummaryAgg(Kll) None ← state has no guarantee + Scan t exact ← no ASAP descendant +``` + +Per node kind: + | Node | Today | After | |---|---|---| -| `SummaryEstimate` | stored at binding: the sketch's own error composed with the child's (`compose_guarantee`) | **derived**: `local_guarantee` composed with the child's. `local_guarantee` is a field set at binding: the sketch's error over an exact input, `None` when the model has no error model for the family | +| `SummaryEstimate` | stored at binding: the sketch's own error composed with the child's (`compose_guarantee`) | **derived**: `local_guarantee` composed with the child's. `local_guarantee` is set at binding: the sketch's error over an exact input, `None` when the model has no error model for the family | | `SummaryAgg` | stored: ExactAggregate family composed with the child's; sketch families `None` | **derived**: ExactAggregate family: exact, composed with the child's under `exact_rule`; sketch families `Set(None)`, state has no guarantee | -| `NonASAPOp` | `KeepPreAsap`: exact; the `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: composed at construction | **derived**: composed from the children (`relational_join_guarantee`, `exact_operation_rule`); exact if no `ASAP` descendant | +| `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: exact
post-ASAP `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: composed at construction | **derived**: composed from the children; exact if no `ASAP` descendant | | `FinalizeExactAccumulator` | copies the child's | **derived**: the child's | | `MaintainPopulation` / `ReadPopulation` | stored: exact | **derived**: exact | -| unexercised variants | `None`: state has no guarantee of its own | unimplemented (§1.3) | - -- **Only the local part is stored.** Today binding stores the composed value, built - from the child it sees. Assembly can fill that child's hole with a different plan, and - a stored composition would go stale (`relink_agg_child` copies it today, relying on the - new child being exact). -- `derive_guarantees` uses the same `AccuracyModel` as binding, and the evidence for - `propagate`'s `PropagationStats`. It reads only the subtree, so search runs it on a - candidate to check its accuracy target (the candidate filter in - `search_workload_with_targets`). -- The value travels with the node through cloning, CSE and serialization, as - `SummaryNode.guarantee` does today. +| unused variants | `None`: state has no guarantee of its own | unimplemented (§1.3) | ### 2.3 Timing: set where position does not decide it +A timing is filled in two steps: + +1. **Binding** sets every `SummaryAgg` to the timing today's fallback derives: + `IngestionTime`, or `QueryTime` when binding built its child at query time (a + query-time `FinalizeExactAccumulator`). The planner sets every + `FinalizeExactAccumulator`. +2. **`derive_timings`** runs on each assembled root, sharing one memo, top-down: a root is query + time, a set node keeps its value, a node of fixed kind takes that kind's time, and + every other node takes its parent's. A node reached at two timings is copied (§4). + +``` + binding derive_timings +Project Unset QueryTime ← root + SummaryEstimate Unset QueryTime ← fixed by kind + SummaryAgg IngestionTime IngestionTime ← kept + Scan t Unset IngestionTime ← from parent +``` + +Exported timings are unchanged. Moving the `SummaryAgg` default to `QueryTime`, and +letting the lifecycle step choose ingestion time, is a separate PR. + +Per node kind: + | Node | Today | After | |---|---|---| -| `NonASAPOp` | `KeepPreAsap`: from the consuming edge; the `ValueOperation` / `BinaryOp` copies: a stored field | **derived** from the consuming edge (§5) | -| `SummaryAgg` | from the child; ingestion time under `KeepPreAsap` | **set**: binding sets `QueryTime`; a lifecycle decision may change it to `IngestionTime` | +| `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: from the consuming edge
post-ASAP `ValueOperation` / `BinaryOp` copies: a stored field | **derived** from the consuming edge (§5) | +| `SummaryAgg` | from the child; ingestion time under `KeepPreAsap` | **set** by binding, as today's fallback: `IngestionTime`, or `QueryTime` over a query-time child | | `FinalizeExactAccumulator` | a stored field, set by the planner | **set** by the planner: the same position allows either time | | `SummaryEstimate` | query time, fixed by the kind | **derived** from the kind: query time | | `MaintainPopulation` / `ReadPopulation` | a stored field, always ingestion / query time | **derived** from the kind: ingestion / query time | -| unexercised variants | `SummaryMerge`: a stored field; `Join` / `Subtract` / `Delete`: ingestion time | unimplemented (§1.3) | +| unused variants | `SummaryMerge`: a stored field; `Join` / `Subtract` / `Delete`: ingestion time | unimplemented (§1.3) | Unlike a guarantee, a timing depends on the parents, so `derive_timings` needs the whole DAG and runs only after assembly. +### 2.4 Workflow of setting up `guarantee` and `timing`: today vs. after + +Today: + +``` +search / binding each SummaryNode's guarantee is composed when the node is built; + BinaryOp / ValueOperation store their timing +selection reads each candidate's stored guarantee against its target +assembly assemble_residual builds kept nodes and composes their guarantee; + relink_summary copies the old guarantee onto a relinked SummaryAgg +lifecycle reads the root's guarantee +export validate_execution_data_states, per root, derives the remaining + timings into a side table and writes them onto the edges +``` + +After: + +``` +search / binding sets only what cannot be derived: SummaryAgg.timing, + local_guarantee, exact_rule; checks accuracy on a derived copy, + then drops the copy +assembly builds each root; kept NonASAP nodes stay as they are +derive_timings per root, one shared memo: fills timings top-down, copies a node + read at two timings +derive_guarantees per root, one shared memo: fills guarantees bottom-up +lifecycle reads the derived guarantee +export reads the slots; rejects an Unset one +``` + --- # II. Changes, in data-flow order @@ -243,8 +322,8 @@ DAG and runs only after assembly. ## 3. Optimizer entry Frontends and `resolve` build `NonASAP` trees only and access children with -`expect_non_asap()`. `ParsedWorkload::new` rejects a tree that `contains_asap()`, -next to its existing entry-count check: +`expect_non_asap()`. `search_cse_workload_with`, which every `search_workload*` entry +reaches, panics on a root that `contains_asap()`: an ASAP node there is a caller bug. ```rust impl Operator { @@ -256,8 +335,7 @@ impl Operator { A compile-time alternative — an associated type on `ColState` with `ColumnRef::ASAP = Never` — only protects frontend code before `resolve`: frontends already return `ColumnId` trees, where `ASAP` is allowed. The entry check covers every -input (frontends after `resolve`, deserialized plans, third-party frontends, test IR) -with simpler types. +input (frontends after `resolve`, deserialized plans, test IR) with simpler types. ## 4. Planner: search and assembly @@ -268,45 +346,57 @@ pub enum Replacement { } ``` -**Holes**: a subtree of a candidate that is `Rc::ptr_eq` to a target is filled by that -target's own choice. A candidate needing a specific child implementation (e.g. -`realize_temporal_average`) inlines a new node, so it is not a hole. -`realize_child_with` falls back to leaving a hole instead of `keep_pre_asap`. +**Candidates stay bottom-up, as today**: a candidate is built on a concrete child plan +(`realize_child_with`, or each child candidate in `prepare_compositions`), so a chosen +plan is complete. Where `realize_child_with` falls back to `keep_pre_asap` today, it +returns the child's original subtree, and assembly keeps it as is. + +**Accuracy check during search**: binding sets no `guarantee` slot (§2.2), so the +candidate filter in `search_workload_with_targets` and `prepare_compositions` run +`derive_guarantees` on the candidate alone, with a fresh memo, then check its accuracy target. The derived +copy is only read, then dropped: PlanSpace keeps the original candidate, whose nodes are +shared with other queries. **Assembly** — one rule replaces `assemble_residual`: ```rust fn assemble(&self, t: &Rc) -> Rc { memo by ptr; // shared children stay one Rc - let body = match self.chosen(t) { - Some(Subtree(r)) => r, // rewrites are assembled further down too; today assemble_target wraps them in keep_pre_asap + match self.chosen(t) { + Some(Subtree(r)) => r, // a complete plan, used as is Some(ExactComposition{..}) => composition.plan, - None => t, // keep the node, recurse — assemble_residual does this for four operators only - }; - body.map_children(|c| if is_target(c) { self.assemble(c) } else { c }) + None => t.map_children(|c| if is_target(c) { self.assemble(c) } else { c }), + // keep the node, assemble its children — assemble_residual does this for four operators only + } } ``` -Then run `derive_timings` (§5) once: - -- **Illegal fill** (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`): - `derive_timings` returns an error, and the hole falls back to its original subtree — - `relink_summary`'s fallback, for every operator. +Then `GlobalSelection::assemble_selected_dag` runs `derive_timings` (§5) → +`derive_guarantees` (§2.2) on each root it assembles. The `DerivationMemo` lives on +`GlobalSelection` next to `assembled_nodes`, so roots assembled one call at a time still +share nodes. `assemble_selected_dag_with_summary_maintenance_lifecycles` plans lifecycles +on that result, as today: lifecycle planning holds `Rc`s into the plan and reads the +root's guarantee, so it must see the derived tree. + +- **Illegal child** (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`): candidates + are checked when built, as `relink_summary` does today + (`validate_execution_data_states_at`). An error from `derive_timings` after assembly is + a bug, and planning fails with that error. - **A shared subtree read at two timings**, e.g. a query-time `Aggregate` and an ingestion-time `SummaryAgg` reading one `Scan`: both choices are legal, only the sharing is not. `derive_timings` memoizes by (pointer, timing), so it builds one copy per - timing; a subtree read at one timing stays one `Rc`. + timing; a subtree read at one timing stays one `Rc`, within a root or across roots. -Finally run `derive_guarantees` (§2.2). `map_children` is `rebuild_children` from +`map_children` is `rebuild_children` from `pre_asap/cse.rs`, dispatching to `NonASAPOp::map_children` / `ASAPOp::map_children`. -Deleted: `assemble_residual`, `relink_summary`, the `query_time_nested_sum` / -`contains_aggregate` special cases, `keep_pre_asap` / `keep_pre_asap_rc`, and the -`KeepPreAsap` branch of `finalize_exact_accumulator`. +Deleted: `assemble_residual`, `keep_pre_asap` / `keep_pre_asap_rc`, and the +`KeepPreAsap` branch of `finalize_exact_accumulator`. Kept: `relink_summary` and the +`query_time_nested_sum` special case, which pick a child after selection today. | #468 problem | Resolution | |---|---| | 1. A `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside | one set of types | -| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds when both run at the same time — always, without lifecycle decisions; otherwise the scan is duplicated (above). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | +| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds when both run at the same time; otherwise the scan is copied (above). With today's defaults the sketch runs at ingestion time and the exact `Aggregate` at query time, so they share only after the `QueryTime` default (separate PR, §2.3). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | | 3. `SetOp` and similar have no post-ASAP copy, so no summary below them | `SetOp` takes `None => t`; both children are assembled | ## 5. Timing: execution data states @@ -362,6 +452,8 @@ lowering plus a `DagInput` arm (an incoming edge as a materialized table); the - **Timing and guarantee** are read from the node slots; an `Unset` slot is rejected. An edge's `data_state` is its producer's timing plus the primitive of its kind (§5). `compile_executable_dag` no longer re-runs data-state validation. +- **`SummaryMerge`** stays a wire payload, although its planner-side variant is + unimplemented (§1.3, §10). - **Phases** become per fragment. Switching phase inside a fragment would need a materialization point, and those are `ASAP` nodes, so nothing is lost. @@ -373,9 +465,9 @@ lowering plus a `DagInput` arm (an incoming edge as a materialized table); the | `dag_export.rs` | delete `build_summary` / `build_summary_hybrid` / `summary_kind_tag`; one exporter with an `ASAP` arm; update the viewer's `node-style.js` and the pin test `viewer_categorizes_exactly_the_exported_node_kinds` | | `summary_maintenance_cost/estimator.rs` (80 sites) | `KeepPreAsap` branches (`query_source_selections`, `retained_queries`) use the §6 fragment; `exact_binary` / `value_operation` costs fold into it. Also fixes the missing `RelationalJoin` arm in `summary_operation_evidence` | | `physical_plan_cost_model.rs::estimate_candidate` | every fragment goes through `lower_query_physical_dag` | -| `summary_maintenance_lifecycle.rs` | `selected_raw_recompute` becomes `!contains_asap(root)`; the `keep_pre_asap(target)` fallback in `assemble_selected_dag_with_summary_maintenance_lifecycles` becomes `target`; the lifecycle decision sets a `SummaryAgg` to `IngestionTime` with `with_timing` (§2.3) | +| `summary_maintenance_lifecycle.rs` | `selected_raw_recompute` becomes `!contains_asap(root)`; the `keep_pre_asap(target)` fallback in `assemble_selected_dag_with_summary_maintenance_lifecycles` becomes `target` | | `maintained_population.rs` | `KeepPreAsap(source)` becomes `source`; `population.matches_input` reads an `NonASAP` child directly | -| `exact_composition.rs` | `ExactOperation::Aggregate` becomes an `NonASAP(Aggregate)` whose child is a hole | +| `exact_composition.rs` | `ExactOperation::Aggregate` becomes a `NonASAP(Aggregate)`, built over each child candidate as `prepare_compositions` does today | | `RelationalJoin.pruning` | never set to `Some` in production; delete. Candidate pruning can return as an `ASAPOp` variant | --- @@ -390,7 +482,7 @@ lowering plus a `DagInput` arm (an incoming edge as a materialized table); the |---|---|---| | 0 Preparation | `Rc` for `Concat.children`; `rebuild_children` → `map_children`; `Column::plain` | `asap-types` | | 1 Split | [decoupling doc](decoupling_op_and_expr.md): `NonASAPOp` + `ScalarExpr`; children stay `Rc` | scalar code ([decoupling doc §3](decoupling_op_and_expr.md#3-changes)) | -| 2 Two levels | §1.1, §1.4: `Operator`, an empty `ASAPOp`, `contains_asap()`, `expect_non_asap()`; child slots become `Rc>`; every variant gets `timing` / `guarantee` slots, left `Unset` | every crate; the same mechanical change everywhere | +| 2 Two levels | §1.1, §1.4: `Operator`, an empty `ASAPOp`, `contains_asap()`, `expect_non_asap()`; child slots become `Rc>`; every variant gets `timing` / `guarantee` slots, and nodes are built through constructors that leave both `Unset` | every crate; the same mechanical change everywhere | | 3 One schema | §2.1: `Column` → `Field` and `Schema.columns` → `fields` (serde keeps the name `columns` until stage 4); `FieldType`, `ASAPType`, `PlainField`, `Schema` everywhere except the `executable_dag.rs` wire types, which keep `SummarySchema` until stage 4. **No wire change** | `asap-types` + schema construction in every crate | | 4 New types | fill `ASAPOp`; `ASAP` arms of `output_schema`; `derive_guarantees` and `derive_timings` (§2.2, §5); the entry check (§3); `flatten(&SummaryNode) -> Rc` so export runs on the new types; wire types become `Schema`, and `Schema.fields` serializes as `fields`. Wire → 6. **The only wire-breaking stage**; merged together with ASAPQuery-backend and ASAPCollector | `asap-types`, `devtools`, viewer | | 5 Planner | §4: candidates and assembly on `Rc`; §7 moves to the new types; delete `flatten` | `asap-aware-mapping` | @@ -403,13 +495,15 @@ the same early flat export, on the final types. - One integration test per #468 problem: 1. `WITH metric AS (SELECT avg(CASE WHEN l_quantity BETWEEN 1 AND 50 THEN 1.0 ELSE 0.0 END) AS in_range FROM lineitem) SELECT in_range, in_range = 1.0 AS ok FROM metric` — no post-ASAP-only node besides `ASAP`; all `Project`s are one variant. - 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` (once a binding rule splits measures). + 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` when both run at the same time (once a binding rule splits measures); with today's ingestion-time default the `Scan` is copied. 3. `SELECT approx_distinct(l_partkey) FROM lineitem UNION ALL SELECT approx_distinct(l_suppkey) FROM lineitem` — each side of the `SetOp` has a `SummaryEstimate`. - A shared `Scan` read at two timings is copied once per timing; read at one timing, it stays one `Rc`. +- A node shared by two roots, assembled in two calls, is still one `Rc` after `derive_*`. - After `derive_*`, no slot is `Unset`; export rejects a tree with one. -- A `SummaryEstimate` over a hole reports the error of the plan that fills the hole. -- Each unexercised variant (§1.3) returns `Unimplemented` from `output_schema`, `derive_*` and export. -- `ParsedWorkload::new` rejects a tree containing `ASAP`. +- Exported timings of today's plans are unchanged. +- A kept `NonASAP` node (e.g. a `SetOp`) reports the guarantee composed from its assembled children. +- Each unused branch (§1.3) returns `Unimplemented` from `output_schema`, `derive_*` and export. +- `search_workload*` panics on a root containing `ASAP`. - Rewrite the 117 `SummaryExpr::` assertions in `sql_to_post_asap.rs` / `promql_to_post_asap.rs` / `exact_composition.rs`. - Wire 6 round trip with a `DagInput` fragment; a version-5 document is rejected. - The 52 `execution_data_state.rs` tests keep their shapes; assertions read the `timing` slot instead of `ExecutionDataStateAssignment`. @@ -420,9 +514,13 @@ the same early flat export, on the final types. - Candidate pruning as an `ASAPOp` variant. - Accuracy through `SummaryMerge` / `Subtract` / `Delete` / `Join` (§2.2): an accuracy descriptor on state, or composing error along the state chain at readout. +- Holes: letting a chosen plan's child be filled by that child target's own choice at + assembly, instead of fixing it when the candidate is built. A search-strategy change, + independent of the types here. - Folding `ExactComposition` into `Subtree` (both of its forms become expressible); deferred until stage 5 is stable. ## 10. Open questions -None at present. +- Does ASAPQuery insert `SummaryMerge` only on the executable DAG, or through ASAPPlanner's + post-ASAP types? The planner-side variant is unimplemented (§1.3). From 7135c989be7fdf35bb968e830028980350f3f1de Mon Sep 17 00:00:00 2001 From: Selvomega Date: Tue, 29 Sep 2026 17:02:42 +0000 Subject: [PATCH 05/46] grilled one more time --- .../proposals/decoupling_op_and_expr.md | 3 +-- .../design_docs/proposals/operator-sharing.md | 23 +++++++++++++------ 2 files changed, 17 insertions(+), 9 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index b28727b4b..e24b87aa0 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -6,8 +6,7 @@ > approximate, measured on `main` at `8acb472`. **The idea.** `QueryExpr` holds two different kinds of node in one enum. This proposal -splits it into `NonASAPOp` (operators) and `ScalarExpr` (scalar expressions), so the -field a node sits in decides its type. +splits it into `NonASAPOp` (operators) and `ScalarExpr` (scalar expressions). ``` Today Proposed diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 44254c4cd..6ad34e65f 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -71,11 +71,18 @@ pub fn derive_guarantees( evidence: &dyn AccuracyEvidenceProvider, // evidence provider used for derivation memo: &mut DerivationMemo, ) -> Result, AccuracyError>; -/// Build the execution timing of one DAG root by derivation +/// Build the execution timing of one DAG root by derivation; the root runs at query time pub fn derive_timings( root: &Rc, memo: &mut DerivationMemo, ) -> Result, ExecutionDataStateError>; +/// Same, with the root's timing given, e.g. `IngestionTime` for a maintenance candidate +/// (like today's `validate_execution_data_states_at`) +pub fn derive_timings_at( + root: &Rc, + root_timing: ExecutionTiming, + memo: &mut DerivationMemo, +) -> Result, ExecutionDataStateError>; ``` - **Derivation recomputes**: `derive_*` keep the values set at construction (§2.2, §2.3) @@ -228,7 +235,7 @@ impl Field { A guarantee is filled in two steps: -1. **Binding** records local accuracy guarantee: a `SummaryEstimate`'s `local_guarantee` (the sketch's error over an exact input) and an exact `SummaryAgg`'s `exact_rule`. No `guarantee` slot is set yet. +1. **Binding** records local accuracy guarantee: a `SummaryEstimate`'s `local_guarantee` (the sketch's error over an exact input) and an exact `SummaryAgg`'s `exact_rule`. No `guarantee` slot is set yet. To size a sketch and check its target, binding still needs the child's error, as today: it runs `derive_guarantees` on the child with a fresh memo, reads the result, and drops it. 2. **`derive_guarantees`** fills every slot bottom-up: a node without an `ASAP` descendant is exact, and every other node composes its children's guarantees by its own rule. ``` @@ -243,7 +250,7 @@ Per node kind: | Node | Today | After | |---|---|---| | `SummaryEstimate` | stored at binding: the sketch's own error composed with the child's (`compose_guarantee`) | **derived**: `local_guarantee` composed with the child's. `local_guarantee` is set at binding: the sketch's error over an exact input, `None` when the model has no error model for the family | -| `SummaryAgg` | stored: ExactAggregate family composed with the child's; sketch families `None` | **derived**: ExactAggregate family: exact, composed with the child's under `exact_rule`; sketch families `Set(None)`, state has no guarantee | +| `SummaryAgg` | stored: ExactAggregate family composed with the child's; sketch families `None` | **derived**: ExactAggregate family: exact, composed with the child's under `exact_rule`, except `ExactKind::Count`, exact whatever the child (as today); sketch families `Set(None)`, state has no guarantee | | `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: exact
post-ASAP `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: composed at construction | **derived**: composed from the children; exact if no `ASAP` descendant | | `FinalizeExactAccumulator` | copies the child's | **derived**: the child's | | `MaintainPopulation` / `ReadPopulation` | stored: exact | **derived**: exact | @@ -362,7 +369,9 @@ shared with other queries. ```rust fn assemble(&self, t: &Rc) -> Rc { memo by ptr; // shared children stay one Rc - match self.chosen(t) { + let chosen = if query_time_nested_sum(t) { None } // as today: keep the outer SUM so the + else { self.chosen(t) }; // inner target's own choice is assembled + match chosen { Some(Subtree(r)) => r, // a complete plan, used as is Some(ExactComposition{..}) => composition.plan, None => t.map_children(|c| if is_target(c) { self.assemble(c) } else { c }), @@ -379,8 +388,8 @@ on that result, as today: lifecycle planning holds `Rc`s into the plan and reads root's guarantee, so it must see the derived tree. - **Illegal child** (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`): candidates - are checked when built, as `relink_summary` does today - (`validate_execution_data_states_at`). An error from `derive_timings` after assembly is + are checked when built with `derive_timings_at(candidate, placement's timing)`, as + `relink_summary` does today with `validate_execution_data_states_at`. An error from `derive_timings` after assembly is a bug, and planning fails with that error. - **A shared subtree read at two timings**, e.g. a query-time `Aggregate` and an ingestion-time `SummaryAgg` reading one `Scan`: both choices are legal, only the sharing @@ -484,7 +493,7 @@ lowering plus a `DagInput` arm (an incoming edge as a materialized table); the | 1 Split | [decoupling doc](decoupling_op_and_expr.md): `NonASAPOp` + `ScalarExpr`; children stay `Rc` | scalar code ([decoupling doc §3](decoupling_op_and_expr.md#3-changes)) | | 2 Two levels | §1.1, §1.4: `Operator`, an empty `ASAPOp`, `contains_asap()`, `expect_non_asap()`; child slots become `Rc>`; every variant gets `timing` / `guarantee` slots, and nodes are built through constructors that leave both `Unset` | every crate; the same mechanical change everywhere | | 3 One schema | §2.1: `Column` → `Field` and `Schema.columns` → `fields` (serde keeps the name `columns` until stage 4); `FieldType`, `ASAPType`, `PlainField`, `Schema` everywhere except the `executable_dag.rs` wire types, which keep `SummarySchema` until stage 4. **No wire change** | `asap-types` + schema construction in every crate | -| 4 New types | fill `ASAPOp`; `ASAP` arms of `output_schema`; `derive_guarantees` and `derive_timings` (§2.2, §5); the entry check (§3); `flatten(&SummaryNode) -> Rc` so export runs on the new types; wire types become `Schema`, and `Schema.fields` serializes as `fields`. Wire → 6. **The only wire-breaking stage**; merged together with ASAPQuery-backend and ASAPCollector | `asap-types`, `devtools`, viewer | +| 4 New types | fill `ASAPOp`; `ASAP` arms of `output_schema`; `derive_guarantees` and `derive_timings` (§2.2, §5); the entry check (§3); `flatten(&SummaryNode) -> Rc` so export runs on the new types, copying each node's guarantee and today's derived timing into the slots, so the export is unchanged; wire types become `Schema`, and `Schema.fields` serializes as `fields`. Wire → 6. **The only wire-breaking stage**; merged together with ASAPQuery-backend and ASAPCollector | `asap-types`, `devtools`, viewer | | 5 Planner | §4: candidates and assembly on `Rc`; §7 moves to the new types; delete `flatten` | `asap-aware-mapping` | | 6 Cleanup | delete `SummaryExpr`, `SummaryNode`, extra `ValueOperation` variants, `ExactOperation`, `post_asap/cse.rs`; update `post-asap-ir.md`, `physical-plan-integration.md`, developer and viewer docs | docs | From 751a0155fe85b9429839c68fb42b2de3789f0345 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 18:59:38 +0000 Subject: [PATCH 06/46] design document for operator sharing updated --- .../proposals/decoupling_op_and_expr.md | 14 +- .../design_docs/proposals/operator-sharing.md | 339 +++++++++++------- 2 files changed, 226 insertions(+), 127 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index 2e9a9342b..dc8d6f6ad 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -1,16 +1,16 @@ # Decoupling Operators From Scalar Expressions -> Status: proposed, not implemented. Companion to [Operator sharing](operator-sharing.md) -> (same PR): this document splits `QueryExpr`; that one builds the shared operator -> language on the result. Code is referenced by file and function; counts are -> approximate, measured on `main` at `8acb472`. +> Status: proposed, not implemented. Companion to [Operator sharing](operator-sharing.md): +> this document splits `QueryExpr`; that one builds the shared operator language on the +> result. Code is referenced by file and function; counts are approximate, measured on +> `main` at `5a32b8b`. **The idea.** `QueryExpr` holds two different kinds of node in one enum. This proposal splits it into `NonASAPOp` (operators) and `ScalarExpr` (scalar expressions). ``` Today Proposed -Filter { pred: Rc, Filter { pred: Predicate(Rc), +Filter { pred: Rc, Filter { pred: Predicate(ScalarExpr), child: Rc } child: Rc } ``` @@ -46,14 +46,14 @@ pub enum NonASAPOp { Scan { .. }, Filter { pred: Predicate, child: Rc> }, Project { cols: Vec>, child }, Aggregate { .. }, Join { .. }, SetOp { .. }, Concat { .. }, Dedup { .. }, Sort { .. }, Limit { .. }, BinaryOp { .. }, SQLWindowFunc { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, - ScalarBridge(Rc>), // formerly PromqlScalarBridge (the `2` in PromQL `v * 2`) + ScalarBridge(ScalarExpr), // formerly PromqlScalarBridge (the `2` in PromQL `v * 2`) EvalTimestamp, // PromQL time() } pub enum ScalarExpr { Column(C), Literal(ScalarValue), Compare { .. }, BoolAnd(..), BoolOr(..), Not(..), IsNull(..), IsNotNull(..), Cast { .. }, InList { .. }, FunctionCall { .. }, Arithmetic { .. }, Case { .. }, CurrentTimestamp, } -pub struct Predicate(pub Rc>); +pub struct Predicate(pub ScalarExpr); // scalars are held by value: CSE hashes them as plain data, nothing shares them pub struct ProjectItem { pub alias: Option, pub expr: ScalarExpr } ``` diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 03b36cf59..2b7f7d082 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -2,12 +2,19 @@ > - Status: proposed, not implemented. > - Problem statement: [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). -> - Builds on [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) (same PR), which splits `QueryExpr` into `NonASAPOp` and `ScalarExpr`. +> - Builds on [Decoupling operators from scalar expressions](decoupling_op_and_expr.md), which splits `QueryExpr` into `NonASAPOp` and `ScalarExpr`. +> - Code is referenced by file and function against `main` at `5a32b8b` (after #472, #478, #510). +> - Timing follows the planner layering of #480 / #509: a node's timing is written from the summary maintenance lifecycle assignment chosen for the DAG, never inferred from the IR (§2.3). **The idea.** Today a post-ASAP plan is glued together from two sets of operator types. This proposal keeps one operator language and makes summary operators extra node kinds in it: any relational operator can sit above a summary, and a summary can read any relational subtree. Nothing is wrapped and nothing is duplicated. +`Operator` has two levels, `NonASAP(NonASAPOp)` and `ASAP(ASAPOp)`, rather than one flat +enum of every variant: frontends, `resolve` and the per-operator export (§6) work on `NonASAP` +trees only, and `NonASAPOp` gives them a precise type for that instead of a run-time check +on each node. + ``` Today Proposed ValueOperation(Project) ← a copy NonASAP(Project) @@ -36,7 +43,7 @@ We define the `Operator` type structure based on the breadth of its attributes. | Applies to | Examples | Defined as | |---|---|---| | every operator | children, schema, timing, guarantee | methods implemented for `Operator` | -| one category | for all `NonASAP` operators, timing is derived from the consuming edge, and guarantee from the children | implementation specified to one enum branch of `Operator` | +| one category | for all `NonASAP` operators, guarantee is derived from the children, and the assigned timing is checked against the consuming edge | implementation specified to one enum branch of `Operator` | | one operator | `Aggregate.measures`, `SummaryAgg.family` | fields of that variant | ```rust @@ -56,11 +63,11 @@ impl Operator { // implemented for every operator } /// `Slot` represents a value that may be unset or set. -/// In the current design, it will be used to wrap the `timing` and `guarantee` values, -/// whose values will only be determined after derivation. +/// `guarantee` is filled by derivation (§2.2), `timing` by applying a lifecycle +/// assignment (§2.3). Both are `Unset` on a freshly built node. pub enum Slot { Unset, Set(T) } -/// Memo of one derivation, keyed by (node pointer, incoming timing). +/// Memo of one pass over a workload, keyed by (node pointer, timing). /// One is shared by every root of a workload, so a node shared by two roots stays one `Rc`. pub struct DerivationMemo { .. } @@ -71,24 +78,27 @@ pub fn derive_guarantees( evidence: &dyn AccuracyEvidenceProvider, // evidence provider used for derivation memo: &mut DerivationMemo, ) -> Result, AccuracyError>; -/// Build the execution timing of one DAG root by derivation; the root runs at query time -pub fn derive_timings( +/// Write the timings of one lifecycle assignment into the `timing` slots of one DAG +/// root, top-down, then validate them (§5). Summary materialization chooses the +/// assignment (§2.3); nothing is inferred from operator kinds. +pub fn apply_lifecycle_timings( root: &Rc, - memo: &mut DerivationMemo, -) -> Result, ExecutionDataStateError>; -/// Same, with the root's timing given, e.g. `IngestionTime` for a maintenance candidate -/// (like today's `validate_execution_data_states_at`) -pub fn derive_timings_at( - root: &Rc, - root_timing: ExecutionTiming, + assignment: &LifecycleAssignment, // per-node timings, expanded from per-state lifecycle choices (§10) memo: &mut DerivationMemo, ) -> Result, ExecutionDataStateError>; ``` -- **Derivation recomputes**: `derive_*` keep the values set at construction (§2.2, §2.3) - and recompute every other slot, so calling them again after a rewrite is safe. -- **Equality**: both slots take part in `PartialEq` and hashing, so CSE never merges two - nodes that differ in timing or guarantee. +- **Passes copy**: `derive_guarantees` and `apply_lifecycle_timings` return a new tree; + nodes are immutable and `with_*` build new ones. Pointers held before a pass + (`assembled_nodes`, planner memos) are not valid into its result; lifecycle selection + reads the guarantee-derived tree, export reads the timed tree (§2.4). +- **Passes recompute**: `derive_guarantees` keeps the values set at construction (§2.2) + and recomputes every other slot; `apply_lifecycle_timings` overwrites every `timing` + slot. Every rewrite runs before them, so running either again gives the same slots. +- **Equality**: the `timing` slot takes part in `PartialEq` and `Hash`, so CSE never merges + two nodes assigned different timings. The `guarantee` slot takes part in neither: it is + a function of the subtree, so equal subtrees derive equal guarantees, and + `ResultGuarantee` holds `f64` and has no `Hash`. Following diagram conceptually displays the structure of `Operator`: ```text @@ -122,7 +132,7 @@ pub enum NonASAPOp { BinaryOp { op, lhs, rhs }, SQLWindowFunc { args: Vec>, order_by: Vec>, child: Rc>, .. }, Dedup { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, - ScalarBridge(Rc>), // the `2` in PromQL `v * 2` + ScalarBridge(ScalarExpr), // the `2` in PromQL `v * 2` EvalTimestamp, // PromQL time() } ``` @@ -172,7 +182,7 @@ be `Rc` to allow free placement: ```rust // After the decoupling doc // After this proposal -Filter { pred: Predicate(Rc), Filter { pred: Predicate(Rc), +Filter { pred: Predicate(ScalarExpr), Filter { pred: Predicate(ScalarExpr), child: Rc } child: Rc } // NonASAP(..) or ASAP(SummaryEstimate ..) ``` @@ -183,13 +193,13 @@ cannot replace a branch — e.g. the branches of SQL `ROLLUP` or PromQL `histogr ## 2. Schema, Guarantee, and Timing -This section discusses three key per-node attributes, `schema`, `guarantee`, and `timing`, as well as how they are stored and derived in the new framework. +This section discusses three key per-node attributes, `schema`, `guarantee`, and `timing`, as well as how they are stored and filled in the new framework. | Field | Meaning | Today | After | |---|---|---|---| | `schema` | output columns and their types | pre-ASAP: computed by `QueryExpr::output_schema()`
post-ASAP: a `SummarySchema` stored on every `SummaryNode` | can be obtained by `output_schema()` | | `guarantee` | accuracy bound | pre-ASAP: none
post-ASAP: stored on every `SummaryNode` | can be obtained by `guarantee()`
binding stores only each operator's own error
complete error bound need to be derived by `derive_guarantees()` | -| `timing` | execution time | pre-ASAP: none
post-ASAP, stored: a field on `BinaryOp` / `ValueOperation` / `SummaryMerge`
post-ASAP, not stored: `KeepPreAsap` from the consuming edge, `SummaryAgg` from the child. | can be obtained by `timing()`
set by binding (`SummaryAgg`) or the planner (`FinalizeExactAccumulator`)
timing of the rest of operators need to be derived by `derive_timings()` | +| `timing` | execution time | pre-ASAP: none
post-ASAP, stored: a field on `BinaryOp` / `ValueOperation` / `SummaryMerge`
post-ASAP, not stored: `KeepPreAsap` from the consuming edge, `SummaryAgg` from the child. | can be obtained by `timing()`
`Unset` after assembly
written for every node by `apply_lifecycle_timings()` from the chosen lifecycle assignment (§2.3) | ### 2.1 Schema: fused into one type @@ -206,6 +216,10 @@ We implement the new schema type based on the original `Schema` type used in pre `SummarySchema` / `SummaryField` are then redundant and deleted. +`Schema` keeps the reserved column `PROMQL_SERIES_IDENTITY` (`"$promql_series_identity"`, +`pre_asap/schema.rs`) and `has_promql_series_identity()`: `maintained_population.rs` uses +it to decide whether a closed PromQL schema still identifies a series (§7). + Detailed code design is shown below. ```rust pub enum FieldType { DataType(DataType), ASAPType(ASAPType) } @@ -256,42 +270,54 @@ Per node kind: | `MaintainPopulation` / `ReadPopulation` | stored: exact | **derived**: exact | | unused variants | `None`: state has no guarantee of its own | unimplemented (§1.3) | -### 2.3 Timing: set where position does not decide it - -A timing is filled in two steps: - -1. **Binding** sets every `SummaryAgg` to the timing today's fallback derives: - `IngestionTime`, or `QueryTime` when binding built its child at query time (a - query-time `FinalizeExactAccumulator`). The planner sets every - `FinalizeExactAccumulator`. -2. **`derive_timings`** runs on each assembled root, sharing one memo, top-down: a root is query - time, a set node keeps its value, a node of fixed kind takes that kind's time, and - every other node takes its parent's. A node reached at two timings is copied (§4). +### 2.3 Timing: written from a lifecycle assignment + +The logical layer — PlanSpace, binding, assembly — decides *what* to compute, not when. +Timing is chosen by summary materialization: for every unique summary state it picks a +lifecycle (maintain at ingestion time or recompute at query time, with window and +retention), and that choice fixes the timing of every node that feeds or reads the state. +One logical DAG can therefore have several assignments; the deployment chooses among +them with its own costs ([Output layers](../architecture/input-output-workflow.md#output-layers), #480). + +Nothing in the IR sets a timing. Binding sets no `SummaryAgg.timing`; the planner sets +none on `FinalizeExactAccumulator` (`finalize_query_candidate`, §4, still inserts the +node, its timing is assigned like any other). After assembly every `timing` slot is +`Unset`. `apply_lifecycle_timings` writes the chosen assignment into the slots, top-down +per root with one shared memo, and validates it (§5): an assignment under which +ingestion work depends on a query-time result, or a node of fixed kind gets the wrong +timing, is rejected. A node assigned two timings is copied (§4). `Unset` means no +assignment was applied; physical compilation and export reject it. ``` - binding derive_timings -Project Unset QueryTime ← root - SummaryEstimate Unset QueryTime ← fixed by kind - SummaryAgg IngestionTime IngestionTime ← kept - Scan t Unset IngestionTime ← from parent + assembly after apply_lifecycle_timings +Project Unset QueryTime + SummaryEstimate Unset QueryTime + SummaryAgg Unset IngestionTime ← the state's lifecycle: maintained + Scan t Unset IngestionTime ← feeds a maintained state ``` -Exported timings are unchanged. Moving the `SummaryAgg` default to `QueryTime`, and -letting the lifecycle step choose ingestion time, is a separate PR. +Some candidates fix a timing when they are built today: exact compositions carry +`OperationPlacement::Read` / `Maintenance` (provenance `ValueOperationAtQueryTime` / +`ValueOperationAtIngestionTime`), maintained populations are built at ingestion time, +and #472's grouped `Rate`→`Sum` pair. Each becomes one logical candidate whose placement +is a lifecycle choice (§8 stage 5). The **default assignment** reproduces today's +timings — a `SummaryAgg` maintained at ingestion time, everything above a readout at +query time — so exported timings do not change until lifecycle selection chooses +otherwise. Per node kind: | Node | Today | After | |---|---|---| -| `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: from the consuming edge
post-ASAP `ValueOperation` / `BinaryOp` copies: a stored field | **derived** from the consuming edge (§5) | -| `SummaryAgg` | from the child; ingestion time under `KeepPreAsap` | **set** by binding, as today's fallback: `IngestionTime`, or `QueryTime` over a query-time child | -| `FinalizeExactAccumulator` | a stored field, set by the planner | **set** by the planner: the same position allows either time | -| `SummaryEstimate` | query time, fixed by the kind | **derived** from the kind: query time | -| `MaintainPopulation` / `ReadPopulation` | a stored field, always ingestion / query time | **derived** from the kind: ingestion / query time | +| `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: from the consuming edge
post-ASAP `ValueOperation` / `BinaryOp` copies: a stored field | **assigned**; validated against its consuming edges (§5) | +| `SummaryAgg` | from the child; ingestion time under `KeepPreAsap` | **assigned** by its state's lifecycle | +| `FinalizeExactAccumulator` | a stored field, set by the planner | **assigned**; the position allows either time | +| `SummaryEstimate` | query time, fixed by the kind | **assigned**; validated: query time only | +| `MaintainPopulation` / `ReadPopulation` | a stored field, always ingestion / query time | **assigned**; validated: ingestion / query time only | | unused variants | `SummaryMerge`: a stored field; `Join` / `Subtract` / `Delete`: ingestion time | unimplemented (§1.3) | -Unlike a guarantee, a timing depends on the parents, so `derive_timings` needs the whole -DAG and runs only after assembly. +Unlike a guarantee, a timing is a property of the whole DAG and its assignment, so it is +applied only after assembly. ### 2.4 Workflow of setting up `guarantee` and `timing`: today vs. after @@ -311,14 +337,14 @@ export validate_execution_data_states, per root, derives the remaini After: ``` -search / binding sets only what cannot be derived: SummaryAgg.timing, - local_guarantee, exact_rule; checks accuracy on a derived copy, - then drops the copy +search / binding sets only what cannot be derived: local_guarantee, exact_rule; + no timing; checks accuracy on a derived copy, keeps the result on + the candidate record (§4), drops the copy assembly builds each root; kept NonASAP nodes stay as they are -derive_timings per root, one shared memo: fills timings top-down, copies a node - read at two timings derive_guarantees per root, one shared memo: fills guarantees bottom-up -lifecycle reads the derived guarantee +lifecycle reads the derived guarantee; chooses a lifecycle per summary state +apply_lifecycle_ per root, one shared memo: writes the assignment's timings + timings top-down, validates them, copies a node assigned two timings export reads the slots; rejects an Unset one ``` @@ -339,6 +365,30 @@ impl Operator { } ``` +Library users do not call `search_workload*` themselves. Since #478 the external +boundary is `asap_aware_mapping::pass::optimize` (reached from `asap_planner::e2e_plan`), +and since #510 a pass always produces lifecycle-aware plans: + +``` +asap_planner::e2e_plan + → asap_aware_mapping::pass::optimize + → MajorPass::optimize + → search_workload_with_targets + → global_selection_with_summary_maintenance_lifecycles + → assemble_selected_dag_with_summary_maintenance_lifecycles (one call per root) +``` + +```rust +pub struct QueryLifecyclePlan { pub entry_index: usize, pub plan: SummaryMaintenanceLifecyclePlan } +pub struct PlanOutput { pub plans: Vec } // one per workload entry +impl PlanOutput { pub fn dags(&self) -> Vec> } // each plan.root; becomes Rc +``` + +`ParsedWorkload` (`asap_types::parsed_workload`) holds the roots as `Vec>` +and becomes `Vec>`. The entry check above stays in `search_cse_workload_with`; +the pass layer adds no check of its own. See +[`updated_interface_with_pluggable_optimization.md`](../architecture/updated_interface_with_pluggable_optimization.md). + A compile-time alternative — an associated type on `ColState` with `ColumnRef::ASAP = Never` — only protects frontend code before `resolve`: frontends already return `ColumnId` trees, where `ASAP` is allowed. The entry check covers every @@ -353,6 +403,12 @@ pub enum Replacement { } ``` +`runtime_support_evidence` (`replacement.rs`) asks the cost model whether the runtime +supports a candidate, per variant: `Summary` → `summary_support_evidence`, `Rewrite` → +`Some(true)`. With one `Subtree` variant it dispatches on the root: an `ASAP` root → +`summary_support_evidence`, a `NonASAP` root → `Some(true)`. `ASAP` nodes below a +`NonASAP` root were each asked when they were candidates themselves. + **Candidates stay bottom-up, as today**: a candidate is built on a concrete child plan (`realize_child_with`, or each child candidate in `prepare_compositions`), so a chosen plan is complete. Where `realize_child_with` falls back to `keep_pre_asap` today, it @@ -361,16 +417,20 @@ returns the child's original subtree, and assembly keeps it as is. **Accuracy check during search**: binding sets no `guarantee` slot (§2.2), so the candidate filter in `search_workload_with_targets` and `prepare_compositions` run `derive_guarantees` on the candidate alone, with a fresh memo, then check its accuracy target. The derived -copy is only read, then dropped: PlanSpace keeps the original candidate, whose nodes are -shared with other queries. +copy is dropped; its root guarantee is stored on the candidate's `ReplacementSubDAG`, where +selection reads it (today's stored `guarantee` moved from the node to the candidate record). +PlanSpace keeps the original candidate, whose nodes are shared with other queries. **Assembly** — one rule replaces `assemble_residual`: ```rust fn assemble(&self, t: &Rc) -> Rc { memo by ptr; // shared children stay one Rc - let chosen = if query_time_nested_sum(t) { None } // as today: keep the outer SUM so the - else { self.chosen(t) }; // inner target's own choice is assembled + let composed = matches!(self.chosen(t), // the chosen summary already folds the + Some(Subtree(r)) if r is ASAP(SummaryAgg { child, .. }) // inner aggregate: nothing is hidden + && !child.contains_asap() && !contains_aggregate(child)); + let chosen = if query_time_nested_sum(t) && !composed { None } // as today: keep the outer SUM so the + else { self.chosen(t) }; // inner target's own choice is assembled match chosen { Some(Subtree(r)) => r, // a complete plan, used as is Some(ExactComposition{..}) => composition.plan, @@ -380,91 +440,117 @@ fn assemble(&self, t: &Rc) -> Rc { } ``` -Then `GlobalSelection::assemble_selected_dag` runs `derive_timings` (§5) → -`derive_guarantees` (§2.2) on each root it assembles. The `DerivationMemo` lives on -`GlobalSelection` next to `assembled_nodes`, so roots assembled one call at a time still -share nodes. `assemble_selected_dag_with_summary_maintenance_lifecycles` plans lifecycles -on that result, as today: lifecycle planning holds `Rc`s into the plan and reads the -root's guarantee, so it must see the derived tree. - -- **Illegal child** (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`): candidates - are checked when built with `derive_timings_at(candidate, placement's timing)`, as - `relink_summary` does today with `validate_execution_data_states_at`. An error from `derive_timings` after assembly is - a bug, and planning fails with that error. -- **A shared subtree read at two timings**, e.g. a query-time `Aggregate` and an - ingestion-time `SummaryAgg` reading one `Scan`: both choices are legal, only the sharing - is not. `derive_timings` memoizes by (pointer, timing), so it builds one copy per - timing; a subtree read at one timing stays one `Rc`, within a root or across roots. +Then `GlobalSelection::assemble_selected_dag` runs `derive_guarantees` (§2.2) on each +root it assembles. The `DerivationMemo` lives on `GlobalSelection` next to +`assembled_nodes`, so roots assembled one call at a time still share nodes. +`assemble_selected_dag_with_summary_maintenance_lifecycles` plans lifecycles on that +result, as today: lifecycle planning holds `Rc`s into the plan and reads the root's +guarantee, so it must see the derived tree. The assignment it chooses is then written +with `apply_lifecycle_timings` (§2.3, §5). + +`assemble_selected_query` stays the boundary for a query result (#472): it runs +`assemble_selected_dag`, then `finalize_query_candidate`, which puts a query-time +`FinalizeExactAccumulator` between maintained exact state and the query-time consumer +(§2.3). The other callers of `finalize_query_candidate` — binding and the two sides of +`BinaryOp` and `Join` — keep calling it, on `Rc`. Today +`assemble_selected_dag_with_summary_maintenance_lifecycles` (`summary_maintenance_lifecycle.rs`) +calls `assemble_selected_dag` directly and skips that step; this proposal leaves that as +it is (§9). + +- **Illegal placement** (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`): + candidates carry no timing, so nothing is checked when they are built. The check moves + to `apply_lifecycle_timings`, which rejects an assignment that places a node illegally. + Summary materialization offers only assignments it has validated (today + `relink_summary` runs the same check through `validate_execution_data_states_at`), so an + error from `apply_lifecycle_timings` on a chosen assignment is a bug, and planning fails + with that error. +- **A shared subtree assigned two timings**, e.g. a query-time `Aggregate` and an + ingestion-time `SummaryAgg` reading one `Scan`: both placements are legal, only the + sharing is not. `apply_lifecycle_timings` memoizes by (pointer, timing), so it builds one + copy per timing; a subtree assigned one timing stays one `Rc`, within a root or across + roots. `map_children` is `rebuild_children` from `pre_asap/cse.rs`, dispatching to `NonASAPOp::map_children` / `ASAPOp::map_children`. Deleted: `assemble_residual`, `keep_pre_asap` / `keep_pre_asap_rc`, and the -`KeepPreAsap` branch of `finalize_exact_accumulator`. Kept: `relink_summary` and the -`query_time_nested_sum` special case, which pick a child after selection today. +`KeepPreAsap` branch of `finalize_exact_accumulator`. Kept: `relink_summary`, +`assemble_selected_query` / `finalize_query_candidate`, and the `query_time_nested_sum` +special case with its #472 exception — the chosen candidate is a `SummaryAgg` whose +child is a `NonASAP` subtree without an `Aggregate` — where `contains_aggregate` takes an +`Operator` instead of a `QueryExpr`. | #468 problem | Resolution | |---|---| | 1. A `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside | one set of types | -| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds when both run at the same time; otherwise the scan is copied (above). With today's defaults the sketch runs at ingestion time and the exact `Aggregate` at query time, so they share only after the `QueryTime` default (separate PR, §2.3). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | +| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds when both run at the same time; otherwise the scan is copied (above). They share under an assignment that recomputes the sketch at query time; the default assignment maintains it at ingestion time, so there the scan is copied (§2.3). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | | 3. `SetOp` and similar have no post-ASAP copy, so no summary below them | `SetOp` takes `None => t`; both children are assembled | -## 5. Timing: execution data states +## 5. Timing: validating the assignment -`validate_execution_data_states` becomes `derive_timings`: instead of returning the -`ExecutionDataStateAssignment` side table (deleted), it writes each node's `timing` slot. -`produced_data_state(KeepPreAsap) = None` (set by the consuming edge) extends to all -`NonASAPOp`s: +`validate_execution_data_states` becomes the validation half of `apply_lifecycle_timings`. +The `ExecutionDataStateAssignment` side table is deleted; instead of deriving a state per +node, the pass checks the timing the assignment wrote into each slot against the node's +kind and its consuming edges. Nothing is derived from the consuming edge any more: | Node | Produced state | |---|---| -| `NonASAPOp` | set by the consuming edge (`QUERY_ROWS` at the root), passed to its children | -| `SummaryAgg` | `{timing, SummaryState}` from its `timing` slot (§2.3); a `NonASAPOp` child takes the same timing | -| other `ASAPOp` | unchanged | +| `NonASAPOp` | `{assigned timing, Raw / …}`; checked against every consuming edge (`QUERY_ROWS` at the root) | +| `SummaryAgg` | `{assigned timing, SummaryState}` (§2.3) | +| other `ASAPOp` | today's rules, checked against the assigned timing | A data state is the `timing` slot plus a primitive (`Raw` / `SummaryState` / …) fixed by the kind; only the timing is stored. The `KeepPreAsap` / `BinaryOp` / `ValueOperation` / `RelationalJoin` arms of today's -`validate_execution_data_states` merge into one `NonASAP` arm of `derive_timings`: +`validate_execution_data_states` merge into one `NonASAP` arm of that validation: -- Pass the state to each child; an `ASAP` child is checked by the `ASAP` edge rules. +- Check the edge to each child; an `ASAP` child is checked by the `ASAP` edge rules. + Ingestion work cannot depend on a query-time result. - `check_plain_operands` stays: referenced columns must be `FieldType::DataType` (`Project` / `Filter` / `Sort` / `Limit` may pass `ExactAggregate` columns through). This rejects `Project(ASAP(SummaryAgg))`. - `BinaryOp`'s ingestion-side constraints move into this arm. -- `AmbiguousKeepPreAsap` is deleted: a subtree read at two timings is copied (§4). +- `AmbiguousKeepPreAsap` is deleted: a subtree assigned two timings is copied (§4). -## 6. Export: fragments in the post-ASAP DAG +## 6. Export: one post-ASAP node per operator The four original-operator payloads (`fallback{expression: QueryExpr}`, `binary`, `value`, `relational_join`) become one: ```rust PostAsapOperatorPayload::Relational { - /// No ASAP node inside. Leaves are Scans, or Scan { source: Source::DagInput { role } } - /// for an incoming edge whose schema is the edge's intermediate_schema. - expression: Operator, + /// One non-ASAP operator: its kind, expressions and parameters, without its child + /// fields. The children are this node's incoming edges, in child-role order. + operator: NonASAPOpKind, } ``` -`compile_post_asap_dag` takes each **largest connected subtree without `ASAP`** as one -fragment, cutting an edge with a `DagInput` leaf wherever it meets an `ASAP` node. -`ASAP` nodes map one-to-one onto the existing summary payloads; +Export emits **one post-ASAP node per `NonASAP` operator**, exactly as it already does +per `ASAP` operator: children become edges, leaves are `Scan` nodes. No subtree is +embedded in a node, so there is no `DagInput` placeholder and no fragment. A `NonASAP` +node shared by two consumers (one `Rc`, kept by CSE, §4) is exported once, with two +outgoing edges. Physical compilation then corresponds node by node: each post-ASAP node +lowers to one physical operator, or to a few helper operators numbered from it. `ASAP` +nodes map one-to-one onto the existing summary payloads; `FinalizeExactAccumulator` / `MaintainPopulation` / `ReadPopulation` stay -`value{operation}`. A backend lowers every fragment with its existing `QueryExpr` -lowering plus a `DagInput` arm (an incoming edge as a materialized table); the -`binary` / `value::Project` / `relational_join` lowerings go. +`value{operation}`. The `fallback` whole-expression lowering and the `binary` / +`value::Project` / `relational_join` special cases go; every `Relational` node is lowered +by one per-operator lowering that reads its inputs from its edges. -- **Wire 5 → 6**: three fewer payloads; `fallback` becomes `relational` with `DagInput` - leaves; `output_schema` / `intermediate_schema` become `Schema`. One cutover (§8 +- **Wire 5 → 6**: three fewer payloads; `fallback` becomes `relational` nodes, one per + operator; `output_schema` / `intermediate_schema` become `Schema`. One cutover (§8 stage 4), together with the downstream readers. -- **Timing and guarantee** are read from the node slots; an `Unset` slot is rejected. - An edge's `data_state` is its producer's timing plus the primitive of its kind (§5). - `compile_post_asap_dag` no longer re-runs data-state validation. +- **Timing and guarantee** are read from the node slots: timing as assigned, guarantee as + derived. An `Unset` slot is rejected; for timing it means no assignment was applied. + An edge's `data_state` is its producer's assigned timing plus the primitive of its kind + (§5). `compile_post_asap_dag` splits the precompute and query DAGs by that timing and + no longer re-runs data-state validation. - **`SummaryMerge`** stays a wire payload, although its planner-side variant is unimplemented (§1.3, §10). -- **Phases** become per fragment. Switching phase inside a fragment would need a - materialization point, and those are `ASAP` nodes, so nothing is lost. +- **Phases** are per node, from the assigned timing. A phase switch between two + `NonASAP` nodes must satisfy §5 (ingestion work cannot read a query-time result); + summary materialization places materialization points only at `ASAP` state, so an + assignment that would need one elsewhere is rejected when applied. ## 7. Other consumers @@ -472,10 +558,15 @@ lowering plus a `DagInput` arm (an incoming edge as a materialized table); the |---|---| | `post_asap/cse.rs` | delete; `share_common_subtrees` covers `ASAPOp` (derives `PartialEq` + serde) | | `dag_export.rs` | delete `build_summary` / `build_summary_hybrid` / `summary_kind_tag`; one exporter with an `ASAP` arm; update the viewer's `node-style.js` and the pin test `viewer_categorizes_exactly_the_exported_node_kinds` | -| `summary_maintenance_cost/estimator.rs` (80 sites) | `KeepPreAsap` branches (`query_source_selections`, `retained_queries`) use the §6 fragment; `exact_binary` / `value_operation` costs fold into it. Also fixes the missing `RelationalJoin` arm in `summary_operation_evidence` | -| `physical_plan_cost_model.rs::estimate_candidate` | every fragment goes through `lower_query_physical_dag` | -| `summary_maintenance_lifecycle.rs` | `selected_raw_recompute` becomes `!contains_asap(root)`; the `keep_pre_asap(target)` fallback in `assemble_selected_dag_with_summary_maintenance_lifecycles` becomes `target` | -| `maintained_population.rs` | `KeepPreAsap(source)` becomes `source`; `population.matches_input` reads an `NonASAP` child directly | +| `summary_maintenance_cost/estimator.rs` (90 `SummaryExpr::` sites, 13 `KeepPreAsap`) | `KeepPreAsap` branches (`query_source_selections`, `retained_queries`) use the §6 `Relational` nodes; `exact_binary` / `value_operation` costs fold into it | +| `summary_maintenance_cost/evidence.rs` | `summary_operation_evidence` gets an `ASAP` arm; also fixes its missing `RelationalJoin` case | +| `physical_plan_cost_model.rs::estimate_candidate` | every `Relational` node goes through the per-operator lowering (§6) in `lower_query_physical_dag` | +| `summary_maintenance_lifecycle.rs` | `SummaryMaintenanceLifecyclePlan.root` becomes `Rc`; `selected_raw_recompute` becomes `!contains_asap(root)`; the `keep_pre_asap(target)` fallback in `assemble_selected_dag_with_summary_maintenance_lifecycles` becomes `target` | +| `pass/mod.rs`, `pass/major.rs` | `PlanOutput::dags()` returns `Vec>`; `MajorPass` otherwise unchanged (§3) | +| `types/parsed_workload.rs` | `ParsedWorkload` roots become `Rc` (§3) | +| `asap-planner` (`planner/src/lib.rs`, `tests/e2e_plan.rs`) | follows `PlanOutput` | +| `maintained_population.rs` | `KeepPreAsap(source)` becomes `source`; `population.matches_input` reads an `NonASAP` child directly and keeps the `has_promql_series_identity()` check (§2.1) | +| `replacement.rs::enumerate_candidate_dags`, `CandidateDagInventory` | walk `Rc` instead of `SummaryNode`; internal and test use only | | `exact_composition.rs` | `ExactOperation::Aggregate` becomes a `NonASAP(Aggregate)`, built over each child candidate as `prepare_compositions` does today | | `RelationalJoin.pruning` | never set to `Some` in production; delete. Candidate pruning can return as an `ASAPOp` variant | @@ -491,11 +582,11 @@ lowering plus a `DagInput` arm (an incoming edge as a materialized table); the |---|---|---| | 0 Preparation | `Rc` for `Concat.children`; `rebuild_children` → `map_children`; `Column::plain` | `asap-types` | | 1 Split | [decoupling doc](decoupling_op_and_expr.md): `NonASAPOp` + `ScalarExpr`; children stay `Rc` | scalar code ([decoupling doc §3](decoupling_op_and_expr.md#3-changes)) | -| 2 Two levels | §1.1, §1.4: `Operator`, an empty `ASAPOp`, `contains_asap()`, `expect_non_asap()`; child slots become `Rc>`; every variant gets `timing` / `guarantee` slots, and nodes are built through constructors that leave both `Unset` | every crate; the same mechanical change everywhere | +| 2 Two levels | §1.1, §1.4: `Operator`, an empty `ASAPOp`, `contains_asap()`, `expect_non_asap()`; child slots become `Rc>`; every variant gets `timing` / `guarantee` slots, and nodes are built through constructors that leave both `Unset` (`timing` stays `Unset` until an assignment is applied) | every crate incl. `asap-planner`; the same mechanical change everywhere | | 3 One schema | §2.1: `Column` → `Field` and `Schema.columns` → `fields` (serde keeps the name `columns` until stage 4); `FieldType`, `ASAPType`, `PlainField`, `Schema` everywhere except the `post_asap_dag.rs` wire types, which keep `SummarySchema` until stage 4. **No wire change** | `asap-types` + schema construction in every crate | -| 4 New types | fill `ASAPOp`; `ASAP` arms of `output_schema`; `derive_guarantees` and `derive_timings` (§2.2, §5); the entry check (§3); `flatten(&SummaryNode) -> Rc` so export runs on the new types, copying each node's guarantee and today's derived timing into the slots, so the export is unchanged; wire types become `Schema`, and `Schema.fields` serializes as `fields`. Wire → 6. **The only wire-breaking stage**; merged together with ASAPQuery-backend and ASAPCollector | `asap-types`, `devtools`, viewer | -| 5 Planner | §4: candidates and assembly on `Rc`; §7 moves to the new types; delete `flatten` | `asap-aware-mapping` | -| 6 Cleanup | delete `SummaryExpr`, `SummaryNode`, extra `ValueOperation` variants, `ExactOperation`, `post_asap/cse.rs`; update `post-asap-ir.md`, `physical-plan-integration.md`, developer and viewer docs | docs | +| 4 New types | fill `ASAPOp`; `ASAP` arms of `output_schema`; `derive_guarantees` and `apply_lifecycle_timings` (§2.2, §2.3, §5); the entry check (§3); `flatten(&SummaryNode) -> Rc` so export runs on the new types, copying each node's guarantee into its slot and applying today's timings as the initial assignment, so the export is unchanged; wire types become `Schema`, and `Schema.fields` serializes as `fields`. Wire → 6. **The only wire-breaking stage**; merged together with ASAPQuery-backend and ASAPCollector | `asap-types`, `devtools`, viewer | +| 5 Planner | §4: candidates and assembly on `Rc`; §7 moves to the new types; delete `flatten`. Timing moves to the lifecycle: binding and candidates set none; the candidates that fix a timing today (§2.3) become one logical candidate each, their placement a lifecycle choice; paths without lifecycle selection apply the default assignment that reproduces today's timings | `asap-aware-mapping`, `asap-planner` | +| 6 Cleanup | delete `SummaryExpr`, `SummaryNode`, extra `ValueOperation` variants, `ExactOperation`, `post_asap/cse.rs`, and today's timing fallbacks (`produced_data_state` defaults, `validate_execution_data_states_at`); `PlanOutput::dags()` and `ParsedWorkload` lose their `SummaryNode` / `QueryExpr` types; update `post-asap-ir.md` (execution phase comes from the assignment), `physical-plan-integration.md`, `updated_interface_with_pluggable_optimization.md`, developer and viewer docs | `asap-planner`, docs | Wrapping pre-ASAP operators in `ValueOperation` first is not planned: stage 4 gives the same early flat export, on the final types. @@ -504,18 +595,19 @@ the same early flat export, on the final types. - One integration test per #468 problem: 1. `WITH metric AS (SELECT avg(CASE WHEN l_quantity BETWEEN 1 AND 50 THEN 1.0 ELSE 0.0 END) AS in_range FROM lineitem) SELECT in_range, in_range = 1.0 AS ok FROM metric` — no post-ASAP-only node besides `ASAP`; all `Project`s are one variant. - 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` when both run at the same time (once a binding rule splits measures); with today's ingestion-time default the `Scan` is copied. + 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` when both run at the same time (once a binding rule splits measures); under the default assignment, which maintains the sketch at ingestion time, the `Scan` is copied. 3. `SELECT approx_distinct(l_partkey) FROM lineitem UNION ALL SELECT approx_distinct(l_suppkey) FROM lineitem` — each side of the `SetOp` has a `SummaryEstimate`. -- A shared `Scan` read at two timings is copied once per timing; read at one timing, it stays one `Rc`. -- A node shared by two roots, assembled in two calls, is still one `Rc` after `derive_*`. -- After `derive_*`, no slot is `Unset`; export rejects a tree with one. -- Exported timings of today's plans are unchanged. +- A shared `Scan` assigned two timings is copied once per timing; assigned one timing, it stays one `Rc`. +- A node shared by two roots, assembled in two calls, is still one `Rc` after `derive_guarantees` and `apply_lifecycle_timings`. +- After both passes no slot is `Unset`; export rejects a tree with one, and a tree with no assignment applied. +- Exported timings of today's plans are unchanged under the default assignment. +- `apply_lifecycle_timings` rejects an assignment that puts ingestion work over a query-time result, and one that gives a `SummaryEstimate` ingestion time. - A kept `NonASAP` node (e.g. a `SetOp`) reports the guarantee composed from its assembled children. -- Each unused branch (§1.3) returns `Unimplemented` from `output_schema`, `derive_*` and export. +- Each unused branch (§1.3) returns `Unimplemented` from `output_schema`, `derive_guarantees`, `apply_lifecycle_timings` and export. - `search_workload*` panics on a root containing `ASAP`. -- Rewrite the 117 `SummaryExpr::` assertions in `sql_to_post_asap.rs` / `promql_to_post_asap.rs` / `exact_composition.rs`. -- Wire 6 round trip with a `DagInput` fragment; a version-5 document is rejected. -- The 52 `execution_data_state.rs` tests keep their shapes; assertions read the `timing` slot instead of `ExecutionDataStateAssignment`. +- Rewrite the 110 `SummaryExpr::` assertions in `sql_to_post_asap.rs` / `promql_to_post_asap.rs` / `exact_composition.rs`. +- Wire 6 round trip of a DAG with `Relational` nodes both above and below `ASAP` nodes; a version-5 document is rejected. +- The 16 `execution_data_state.rs` tests keep their shapes; they apply an assignment and read the `timing` slot instead of `ExecutionDataStateAssignment`. ## 9. Out of scope @@ -528,8 +620,15 @@ the same early flat export, on the final types. independent of the types here. - Folding `ExactComposition` into `Subtree` (both of its forms become expressible); deferred until stage 5 is stable. +- `assemble_selected_dag_with_summary_maintenance_lifecycles`, the only assembly `MajorPass` + uses, calls `assemble_selected_dag` rather than `assemble_selected_query` (§4), so pass + output carries no root `FinalizeExactAccumulator`. Not changed here. ## 10. Open questions +- `LifecycleAssignment` (§1.1): #482 adds `SummaryMaintenanceLifecyclePlan::execution_timed_dag()`, + which expands per-state lifecycle choices into per-node timings. This proposal should + take that type as the assignment rather than define its own. + - Does ASAPQuery insert `SummaryMerge` only on the exported post-ASAP DAG, or through ASAPPlanner's post-ASAP types? The planner-side variant is unimplemented (§1.3). From be5c43d25642565926b0bad4b2987ca69a330ff8 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 19:32:00 +0000 Subject: [PATCH 07/46] updated the document --- .../design_docs/proposals/operator-sharing.md | 40 ++++++++++++------- 1 file changed, 26 insertions(+), 14 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 2b7f7d082..cc9420c57 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -67,7 +67,7 @@ impl Operator { // implemented for every operator /// assignment (§2.3). Both are `Unset` on a freshly built node. pub enum Slot { Unset, Set(T) } -/// Memo of one pass over a workload, keyed by (node pointer, timing). +/// Memo of one pass over a workload, keyed by node pointer. /// One is shared by every root of a workload, so a node shared by two roots stays one `Rc`. pub struct DerivationMemo { .. } @@ -80,7 +80,8 @@ pub fn derive_guarantees( ) -> Result, AccuracyError>; /// Write the timings of one lifecycle assignment into the `timing` slots of one DAG /// root, top-down, then validate them (§5). Summary materialization chooses the -/// assignment (§2.3); nothing is inferred from operator kinds. +/// assignment (§2.3); nothing is inferred from operator kinds. A shared node reached +/// with two different timings is an error (§4). pub fn apply_lifecycle_timings( root: &Rc, assignment: &LifecycleAssignment, // per-node timings, expanded from per-state lifecycle choices (§10) @@ -254,7 +255,7 @@ A guarantee is filled in two steps: ``` Project p99 ±1% ← the child's - SummaryEstimate p99 ±1% ← local ±1%, composed with the child's + SummaryEstimate p99 ±1% ← local ±1%, composed with Scan t's: looks through the SummaryAgg SummaryAgg(Kll) None ← state has no guarantee Scan t exact ← no ASAP descendant ``` @@ -263,7 +264,7 @@ Per node kind: | Node | Today | After | |---|---|---| -| `SummaryEstimate` | stored at binding: the sketch's own error composed with the child's (`compose_guarantee`) | **derived**: `local_guarantee` composed with the child's. `local_guarantee` is set at binding: the sketch's error over an exact input, `None` when the model has no error model for the family | +| `SummaryEstimate` | stored at binding: the sketch's own error composed with the child's (`compose_guarantee`) | **derived**: `local_guarantee` composed with the guarantee of the state's input, i.e. the child of the `SummaryAgg` below (the `SummaryAgg` itself has none). `local_guarantee` is set at binding: the sketch's error over an exact input, `None` when the model has no error model for the family | | `SummaryAgg` | stored: ExactAggregate family composed with the child's; sketch families `None` | **derived**: ExactAggregate family: exact, composed with the child's under `exact_rule`, except `ExactKind::Count`, exact whatever the child (as today); sketch families `Set(None)`, state has no guarantee | | `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: exact
post-ASAP `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: composed at construction | **derived**: composed from the children; exact if no `ASAP` descendant | | `FinalizeExactAccumulator` | copies the child's | **derived**: the child's | @@ -285,8 +286,8 @@ node, its timing is assigned like any other). After assembly every `timing` slot `Unset`. `apply_lifecycle_timings` writes the chosen assignment into the slots, top-down per root with one shared memo, and validates it (§5): an assignment under which ingestion work depends on a query-time result, or a node of fixed kind gets the wrong -timing, is rejected. A node assigned two timings is copied (§4). `Unset` means no -assignment was applied; physical compilation and export reject it. +timing, is rejected, as is one that gives a shared node two timings (§4). `Unset` means +no assignment was applied; physical compilation and export reject it. ``` assembly after apply_lifecycle_timings @@ -344,7 +345,7 @@ assembly builds each root; kept NonASAP nodes stay as they are derive_guarantees per root, one shared memo: fills guarantees bottom-up lifecycle reads the derived guarantee; chooses a lifecycle per summary state apply_lifecycle_ per root, one shared memo: writes the assignment's timings - timings top-down, validates them, copies a node assigned two timings + timings top-down, validates them, rejects a shared node assigned two timings export reads the slots; rejects an Unset one ``` @@ -466,9 +467,11 @@ it is (§9). with that error. - **A shared subtree assigned two timings**, e.g. a query-time `Aggregate` and an ingestion-time `SummaryAgg` reading one `Scan`: both placements are legal, only the - sharing is not. `apply_lifecycle_timings` memoizes by (pointer, timing), so it builds one - copy per timing; a subtree assigned one timing stays one `Rc`, within a root or across - roots. + sharing is not. `apply_lifecycle_timings` memoizes by pointer and records the timing it + wrote; reaching the node again with another timing is an error, and the assignment is + rejected like any other illegal one. A subtree assigned one timing stays one `Rc`, + within a root or across roots. Nothing is copied: a plan that needs the same subtree + in both phases must hold two `Rc`s before the pass (§10). `map_children` is `rebuild_children` from `pre_asap/cse.rs`, dispatching to `NonASAPOp::map_children` / `ASAPOp::map_children`. @@ -482,7 +485,7 @@ child is a `NonASAP` subtree without an `Aggregate` — where `contains_aggregat | #468 problem | Resolution | |---|---| | 1. A `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside | one set of types | -| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds when both run at the same time; otherwise the scan is copied (above). They share under an assignment that recomputes the sketch at query time; the default assignment maintains it at ingestion time, so there the scan is copied (§2.3). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | +| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds only when both run at the same time, e.g. under an assignment that recomputes the sketch at query time; an assignment that runs them in different phases over one `Rc` is rejected (above). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | | 3. `SetOp` and similar have no post-ASAP copy, so no summary below them | `SetOp` takes `None => t`; both children are assembled | ## 5. Timing: validating the assignment @@ -510,7 +513,8 @@ The `KeepPreAsap` / `BinaryOp` / `ValueOperation` / `RelationalJoin` arms of tod (`Project` / `Filter` / `Sort` / `Limit` may pass `ExactAggregate` columns through). This rejects `Project(ASAP(SummaryAgg))`. - `BinaryOp`'s ingestion-side constraints move into this arm. -- `AmbiguousKeepPreAsap` is deleted: a subtree assigned two timings is copied (§4). +- `AmbiguousKeepPreAsap` is deleted: a shared subtree assigned two timings is rejected by + `apply_lifecycle_timings` (§4). ## 6. Export: one post-ASAP node per operator @@ -595,9 +599,9 @@ the same early flat export, on the final types. - One integration test per #468 problem: 1. `WITH metric AS (SELECT avg(CASE WHEN l_quantity BETWEEN 1 AND 50 THEN 1.0 ELSE 0.0 END) AS in_range FROM lineitem) SELECT in_range, in_range = 1.0 AS ok FROM metric` — no post-ASAP-only node besides `ASAP`; all `Project`s are one variant. - 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` when both run at the same time (once a binding rule splits measures); under the default assignment, which maintains the sketch at ingestion time, the `Scan` is copied. + 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` under an assignment that runs both at query time (once a binding rule splits measures). 3. `SELECT approx_distinct(l_partkey) FROM lineitem UNION ALL SELECT approx_distinct(l_suppkey) FROM lineitem` — each side of the `SetOp` has a `SummaryEstimate`. -- A shared `Scan` assigned two timings is copied once per timing; assigned one timing, it stays one `Rc`. +- A shared `Scan` assigned two timings is rejected; assigned one timing, it stays one `Rc`. - A node shared by two roots, assembled in two calls, is still one `Rc` after `derive_guarantees` and `apply_lifecycle_timings`. - After both passes no slot is `Unset`; export rejects a tree with one, and a tree with no assignment applied. - Exported timings of today's plans are unchanged under the default assignment. @@ -632,3 +636,11 @@ the same early flat export, on the final types. - Does ASAPQuery insert `SummaryMerge` only on the exported post-ASAP DAG, or through ASAPPlanner's post-ASAP types? The planner-side variant is unimplemented (§1.3). + +- Workload CSE can share one `Scan` between a query whose `SummaryAgg` is maintained at + ingestion time and another whose `Aggregate` runs at query time. Today `KeepPreAsap` + takes its timing from the consuming edge, so the two sides are effectively separate; + after this proposal the default assignment is rejected on that `Rc` (§4). Either + assembly un-shares a `NonASAP` subtree that `SummaryAgg` reads, or summary + materialization must treat the shared scan as one state with one lifecycle. Not + decided here. From f9408ac153788db6792f14b5619775df646dec5e Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 19:39:04 +0000 Subject: [PATCH 08/46] renaming in doc --- docs/design_docs/proposals/operator-sharing.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index cc9420c57..043965c72 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -12,7 +12,7 @@ Nothing is wrapped and nothing is duplicated. `Operator` has two levels, `NonASAP(NonASAPOp)` and `ASAP(ASAPOp)`, rather than one flat enum of every variant: frontends, `resolve` and the per-operator export (§6) work on `NonASAP` -trees only, and `NonASAPOp` gives them a precise type for that instead of a run-time check +DAGs only, and `NonASAPOp` gives them a precise type for that instead of a run-time check on each node. ``` From ccbf7b53750e9199469200192025c4bd63d06616 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 19:52:43 +0000 Subject: [PATCH 09/46] removed some verbosity --- .../design_docs/proposals/operator-sharing.md | 47 +++++++++---------- 1 file changed, 23 insertions(+), 24 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 043965c72..87d9ad340 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -273,21 +273,22 @@ Per node kind: ### 2.3 Timing: written from a lifecycle assignment -The logical layer — PlanSpace, binding, assembly — decides *what* to compute, not when. -Timing is chosen by summary materialization: for every unique summary state it picks a -lifecycle (maintain at ingestion time or recompute at query time, with window and -retention), and that choice fixes the timing of every node that feeds or reads the state. -One logical DAG can therefore have several assignments; the deployment chooses among -them with its own costs ([Output layers](../architecture/input-output-workflow.md#output-layers), #480). - -Nothing in the IR sets a timing. Binding sets no `SummaryAgg.timing`; the planner sets -none on `FinalizeExactAccumulator` (`finalize_query_candidate`, §4, still inserts the -node, its timing is assigned like any other). After assembly every `timing` slot is -`Unset`. `apply_lifecycle_timings` writes the chosen assignment into the slots, top-down -per root with one shared memo, and validates it (§5): an assignment under which -ingestion work depends on a query-time result, or a node of fixed kind gets the wrong -timing, is rejected, as is one that gives a shared node two timings (§4). `Unset` means -no assignment was applied; physical compilation and export reject it. +The logical layer (PlanSpace, binding, assembly) decides *what* to compute. Summary +materialization decides *when*: it picks a lifecycle per summary state (maintain at +ingestion time or recompute at query time, with window and retention), and that choice +fixes the timing of every node that feeds or reads the state. One logical DAG can have +several assignments; the deployment picks one by its own costs +([Output layers](../architecture/input-output-workflow.md#output-layers), #480). + +The IR never sets a timing: after assembly every `timing` slot is `Unset`. +`apply_lifecycle_timings` writes the chosen assignment into the slots, top-down per root +with one shared memo, and validates it (§5). It rejects an assignment that + +- makes ingestion work depend on a query-time result, +- gives a node of fixed kind the wrong timing, e.g. an ingestion-time `SummaryEstimate`, +- gives a shared node two timings (§4). + +Export and physical compilation reject an `Unset` slot: no assignment was applied. ``` assembly after apply_lifecycle_timings @@ -297,14 +298,12 @@ Project Unset QueryTime Scan t Unset IngestionTime ← feeds a maintained state ``` -Some candidates fix a timing when they are built today: exact compositions carry -`OperationPlacement::Read` / `Maintenance` (provenance `ValueOperationAtQueryTime` / -`ValueOperationAtIngestionTime`), maintained populations are built at ingestion time, -and #472's grouped `Rate`→`Sum` pair. Each becomes one logical candidate whose placement -is a lifecycle choice (§8 stage 5). The **default assignment** reproduces today's -timings — a `SummaryAgg` maintained at ingestion time, everything above a readout at -query time — so exported timings do not change until lifecycle selection chooses -otherwise. +Candidates that fix a timing at construction today — exact compositions +(`OperationPlacement::Read` / `Maintenance`), maintained populations, #472's grouped +`Rate`→`Sum` pair — each become one logical candidate whose placement is a lifecycle +choice (§8 stage 5). The **default assignment** reproduces today's timings (`SummaryAgg` +at ingestion time, everything above a readout at query time), so exported timings do not +change until lifecycle selection chooses otherwise. Per node kind: @@ -317,7 +316,7 @@ Per node kind: | `MaintainPopulation` / `ReadPopulation` | a stored field, always ingestion / query time | **assigned**; validated: ingestion / query time only | | unused variants | `SummaryMerge`: a stored field; `Join` / `Subtract` / `Delete`: ingestion time | unimplemented (§1.3) | -Unlike a guarantee, a timing is a property of the whole DAG and its assignment, so it is +A timing, unlike a guarantee, depends on the whole DAG and its assignment, so it is applied only after assembly. ### 2.4 Workflow of setting up `guarantee` and `timing`: today vs. after From bb2403ae844a9fb6941395a6eb762b51f8690990 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 19:57:45 +0000 Subject: [PATCH 10/46] docs: simplify operator-sharing proposals and clarify timing ownership --- .../proposals/decoupling_op_and_expr.md | 95 ++-- .../design_docs/proposals/operator-sharing.md | 423 ++++++++---------- 2 files changed, 233 insertions(+), 285 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index dc8d6f6ad..f593ea8bf 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -1,16 +1,16 @@ # Decoupling Operators From Scalar Expressions -> Status: proposed, not implemented. Companion to [Operator sharing](operator-sharing.md): -> this document splits `QueryExpr`; that one builds the shared operator language on the -> result. Code is referenced by file and function; counts are approximate, measured on -> `main` at `5a32b8b`. +> Status: proposal, not implemented. Audience: planner developers and designers. +> Companion to [Operator sharing](operator-sharing.md). Code references use `main` +> at `5a32b8b`. -**The idea.** `QueryExpr` holds two different kinds of node in one enum. This proposal -splits it into `NonASAPOp` (operators) and `ScalarExpr` (scalar expressions). +Split `QueryExpr` into `NonASAPOp` for table-producing operators and `ScalarExpr` +for expressions evaluated within an operator. This makes invalid combinations, +such as a literal used as a filter's table input, unrepresentable. ``` Today Proposed -Filter { pred: Rc, Filter { pred: Predicate(ScalarExpr), +Filter { pred: Rc, Filter { pred: Predicate(ScalarExpr), child: Rc } child: Rc } ``` @@ -26,18 +26,14 @@ Project ← operator: outputs a table child: Scan lineitem ← operator ``` -- An **operator** outputs a table. It is a node of the plan: the planner can replace it, - share it, or put a summary under it. -- A **scalar expression** has no table of its own. `Column(4)` means "column 4 of the - input of the operator I sit in"; outside that operator it means nothing. +Operators produce tables and can be replaced or shared by the planner. Scalars +compute values within the schema chosen by their operator; `Column(4)` has no meaning without +that context. -Today both are `QueryExpr` variants, told apart only by field position. The code already -separates them, but only by convention: +Today both are `QueryExpr` variants, separated only by convention: -- `QueryExpr::output_schema` returns `ScalarHasNoRowSchema` for all 13 scalar variants - (`query_expr.rs`), so `Filter { child: Literal(2) }` compiles and fails at run time. -- `pre_asap/cse.rs` never descends into a scalar, and repeats a "scalar: nothing to do" - arm in each of its three traversals; `canonicalize` likewise never rewrites one. +- `Filter { child: Literal(2) }` compiles, then fails with `ScalarHasNoRowSchema`. +- CSE and canonicalization skip scalars using repeated special-case branches. ## 2. Types @@ -53,10 +49,14 @@ pub enum ScalarExpr { Column(C), Literal(ScalarValue), Compare { .. }, BoolAnd(..), BoolOr(..), Not(..), IsNull(..), IsNotNull(..), Cast { .. }, InList { .. }, FunctionCall { .. }, Arithmetic { .. }, Case { .. }, CurrentTimestamp, } -pub struct Predicate(pub ScalarExpr); // scalars are held by value: CSE hashes them as plain data, nothing shares them +pub struct Predicate(pub ScalarExpr); pub struct ProjectItem { pub alias: Option, pub expr: ScalarExpr } ``` +Operators own their scalar fields by value. CSE hashes these fields as data; scalars +are not shared DAG nodes. Recursive scalar fields still require indirection such as +`Box` or `Vec`; the type sketches omit those details. + ```text NonASAPOp ├─ children: Rc> (Rc> after operator sharing) @@ -65,49 +65,50 @@ NonASAPOp └─ children: ScalarExpr only, never an operator ``` -- **Naming**: `NonASAPOp` is named for [Operator sharing](operator-sharing.md), where it - becomes the non-ASAP category of `Operator`. `QueryExpr` goes away. -- **Scalar fields**: `Filter.pred`, `Join.pred`, `Aggregate.having` (`Predicate`); - `Project.cols` (`ProjectItem`); `Sort` / `SQLWindowFunc` sort keys (`SortKey`); - `SQLWindowFunc.args`; `PromqlRelabel.value`. -- **Borderline variants** go by position, not by look. `PromqlScalarBridge`, - `EvalTimestamp`, `PromqlScalarFromVector` and `PromqlVectorFromScalar` sit in operator - position with a row schema (e.g. a `BinaryOp` operand) → `NonASAPOp`. `CurrentTimestamp` - (SQL `NOW()`) is produced only by scalar lowering (`df_expr_to_unresolved`); its only - operator-position use is in a unit test → `ScalarExpr`. -- **Gain**: neither a scalar in operator position nor an operator in scalar position is - expressible, and `ScalarHasNoRowSchema` is deleted. +`QueryExpr` is removed. `NonASAPOp` is named for its role in the companion proposal. +Scalar fields include predicates (`Filter`, `Join`, `HAVING`), project items, sort +keys, window-function arguments and relabel values. + +Classify ambiguous variants by their position and schema: + +- `ScalarBridge`, `EvalTimestamp`, `PromqlScalarFromVector` and + `PromqlVectorFromScalar` remain operators because they have row schemas. +- `CurrentTimestamp` (SQL `NOW()`) is a scalar. Its only operator-position use today + is a unit test. + +The split prevents scalars in operator positions and operators in scalar positions, +eliminating `ScalarHasNoRowSchema`. ## 3. Changes | Location | Change | |---|---| | frontend expression lowering (`df_expr_to_unresolved`, PromQL `walk`) | scalar positions build `ScalarExpr`, operator positions `NonASAPOp` | -| `resolve`, `column_resolution.rs` | already separate: `resolve` walks operators and calls `resolve_expr` for scalars. Each scalar resolves against one schema its operator picks (usually the child's output; the `Aggregate`'s output for `HAVING`, left + right for a `Join` predicate, the `Scan`'s own schema for `Scan` predicates). The split only changes their signatures: `resolve` takes `NonASAPOp`, `resolve_expr` takes `ScalarExpr`. Leaf schemas are still inferred from scalar column references across the whole tree | +| `resolve`, `column_resolution.rs` | `resolve` takes `NonASAPOp`; `resolve_expr` takes `ScalarExpr`. Schema selection and leaf-schema inference are unchanged. | | `canonicalize`, `pre_asap/cse.rs` | the "scalar: nothing to do" arms go; scalars are hashed as plain data | | `scalar_signature.rs`, `infer_expr_type` | take `ScalarExpr` | | `QueryExpr::output_schema` | becomes `NonASAPOp::output_schema`; the scalar arms and `ScalarHasNoRowSchema` go | +Scalars resolve against the schema chosen by their operator: usually the child's +output, the aggregate output for `HAVING`, both inputs for a join predicate, or the +scan schema for a scan predicate. Leaf schemas still use column references across +the whole tree. + ## 4. Implementation and tests -This is stage 1 of the joint plan ([Operator sharing §8](operator-sharing.md#8-stages-and-tests)): -children stay `Rc`; operator sharing widens them to `Rc` in its -stage 2. +This is [stage 1 of operator sharing](operator-sharing.md#8-stages-and-tests). +Children stay `Rc` until stage 2 widens them to `Rc`. -**No wire change.** The `fallback` payload serializes a `QueryExpr`, externally tagged. -Variant names are kept, so a tree serializes the same; `ScalarBridge` keeps the name -`PromqlScalarBridge` with `#[serde(rename)]`. +The wire format stays unchanged: preserve the externally tagged variant names, +including `PromqlScalarBridge` via `#[serde(rename)]`. -- Existing tests pass unchanged apart from construction syntax. -- Tests that place a scalar in operator position no longer compile and are rewritten or - deleted: the `CurrentTimestamp` unit test, the `ScalarHasNoRowSchema` tests, and the - `post_asap_dag.rs` tests using `QueryExpr::Literal` as a `fallback` expression. +Update construction syntax in existing tests. Rewrite or remove tests that put +scalars in operator positions: the `CurrentTimestamp`, `ScalarHasNoRowSchema`, and +literal-as-`fallback` cases. ## 5. Limits -The split relies on no scalar containing an operator. That holds today: SQL -`IN (SELECT …)` / `EXISTS` in a filter lower to a semi-join (`lower_filter`), and every -other subquery-valued expression is rejected (`frontend-sql/src/sql/expr.rs`). Supporting -a scalar subquery (`WHERE x > (SELECT avg(x) …)`) would add `ScalarExpr::Subquery(Rc<..>)`, -make the two types mutually recursive, and require CSE and the planner to look inside -scalars. +Scalars currently contain no operators: filter `IN (SELECT …)` and `EXISTS` lower +to semi-joins; other subquery-valued expressions are rejected +(`frontend-sql/src/sql/expr.rs`). Supporting scalar subqueries later would require +mutually recursive types and traversal into scalars by CSE and the planner. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 87d9ad340..a95fa28b1 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -1,19 +1,20 @@ # Sharing Operators Between Pre-ASAP IR and Post-ASAP IR -> - Status: proposed, not implemented. -> - Problem statement: [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). -> - Builds on [Decoupling operators from scalar expressions](decoupling_op_and_expr.md), which splits `QueryExpr` into `NonASAPOp` and `ScalarExpr`. -> - Code is referenced by file and function against `main` at `5a32b8b` (after #472, #478, #510). -> - Timing follows the planner layering of #480 / #509: a node's timing is written from the summary maintenance lifecycle assignment chosen for the DAG, never inferred from the IR (§2.3). +> Status: proposal, not implemented. Audience: planner developers and designers. +> Addresses [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468) and builds on +> [Decoupling operators from scalar expressions](decoupling_op_and_expr.md). +> Code references use `main` at `5a32b8b` (after #472, #478 and #510). -**The idea.** Today a post-ASAP plan is glued together from two sets of operator types. -This proposal keeps one operator language and makes summary operators extra node kinds in it: any relational operator can sit above a summary, and a summary can read any relational subtree. -Nothing is wrapped and nothing is duplicated. +Use one operator language before and after ASAP optimization. Relational operators +and summary operators can be children of each other, without `KeepPreAsap` wrappers +or duplicate relational variants. -`Operator` has two levels, `NonASAP(NonASAPOp)` and `ASAP(ASAPOp)`, rather than one flat -enum of every variant: frontends, `resolve` and the per-operator export (§6) work on `NonASAP` -DAGs only, and `NonASAPOp` gives them a precise type for that instead of a run-time check -on each node. +`Operator` has two categories: `NonASAP(NonASAPOp)` and `ASAP(ASAPOp)`. This separates +relational and summary operations while giving both the same child type, +`Rc`. Frontend inputs must contain only `NonASAP` nodes (§3). + +Guarantees are derived from the DAG. Execution timings come from an explicit +lifecycle assignment and are validated before export (§2). ``` Today Proposed @@ -37,14 +38,9 @@ ValueOperation(Project) ← a copy NonASAP(Project) ### 1.1 Unified `Operator` type -Operator attributes differ in how widely they apply. -We define the `Operator` type structure based on the breadth of its attributes. - -| Applies to | Examples | Defined as | -|---|---|---| -| every operator | children, schema, timing, guarantee | methods implemented for `Operator` | -| one category | for all `NonASAP` operators, guarantee is derived from the children, and the assigned timing is checked against the consuming edge | implementation specified to one enum branch of `Operator` | -| one operator | `Aggregate.measures`, `SummaryAgg.family` | fields of that variant | +Common operations—children, schema, guarantee and timing—are methods on `Operator`. +Each enum branch implements them for its category; variant-specific data, such as +`Aggregate.measures` and `SummaryAgg.family`, stays on the variant. ```rust pub enum Operator { @@ -58,8 +54,8 @@ impl Operator { // implemented for every operator pub fn output_schema(&self) -> Result; // schema: §2.1 pub fn guarantee(&self) -> &Slot>; // accuracy guarantee: §2.2 pub fn timing(&self) -> &Slot; // execution timing: §2.3 - pub fn with_guarantee(&self, guarantee: Option) -> Self; // Setter of accuracy guarantee - pub fn with_timing(&self, timing: ExecutionTiming) -> Self; // Setter of execution timing + pub fn with_guarantee(&self, guarantee: Option) -> Self; + pub fn with_timing(&self, timing: ExecutionTiming) -> Self; } /// `Slot` represents a value that may be unset or set. @@ -89,24 +85,22 @@ pub fn apply_lifecycle_timings( ) -> Result, ExecutionDataStateError>; ``` -- **Passes copy**: `derive_guarantees` and `apply_lifecycle_timings` return a new tree; - nodes are immutable and `with_*` build new ones. Pointers held before a pass - (`assembled_nodes`, planner memos) are not valid into its result; lifecycle selection - reads the guarantee-derived tree, export reads the timed tree (§2.4). -- **Passes recompute**: `derive_guarantees` keeps the values set at construction (§2.2) - and recomputes every other slot; `apply_lifecycle_timings` overwrites every `timing` - slot. Every rewrite runs before them, so running either again gives the same slots. -- **Equality**: the `timing` slot takes part in `PartialEq` and `Hash`, so CSE never merges - two nodes assigned different timings. The `guarantee` slot takes part in neither: it is - a function of the subtree, so equal subtrees derive equal guarantees, and - `ResultGuarantee` holds `f64` and has no `Hash`. - -Following diagram conceptually displays the structure of `Operator`: +- **Immutable passes:** both passes return a new DAG and preserve sharing with one + memo per pass across all workload roots. Lifecycle planning uses the + guarantee-derived DAG; export uses the timed DAG. Earlier pointers do not identify + nodes in those results. +- **Recomputation:** `derive_guarantees` fills guarantee slots from local evidence + and child guarantees. `apply_lifecycle_timings` overwrites timing slots from the + assignment. Run rewrites before these passes. +- **Equality:** timing participates in `PartialEq` and `Hash`; derived guarantees do + not. Equal subtrees derive equal guarantees under the same accuracy model and + evidence. `ResultGuarantee` contains `f64` and has no `Hash`. + ```text Operator ├─ NonASAP(NonASAPOp) │ ├─ children: Rc> → back to Operator: NonASAP or ASAP -│ ├─ timing / guarantee +│ ├─ timing / guarantee │ └─ scalar expressions: Predicate / ProjectItem / SortKey / ScalarBridge / ... │ └─ ScalarExpr: never contains an Operator └─ ASAP(ASAPOp) @@ -116,8 +110,8 @@ Operator ### 1.2 `NonASAPOp` -`NonASAPOp` is the non-ASAP category of `Operator`. -It comes from splitting `QueryExpr` into "operator" and "scalar expression" parts ([decoupling doc](decoupling_op_and_expr.md#2-types)). +`NonASAPOp` contains the operator variants split from `QueryExpr` in the +[decoupling proposal](decoupling_op_and_expr.md#2-types). ```rust pub enum NonASAPOp { @@ -138,12 +132,12 @@ pub enum NonASAPOp { } ``` -Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. +Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. ### 1.3 `ASAPOp` -`ASAPOp` is the ASAP category of `Operator`. -`ASAPOp` comes from today's `SummaryExpr`: its summary variants, and the summary-specific `ValueOperation` variants. +`ASAPOp` contains today's summary variants from `SummaryExpr` and the +summary-specific variants of `ValueOperation`. ```rust pub enum ASAPOp { @@ -164,9 +158,11 @@ pub enum ASAPOp { Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. -**Unused branches**: `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` are built only in tests today. They are migrated, but for safety, we have all their methods return `Unimplemented`. +`SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` +are currently built only in tests. Migrate their variants, but return `Unimplemented` +from schema, guarantee, timing and export operations. -Following table shows how some legacy types get expressed in the new framework. +Legacy types map to the new IR as follows: | Legacy types | Expressed as | |---|---| @@ -183,7 +179,7 @@ be `Rc` to allow free placement: ```rust // After the decoupling doc // After this proposal -Filter { pred: Predicate(ScalarExpr), Filter { pred: Predicate(ScalarExpr), +Filter { pred: Predicate(ScalarExpr), Filter { pred: Predicate(ScalarExpr), child: Rc } child: Rc } // NonASAP(..) or ASAP(SummaryEstimate ..) ``` @@ -194,34 +190,28 @@ cannot replace a branch — e.g. the branches of SQL `ROLLUP` or PromQL `histogr ## 2. Schema, Guarantee, and Timing -This section discusses three key per-node attributes, `schema`, `guarantee`, and `timing`, as well as how they are stored and filled in the new framework. +| Attribute | Proposed rule | +|---|---| +| `schema` | Computed by `output_schema()` using one schema type for all nodes | +| `guarantee` | Binding records local evidence; `derive_guarantees()` fills the slots bottom-up | +| `timing` | Unset during assembly; `apply_lifecycle_timings()` fills the slots from an explicit assignment | -| Field | Meaning | Today | After | -|---|---|---|---| -| `schema` | output columns and their types | pre-ASAP: computed by `QueryExpr::output_schema()`
post-ASAP: a `SummarySchema` stored on every `SummaryNode` | can be obtained by `output_schema()` | -| `guarantee` | accuracy bound | pre-ASAP: none
post-ASAP: stored on every `SummaryNode` | can be obtained by `guarantee()`
binding stores only each operator's own error
complete error bound need to be derived by `derive_guarantees()` | -| `timing` | execution time | pre-ASAP: none
post-ASAP, stored: a field on `BinaryOp` / `ValueOperation` / `SummaryMerge`
post-ASAP, not stored: `KeepPreAsap` from the consuming edge, `SummaryAgg` from the child. | can be obtained by `timing()`
`Unset` after assembly
written for every node by `apply_lifecycle_timings()` from the chosen lifecycle assignment (§2.3) | +A set guarantee of `None` means no guarantee is available. It is different from an +`Unset` slot, which means the pass has not run. ### 2.1 Schema: fused into one type -Today schemas of pre-ASAP operators and post-ASAP operators are different: -- pre-ASAP uses `Schema { columns: Vec, time_index, unique_keys, closed }` with `Column.dtype: DataType` (plain values only), -- post-ASAP stores a `SummarySchema { fields: Vec, time_index }` on every node, with `SummaryField.dtype: SummaryFamilyType` (`Plain(DataType)` or summary state). -Now since the two operators types are unified into one, we need a unified schema type as well. - -We implement the new schema type based on the original `Schema` type used in pre-ASAP operators, with two changes: - -- `Column` is renamed `Field`, and `Schema.columns` `Schema.fields`: the struct describes - a column and holds none of its data. (Arrow and DataFusion use the same names.) -- `Field.dtype` widens from `DataType` to an enum `FieldType`, which covers both plain data types and ASAP summary types. +Pre-ASAP uses `Schema` with plain `DataType` columns; post-ASAP stores a separate +`SummarySchema` that also supports summary state. Use `Schema` for both: -`SummarySchema` / `SummaryField` are then redundant and deleted. +- Rename `Column` to `Field` and `Schema.columns` to `Schema.fields`. +- Widen `Field.dtype` to `FieldType`, covering plain values and ASAP state. +- Delete `SummarySchema` and `SummaryField`. `Schema` keeps the reserved column `PROMQL_SERIES_IDENTITY` (`"$promql_series_identity"`, `pre_asap/schema.rs`) and `has_promql_series_identity()`: `maintained_population.rs` uses it to decide whether a closed PromQL schema still identifies a series (§7). -Detailed code design is shown below. ```rust pub enum FieldType { DataType(DataType), ASAPType(ASAPType) } pub enum ASAPType { // SummaryFamilyType without Plain @@ -248,10 +238,10 @@ impl Field { ### 2.2 Guarantee: always derived -A guarantee is filled in two steps: - -1. **Binding** records local accuracy guarantee: a `SummaryEstimate`'s `local_guarantee` (the sketch's error over an exact input) and an exact `SummaryAgg`'s `exact_rule`. No `guarantee` slot is set yet. To size a sketch and check its target, binding still needs the child's error, as today: it runs `derive_guarantees` on the child with a fresh memo, reads the result, and drops it. -2. **`derive_guarantees`** fills every slot bottom-up: a node without an `ASAP` descendant is exact, and every other node composes its children's guarantees by its own rule. +1. Binding records `SummaryEstimate.local_guarantee` (error over exact input) and + exact `SummaryAgg.exact_rule`, leaving guarantee slots unset. To size a sketch, it derives + the child's guarantee on a temporary copy, reads it, then drops the copy. +2. `derive_guarantees` fills slots bottom-up using the rules below. ``` Project p99 ±1% ← the child's @@ -260,35 +250,27 @@ Project p99 ±1% ← the child's Scan t exact ← no ASAP descendant ``` -Per node kind: +| Node | Guarantee rule | +|---|---| +| `SummaryEstimate` | Compose `local_guarantee` with the guarantee of the input below `SummaryAgg`; sketch state itself has no guarantee. The local guarantee is `None` if the family has no error model. | +| `SummaryAgg` | ExactAggregate: compose with the child under `exact_rule`; `ExactKind::Count` stays exact regardless of the child. Sketch state: `Set(None)`. | +| `NonASAPOp` | Compose child guarantees; exact when there is no ASAP descendant. | +| `FinalizeExactAccumulator` | Copy the child's guarantee. | +| `MaintainPopulation` / `ReadPopulation` | Exact. | +| Unused variants | Return `Unimplemented` (§1.3). | -| Node | Today | After | -|---|---|---| -| `SummaryEstimate` | stored at binding: the sketch's own error composed with the child's (`compose_guarantee`) | **derived**: `local_guarantee` composed with the guarantee of the state's input, i.e. the child of the `SummaryAgg` below (the `SummaryAgg` itself has none). `local_guarantee` is set at binding: the sketch's error over an exact input, `None` when the model has no error model for the family | -| `SummaryAgg` | stored: ExactAggregate family composed with the child's; sketch families `None` | **derived**: ExactAggregate family: exact, composed with the child's under `exact_rule`, except `ExactKind::Count`, exact whatever the child (as today); sketch families `Set(None)`, state has no guarantee | -| `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: exact
post-ASAP `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: composed at construction | **derived**: composed from the children; exact if no `ASAP` descendant | -| `FinalizeExactAccumulator` | copies the child's | **derived**: the child's | -| `MaintainPopulation` / `ReadPopulation` | stored: exact | **derived**: exact | -| unused variants | `None`: state has no guarantee of its own | unimplemented (§1.3) | - ### 2.3 Timing: written from a lifecycle assignment -The logical layer (PlanSpace, binding, assembly) decides *what* to compute. Summary -materialization decides *when*: it picks a lifecycle per summary state (maintain at -ingestion time or recompute at query time, with window and retention), and that choice -fixes the timing of every node that feeds or reads the state. One logical DAG can have -several assignments; the deployment picks one by its own costs -([Output layers](../architecture/input-output-workflow.md#output-layers), #480). +Logical planning decides what to compute. Summary materialization chooses a lifecycle +for each summary state, including execution timing and retention. One logical DAG can +have several assignments; the planner selects among them using deployment-provided +costs, consistent with the [planning-stages proposal](https://github.com/ProjectASAP/ASAPPlanner/pull/509). -The IR never sets a timing: after assembly every `timing` slot is `Unset`. -`apply_lifecycle_timings` writes the chosen assignment into the slots, top-down per root -with one shared memo, and validates it (§5). It rejects an assignment that - -- makes ingestion work depend on a query-time result, -- gives a node of fixed kind the wrong timing, e.g. an ingestion-time `SummaryEstimate`, -- gives a shared node two timings (§4). - -Export and physical compilation reject an `Unset` slot: no assignment was applied. +Assembly leaves every timing slot `Unset`. `apply_lifecycle_timings` writes the +assignment with one memo across all roots and validates node and edge constraints +(§5). It rejects ingestion work that depends on query-time results, fixed-kind timing +violations, and conflicting timings on a shared node. Export and physical compilation +reject unset slots. ``` assembly after apply_lifecycle_timings @@ -298,54 +280,39 @@ Project Unset QueryTime Scan t Unset IngestionTime ← feeds a maintained state ``` -Candidates that fix a timing at construction today — exact compositions -(`OperationPlacement::Read` / `Maintenance`), maintained populations, #472's grouped -`Rate`→`Sum` pair — each become one logical candidate whose placement is a lifecycle -choice (§8 stage 5). The **default assignment** reproduces today's timings (`SummaryAgg` -at ingestion time, everything above a readout at query time), so exported timings do not -change until lifecycle selection chooses otherwise. +Exact compositions (`OperationPlacement::Read` / `Maintenance`), maintained +populations and #472's grouped `Rate`→`Sum` pair currently fix timing during +construction. Each becomes a logical candidate whose placement is a lifecycle choice +(§8, stage 5). -Per node kind: +The default assignment aims to preserve today's timings. It cannot do so when one +shared node requires both phases; resolving that conflict remains open (§10). -| Node | Today | After | -|---|---|---| -| `NonASAPOp` | pre-ASAP `QueryExpr`: none
post-ASAP `KeepPreAsap`: from the consuming edge
post-ASAP `ValueOperation` / `BinaryOp` copies: a stored field | **assigned**; validated against its consuming edges (§5) | -| `SummaryAgg` | from the child; ingestion time under `KeepPreAsap` | **assigned** by its state's lifecycle | -| `FinalizeExactAccumulator` | a stored field, set by the planner | **assigned**; the position allows either time | -| `SummaryEstimate` | query time, fixed by the kind | **assigned**; validated: query time only | -| `MaintainPopulation` / `ReadPopulation` | a stored field, always ingestion / query time | **assigned**; validated: ingestion / query time only | -| unused variants | `SummaryMerge`: a stored field; `Join` / `Subtract` / `Delete`: ingestion time | unimplemented (§1.3) | +| Node | Timing constraint | +|---|---| +| `NonASAPOp` | Assigned timing must satisfy every consuming edge. | +| `SummaryAgg` | Assigned by the state's lifecycle. | +| `FinalizeExactAccumulator` | Either phase, subject to its inputs and consumers. | +| `SummaryEstimate` | Query time only. | +| `MaintainPopulation` / `ReadPopulation` | Ingestion time / query time only. | +| Unused variants | Return `Unimplemented` (§1.3). | -A timing, unlike a guarantee, depends on the whole DAG and its assignment, so it is -applied only after assembly. +Apply timing after assembly because it depends on the whole DAG and its lifecycle. -### 2.4 Workflow of setting up `guarantee` and `timing`: today vs. after +### 2.4 Workflow -Today: +Today, construction and assembly compose guarantees, and export derives remaining +timings into a side table. The proposed sequence is: -``` -search / binding each SummaryNode's guarantee is composed when the node is built; - BinaryOp / ValueOperation store their timing -selection reads each candidate's stored guarantee against its target -assembly assemble_residual builds kept nodes and composes their guarantee; - relink_summary copies the old guarantee onto a relinked SummaryAgg -lifecycle reads the root's guarantee -export validate_execution_data_states, per root, derives the remaining - timings into a side table and writes them onto the edges -``` - -After: - -``` -search / binding sets only what cannot be derived: local_guarantee, exact_rule; - no timing; checks accuracy on a derived copy, keeps the result on - the candidate record (§4), drops the copy -assembly builds each root; kept NonASAP nodes stay as they are -derive_guarantees per root, one shared memo: fills guarantees bottom-up -lifecycle reads the derived guarantee; chooses a lifecycle per summary state -apply_lifecycle_ per root, one shared memo: writes the assignment's timings - timings top-down, validates them, rejects a shared node assigned two timings -export reads the slots; rejects an Unset one +```text +search / binding record local accuracy evidence; leave timing unset + check accuracy on a temporary derived copy (§4) +assembly build the selected roots, preserving shared nodes +derive_guarantees fill guarantees bottom-up, with one memo across roots +lifecycle choose per-state lifecycles using the derived DAG +apply_lifecycle_timings + write and validate timing, with one memo across roots +export read filled slots; reject any Unset slot ``` --- @@ -403,23 +370,17 @@ pub enum Replacement { } ``` -`runtime_support_evidence` (`replacement.rs`) asks the cost model whether the runtime -supports a candidate, per variant: `Summary` → `summary_support_evidence`, `Rewrite` → -`Some(true)`. With one `Subtree` variant it dispatches on the root: an `ASAP` root → -`summary_support_evidence`, a `NonASAP` root → `Some(true)`. `ASAP` nodes below a -`NonASAP` root were each asked when they were candidates themselves. +**Runtime support.** `runtime_support_evidence` dispatches on the replacement root: +`ASAP` uses `summary_support_evidence`; `NonASAP` returns `Some(true)`. Nested ASAP +nodes were checked when their own candidates were built. -**Candidates stay bottom-up, as today**: a candidate is built on a concrete child plan -(`realize_child_with`, or each child candidate in `prepare_compositions`), so a chosen -plan is complete. Where `realize_child_with` falls back to `keep_pre_asap` today, it -returns the child's original subtree, and assembly keeps it as is. +**Candidate construction.** Candidates remain bottom-up and contain concrete child +plans (`realize_child_with`, `prepare_compositions`). The old `keep_pre_asap` fallback +returns the original child subtree directly. -**Accuracy check during search**: binding sets no `guarantee` slot (§2.2), so the -candidate filter in `search_workload_with_targets` and `prepare_compositions` run -`derive_guarantees` on the candidate alone, with a fresh memo, then check its accuracy target. The derived -copy is dropped; its root guarantee is stored on the candidate's `ReplacementSubDAG`, where -selection reads it (today's stored `guarantee` moved from the node to the candidate record). -PlanSpace keeps the original candidate, whose nodes are shared with other queries. +**Accuracy checks.** Search derives each candidate's guarantee with a fresh memo, +checks its target, and stores the root guarantee on `ReplacementSubDAG` for selection. +It drops the derived copy and keeps the original shared candidate in `PlanSpace`. **Assembly** — one rule replaces `assemble_residual`: @@ -440,37 +401,22 @@ fn assemble(&self, t: &Rc) -> Rc { } ``` -Then `GlobalSelection::assemble_selected_dag` runs `derive_guarantees` (§2.2) on each -root it assembles. The `DerivationMemo` lives on `GlobalSelection` next to -`assembled_nodes`, so roots assembled one call at a time still share nodes. -`assemble_selected_dag_with_summary_maintenance_lifecycles` plans lifecycles on that -result, as today: lifecycle planning holds `Rc`s into the plan and reads the root's -guarantee, so it must see the derived tree. The assignment it chooses is then written -with `apply_lifecycle_timings` (§2.3, §5). - -`assemble_selected_query` stays the boundary for a query result (#472): it runs -`assemble_selected_dag`, then `finalize_query_candidate`, which puts a query-time -`FinalizeExactAccumulator` between maintained exact state and the query-time consumer -(§2.3). The other callers of `finalize_query_candidate` — binding and the two sides of -`BinaryOp` and `Join` — keep calling it, on `Rc`. Today -`assemble_selected_dag_with_summary_maintenance_lifecycles` (`summary_maintenance_lifecycle.rs`) -calls `assemble_selected_dag` directly and skips that step; this proposal leaves that as -it is (§9). - -- **Illegal placement** (e.g. a query-time `SummaryEstimate` under a `SummaryAgg`): - candidates carry no timing, so nothing is checked when they are built. The check moves - to `apply_lifecycle_timings`, which rejects an assignment that places a node illegally. - Summary materialization offers only assignments it has validated (today - `relink_summary` runs the same check through `validate_execution_data_states_at`), so an - error from `apply_lifecycle_timings` on a chosen assignment is a bug, and planning fails - with that error. -- **A shared subtree assigned two timings**, e.g. a query-time `Aggregate` and an - ingestion-time `SummaryAgg` reading one `Scan`: both placements are legal, only the - sharing is not. `apply_lifecycle_timings` memoizes by pointer and records the timing it - wrote; reaching the node again with another timing is an error, and the assignment is - rejected like any other illegal one. A subtree assigned one timing stays one `Rc`, - within a root or across roots. Nothing is copied: a plan that needs the same subtree - in both phases must hold two `Rc`s before the pass (§10). +`GlobalSelection::assemble_selected_dag` derives guarantees after assembly, sharing +one `DerivationMemo` across roots. Lifecycle planning reads that derived DAG, then +`apply_lifecycle_timings` applies the chosen assignment. + +`assemble_selected_query` remains the query-result boundary (#472): it calls +`assemble_selected_dag`, then `finalize_query_candidate` to insert a +`FinalizeExactAccumulator` where needed. Binding and the `BinaryOp` / `Join` callers +keep this finalization step. The lifecycle assembly path still bypasses it (§9). + +Timing validation moves from candidate construction to lifecycle assignment: + +- Materialization must offer only valid assignments. Failure when applying a selected + assignment is a planning error. +- A shared node must have one timing across all roots. The pass rejects a conflict; + it does not split the node by phase. Plans needing both phases must already contain + distinct nodes (§10). `map_children` is `rebuild_children` from `pre_asap/cse.rs`, dispatching to `NonASAPOp::map_children` / `ASAPOp::map_children`. @@ -484,15 +430,14 @@ child is a `NonASAP` subtree without an `Aggregate` — where `contains_aggregat | #468 problem | Resolution | |---|---| | 1. A `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside | one set of types | -| 2. Nothing outside `KeepPreAsap` can reference the `Scan` inside, so an exact aggregate and a sketch cannot share a scan | `Aggregate` and `SummaryAgg` can point to the same `Scan`. This holds only when both run at the same time, e.g. under an assignment that recomputes the sketch at query time; an assignment that runs them in different phases over one `Rc` is rejected (above). Splitting a multi-measure `Aggregate` into exact + sketch is a binding rule, out of scope (§9) | +| 2. Exact and summary operators cannot share a scan hidden in `KeepPreAsap` | Both can reference one `Scan` when their timings agree. Different phases require distinct nodes. Splitting a multi-measure aggregate remains out of scope (§9). | | 3. `SetOp` and similar have no post-ASAP copy, so no summary below them | `SetOp` takes `None => t`; both children are assembled | ## 5. Timing: validating the assignment -`validate_execution_data_states` becomes the validation half of `apply_lifecycle_timings`. -The `ExecutionDataStateAssignment` side table is deleted; instead of deriving a state per -node, the pass checks the timing the assignment wrote into each slot against the node's -kind and its consuming edges. Nothing is derived from the consuming edge any more: +`apply_lifecycle_timings` replaces timing derivation with validation. It checks the +assigned timing against node kinds and consuming edges; the old +`ExecutionDataStateAssignment` side table is removed. | Node | Produced state | |---|---| @@ -517,8 +462,8 @@ The `KeepPreAsap` / `BinaryOp` / `ValueOperation` / `RelationalJoin` arms of tod ## 6. Export: one post-ASAP node per operator -The four original-operator payloads (`fallback{expression: QueryExpr}`, `binary`, -`value`, `relational_join`) become one: +The relational cases of `fallback{expression: QueryExpr}`, `binary`, `value` and +`relational_join` become one payload: ```rust PostAsapOperatorPayload::Relational { @@ -528,32 +473,26 @@ PostAsapOperatorPayload::Relational { } ``` -Export emits **one post-ASAP node per `NonASAP` operator**, exactly as it already does -per `ASAP` operator: children become edges, leaves are `Scan` nodes. No subtree is -embedded in a node, so there is no `DagInput` placeholder and no fragment. A `NonASAP` -node shared by two consumers (one `Rc`, kept by CSE, §4) is exported once, with two -outgoing edges. Physical compilation then corresponds node by node: each post-ASAP node -lowers to one physical operator, or to a few helper operators numbered from it. `ASAP` -nodes map one-to-one onto the existing summary payloads; -`FinalizeExactAccumulator` / `MaintainPopulation` / `ReadPopulation` stay -`value{operation}`. The `fallback` whole-expression lowering and the `binary` / -`value::Project` / `relational_join` special cases go; every `Relational` node is lowered -by one per-operator lowering that reads its inputs from its edges. - -- **Wire 5 → 6**: three fewer payloads; `fallback` becomes `relational` nodes, one per - operator; `output_schema` / `intermediate_schema` become `Schema`. One cutover (§8 - stage 4), together with the downstream readers. -- **Timing and guarantee** are read from the node slots: timing as assigned, guarantee as - derived. An `Unset` slot is rejected; for timing it means no assignment was applied. - An edge's `data_state` is its producer's assigned timing plus the primitive of its kind - (§5). `compile_post_asap_dag` splits the precompute and query DAGs by that timing and - no longer re-runs data-state validation. -- **`SummaryMerge`** stays a wire payload, although its planner-side variant is - unimplemented (§1.3, §10). -- **Phases** are per node, from the assigned timing. A phase switch between two - `NonASAP` nodes must satisfy §5 (ingestion work cannot read a query-time result); - summary materialization places materialization points only at `ASAP` state, so an - assignment that would need one elsewhere is rejected when applied. +Export emits one node per operator, with children represented by edges. A shared +operator is exported once. Each relational node stores its kind, scalar expressions +and parameters in `NonASAPOpKind`; it embeds no child subtree or `DagInput` placeholder. + +Physical compilation lowers each node to one operator or a few helper operators. +Relational lowering reads inputs from edges, replacing the old whole-expression +`fallback` and the `binary`, `value::Project` and `relational_join` special cases. +ASAP nodes retain their summary payloads; `FinalizeExactAccumulator`, +`MaintainPopulation` and `ReadPopulation` retain `value{operation}`. + +- **Wire version 6:** replace relational payloads with per-operator `relational` + nodes and use `Schema` for output and intermediate schemas. Update downstream + readers together in stage 4 (§8). +- **Attributes:** export reads derived guarantees and assigned timings, rejecting + unset slots. Edge data state combines the producer's timing with its primitive (§5). +- **Compilation:** `compile_post_asap_dag` splits precompute and query DAGs by assigned + timing without repeating validation. +- **Phase boundaries:** assignments must satisfy §5. This proposal materializes only + ASAP state; assignments requiring materialization elsewhere are rejected. +- **`SummaryMerge`:** keep its wire payload; planner-side support remains open (§10). ## 7. Other consumers @@ -581,18 +520,31 @@ by one per-operator lowering that reads its inputs from its edges. `main` builds and passes all tests after every stage. -| Stage | Content | Touches | +| Stage | Change | Main consumers | |---|---|---| -| 0 Preparation | `Rc` for `Concat.children`; `rebuild_children` → `map_children`; `Column::plain` | `asap-types` | -| 1 Split | [decoupling doc](decoupling_op_and_expr.md): `NonASAPOp` + `ScalarExpr`; children stay `Rc` | scalar code ([decoupling doc §3](decoupling_op_and_expr.md#3-changes)) | -| 2 Two levels | §1.1, §1.4: `Operator`, an empty `ASAPOp`, `contains_asap()`, `expect_non_asap()`; child slots become `Rc>`; every variant gets `timing` / `guarantee` slots, and nodes are built through constructors that leave both `Unset` (`timing` stays `Unset` until an assignment is applied) | every crate incl. `asap-planner`; the same mechanical change everywhere | -| 3 One schema | §2.1: `Column` → `Field` and `Schema.columns` → `fields` (serde keeps the name `columns` until stage 4); `FieldType`, `ASAPType`, `PlainField`, `Schema` everywhere except the `post_asap_dag.rs` wire types, which keep `SummarySchema` until stage 4. **No wire change** | `asap-types` + schema construction in every crate | -| 4 New types | fill `ASAPOp`; `ASAP` arms of `output_schema`; `derive_guarantees` and `apply_lifecycle_timings` (§2.2, §2.3, §5); the entry check (§3); `flatten(&SummaryNode) -> Rc` so export runs on the new types, copying each node's guarantee into its slot and applying today's timings as the initial assignment, so the export is unchanged; wire types become `Schema`, and `Schema.fields` serializes as `fields`. Wire → 6. **The only wire-breaking stage**; merged together with ASAPQuery-backend and ASAPCollector | `asap-types`, `devtools`, viewer | -| 5 Planner | §4: candidates and assembly on `Rc`; §7 moves to the new types; delete `flatten`. Timing moves to the lifecycle: binding and candidates set none; the candidates that fix a timing today (§2.3) become one logical candidate each, their placement a lifecycle choice; paths without lifecycle selection apply the default assignment that reproduces today's timings | `asap-aware-mapping`, `asap-planner` | -| 6 Cleanup | delete `SummaryExpr`, `SummaryNode`, extra `ValueOperation` variants, `ExactOperation`, `post_asap/cse.rs`, and today's timing fallbacks (`produced_data_state` defaults, `validate_execution_data_states_at`); `PlanOutput::dags()` and `ParsedWorkload` lose their `SummaryNode` / `QueryExpr` types; update `post-asap-ir.md` (execution phase comes from the assignment), `physical-plan-integration.md`, `updated_interface_with_pluggable_optimization.md`, developer and viewer docs | `asap-planner`, docs | - -Wrapping pre-ASAP operators in `ValueOperation` first is not planned: stage 4 gives -the same early flat export, on the final types. +| 0. Prepare | Give `Concat.children` `Rc` identity; rename `rebuild_children` to `map_children`; add `Column::plain`. | `asap-types` | +| 1. Split | Separate `NonASAPOp` and `ScalarExpr`; keep `Rc` children. | [Scalar consumers](decoupling_op_and_expr.md#3-changes) | +| 2. Unify operators | Add `Operator`, an empty `ASAPOp`, entry helpers, and unset attribute slots. Widen children to `Rc`. | All crates | +| 3. Unify schemas | Introduce `Field`, `FieldType`, `ASAPType` and the shared `Schema`. Keep wire types and serialized names unchanged. | Schema consumers | +| 4. Switch export | Fill `ASAPOp`, add attribute passes and entry validation, and export the new IR with wire version 6. | Types, devtools, viewer, backend, collector | +| 5. Migrate planning | Build candidates and assemble directly on `Rc`; move timing choices to lifecycle planning; migrate §7 consumers. | Mapping and planner crates | +| 6. Remove legacy code | Delete old IR types, duplicate operators, CSE and timing fallbacks; update architecture and developer docs. | All remaining consumers | + +Stage 4 is the only wire-breaking stage and must land with ASAPQuery-backend and +ASAPCollector updates. A temporary `flatten(&SummaryNode) -> Rc` adapter +copies existing guarantees and applies existing timings where valid. Wire schemas +switch from `SummarySchema` to `Schema`, and `Schema.fields` serializes as `fields` +instead of `columns`. + +Stage 5 deletes `flatten`. Binding leaves timing unset; lifecycle planning handles +the placement choices listed in §2.3. Paths without lifecycle selection use the +default assignment, subject to the shared-node conflict in §10. + +Stage 6 removes `SummaryExpr`, `SummaryNode`, duplicate `ValueOperation` variants, +`ExactOperation`, `post_asap/cse.rs`, `produced_data_state` defaults and +`validate_execution_data_states_at`. `PlanOutput::dags()` and `ParsedWorkload` finish +moving to `Operator`. Update `post-asap-ir.md`, `physical-plan-integration.md`, +`updated_interface_with_pluggable_optimization.md`, and developer/viewer docs. **Tests**: @@ -603,7 +555,7 @@ the same early flat export, on the final types. - A shared `Scan` assigned two timings is rejected; assigned one timing, it stays one `Rc`. - A node shared by two roots, assembled in two calls, is still one `Rc` after `derive_guarantees` and `apply_lifecycle_timings`. - After both passes no slot is `Unset`; export rejects a tree with one, and a tree with no assignment applied. -- Exported timings of today's plans are unchanged under the default assignment. +- The default assignment preserves existing timings where valid; mixed-phase sharing is covered by the open question in §10. - `apply_lifecycle_timings` rejects an assignment that puts ingestion work over a query-time result, and one that gives a `SummaryEstimate` ingestion time. - A kept `NonASAP` node (e.g. a `SetOp`) reports the guarantee composed from its assembled children. - Each unused branch (§1.3) returns `Unimplemented` from `output_schema`, `derive_guarantees`, `apply_lifecycle_timings` and export. @@ -629,17 +581,12 @@ the same early flat export, on the final types. ## 10. Open questions -- `LifecycleAssignment` (§1.1): #482 adds `SummaryMaintenanceLifecyclePlan::execution_timed_dag()`, - which expands per-state lifecycle choices into per-node timings. This proposal should - take that type as the assignment rather than define its own. - -- Does ASAPQuery insert `SummaryMerge` only on the exported post-ASAP DAG, or through ASAPPlanner's - post-ASAP types? The planner-side variant is unimplemented (§1.3). - -- Workload CSE can share one `Scan` between a query whose `SummaryAgg` is maintained at - ingestion time and another whose `Aggregate` runs at query time. Today `KeepPreAsap` - takes its timing from the consuming edge, so the two sides are effectively separate; - after this proposal the default assignment is rejected on that `Rc` (§4). Either - assembly un-shares a `NonASAP` subtree that `SummaryAgg` reads, or summary - materialization must treat the shared scan as one state with one lifecycle. Not - decided here. +- **Assignment type:** reuse the type returned by #482's + `SummaryMaintenanceLifecyclePlan::execution_timed_dag()` to expand state lifecycles + into per-node timings, rather than define a competing type. +- **Summary merge:** does ASAPQuery insert `SummaryMerge` only in the exported DAG, + or through planner-side types? The latter remains unimplemented. +- **Mixed-phase sharing:** one shared `Scan` may feed an ingestion-time `SummaryAgg` + and a query-time `Aggregate`. The default assignment then conflicts. Decide whether + assembly creates separate subtrees or materialization gives the shared state one + lifecycle. The timing pass itself must not silently split it. From f9843d6b9d12ece8d0241e034c29a326f9ac1bb7 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 19:57:50 +0000 Subject: [PATCH 11/46] some wording fixed in 2.3 --- docs/design_docs/proposals/operator-sharing.md | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 87d9ad340..f4e99449d 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -277,10 +277,9 @@ The logical layer (PlanSpace, binding, assembly) decides *what* to compute. Summ materialization decides *when*: it picks a lifecycle per summary state (maintain at ingestion time or recompute at query time, with window and retention), and that choice fixes the timing of every node that feeds or reads the state. One logical DAG can have -several assignments; the deployment picks one by its own costs -([Output layers](../architecture/input-output-workflow.md#output-layers), #480). +several assignments; the deployment picks one by its own costs. -The IR never sets a timing: after assembly every `timing` slot is `Unset`. +Operators never set a timing by themselves: after assembly every `timing` slot is `Unset`. `apply_lifecycle_timings` writes the chosen assignment into the slots, top-down per root with one shared memo, and validates it (§5). It rejects an assignment that From ed62d1ed6c6a05e68f90790688425960cd07beb6 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:06:01 +0000 Subject: [PATCH 12/46] docs: explain why attribute passes must preserve DAG sharing --- .../design_docs/proposals/operator-sharing.md | 46 +++++++++++++++---- 1 file changed, 38 insertions(+), 8 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index a95fa28b1..12fca42ec 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -63,8 +63,39 @@ impl Operator { // implemented for every operator /// assignment (§2.3). Both are `Unset` on a freshly built node. pub enum Slot { Unset, Set(T) } -/// Memo of one pass over a workload, keyed by node pointer. -/// One is shared by every root of a workload, so a node shared by two roots stays one `Rc`. +``` + +**Preserve sharing while filling attributes.** Guarantee derivation and timing +assignment return new nodes because the IR is immutable. If two queries share a +`Scan`, processing their roots independently could create two replacement scans and +lose that sharing. The design therefore requires each input node to map to one +output node within a pass, across every root in the workload: + +```text +Before the pass After the pass +query A ─┐ query A′ ─┐ + ├─ shared Scan ├─ shared Scan′ +query B ─┘ query B′ ─┘ +``` + +A temporary lookup table can enforce this: record `input node → output node` on the +first visit and reuse the output on later visits. Key it by node identity, not +structural equality: the pass preserves existing sharing; it does not introduce +new sharing between separate computations. For timing, a repeated visit must also +agree with the timing already assigned, or the assignment is rejected (§4). + +Use a fresh table for each pass invocation, shared across its workload roots. +Reusing it after a rewrite or for another assignment could return stale nodes. +Today's `GlobalSelection.assembled_nodes` already uses this approach during assembly +(`replacement.rs`); this proposal extends it to the attribute passes. + +`DerivationMemo` below names that temporary table. Its name, storage and exposure in +these illustrative signatures are implementation details, not a new plan attribute +or a required public API. The design requirement is to preserve sharing and detect +timing conflicts. + +```rust +/// Reuses each input node's output across all roots in one attribute pass. pub struct DerivationMemo { .. } /// Build the accuracy guarantee of one DAG root by derivation @@ -85,10 +116,8 @@ pub fn apply_lifecycle_timings( ) -> Result, ExecutionDataStateError>; ``` -- **Immutable passes:** both passes return a new DAG and preserve sharing with one - memo per pass across all workload roots. Lifecycle planning uses the - guarantee-derived DAG; export uses the timed DAG. Earlier pointers do not identify - nodes in those results. +- **Pass results:** lifecycle planning uses the guarantee-derived DAG; export uses + the timed DAG. Earlier pointers do not identify nodes in those results. - **Recomputation:** `derive_guarantees` fills guarantee slots from local evidence and child guarantees. `apply_lifecycle_timings` overwrites timing slots from the assignment. Run rewrites before these passes. @@ -401,8 +430,9 @@ fn assemble(&self, t: &Rc) -> Rc { } ``` -`GlobalSelection::assemble_selected_dag` derives guarantees after assembly, sharing -one `DerivationMemo` across roots. Lifecycle planning reads that derived DAG, then +`GlobalSelection::assemble_selected_dag` derives guarantees after assembly, +preserving shared nodes across roots as described in §1.1. Lifecycle planning reads +that derived DAG, then `apply_lifecycle_timings` applies the chosen assignment. `assemble_selected_query` remains the query-result boundary (#472): it calls From b71e4ff6a14d8b8f10a060dec8c1be4f1c01422e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:08:01 +0000 Subject: [PATCH 13/46] docs: explain operator design choices and each common method --- .../design_docs/proposals/operator-sharing.md | 100 +++++++++++++++--- 1 file changed, 86 insertions(+), 14 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 12fca42ec..606826511 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -38,9 +38,57 @@ ValueOperation(Project) ← a copy NonASAP(Project) ### 1.1 Unified `Operator` type -Common operations—children, schema, guarantee and timing—are methods on `Operator`. -Each enum branch implements them for its category; variant-specific data, such as -`Aggregate.measures` and `SummaryAgg.family`, stays on the variant. +**Why one operator type?** Today a relational subtree inside `KeepPreAsap` is +opaque to the surrounding summary plan. To put a `Project` above a summary, the +planner needs a second project representation; to share the enclosed `Scan` with +another computation, it must cross that wrapper. Both problems come from giving +relational and summary plans different node types. + +Use `Operator` as the common node type, with `Rc` children. Then +`Project → SummaryEstimate → SummaryAgg → Scan` is one traversable DAG. Assembly, +sharing and export can follow every edge using the same interface. Each relational +operator keeps one definition wherever it appears in the plan. + +**Why keep two categories?** Relational operators describe ordinary query semantics; +ASAP operators introduce summary state and its accuracy and lifecycle rules. Keeping +`NonASAPOp` and `ASAPOp` separate lets schema and validation code handle those rules +by category while common traversals work on `Operator`. A single flat enum could +also express the DAG; the two-level enum keeps that semantic distinction explicit, +at the cost of another match when dispatching to a variant. It does not prove that +a whole subtree is non-ASAP: children can contain either category, so optimizer +entry still validates frontend DAGs (§3). + +**Why a common interface?** Traversals need children, schema, guarantee and timing +for every node. These are methods on `Operator`; each category implements its own +rules. Data needed by only one operator, such as `Aggregate.measures` or +`SummaryAgg.family`, stays on that variant. `map_children` lets assembly replace +inputs while retaining the parent operator, including operators with no special +summary-planning logic. + +The methods below expose a node's inputs and attributes. They are illustrative +signatures, not a finalized API. `C` describes how column references are represented +(e.g. `ColumnId`); `C: ColState` requires that representation to support the column +operations used by the IR. + +| Method | Meaning and purpose | Example | +|---|---|---| +| `children()` | Read the node's immediate input operators. Generic traversals use this to walk the DAG without matching every operator kind. It does not return scalar expressions or all descendants. | A `Filter` has one input; a `Join` has two; a `Scan` has none. | +| `map_children(f)` | Return a copy of this operator with `f` applied to each immediate input, keeping its other fields. Assembly uses it to connect selected child plans under an existing parent. It does not recursively rewrite the DAG by itself. | Keep a `Project` and its expressions, but replace its input aggregate with a selected summary-estimation subtree. | +| `output_schema()` | Compute the output fields and their types, or report a schema error. Parents and export need this to interpret the node's result. | `SummaryAgg` outputs grouping fields and summary state; `SummaryEstimate` outputs grouping fields and estimated values. | +| `guarantee()` | Read the stored accuracy result without computing it. Selection and validation use the result after guarantee derivation. | `Unset`: not derived; `Set(Some(g))`: known guarantee, including exactness; `Set(None)`: no guarantee established. | +| `timing()` | Read the stored execution phase without choosing it. Export uses this after a lifecycle assignment has been applied. | `Unset`, `Set(IngestionTime)` or `Set(QueryTime)`. | +| `with_guarantee(g)` | Return a copy of this node with its guarantee slot set to `Set(g)`. The guarantee pass uses it to record its calculation. | `with_guarantee(None)` records that derivation found no guarantee; it does not leave the slot unset. | +| `with_timing(t)` | Return a copy of this node with its timing slot set to `Set(t)`. The timing pass uses it to record the lifecycle's choice. | `with_timing(QueryTime)` marks the copied node for query-time execution. | + +`&self` means the method reads the current node. `Self` means it returns a node of +this same type; `with_*` and `map_children` do not mutate the original or rewrite its +descendants. The passes are responsible for traversal and validation: setting a +slot alone does not prove that its value is valid. After replacing children, rerun +the attribute passes before export because the copied slots may be stale. + +`Rc` is a shared reference to a node. Thus `children()` returns borrowed references +to existing inputs, while the callback passed to `map_children` returns the shared +reference to use for each replacement input. ```rust pub enum Operator { @@ -57,14 +105,30 @@ impl Operator { // implemented for every operator pub fn with_guarantee(&self, guarantee: Option) -> Self; pub fn with_timing(&self, timing: ExecutionTiming) -> Self; } +``` -/// `Slot` represents a value that may be unset or set. -/// `guarantee` is filled by derivation (§2.2), `timing` by applying a lifecycle -/// assignment (§2.3). Both are `Unset` on a freshly built node. -pub enum Slot { Unset, Set(T) } +**Why defer attributes?** A node's complete accuracy guarantee depends on its +assembled inputs; its execution timing depends on the chosen lifecycle. Construction +therefore leaves both unresolved. Explicit passes fill them once that context is +available (§2), so constructors do not guess timings or duplicate error-composition +logic. Export reads the completed attributes and rejects an unfinished plan. + +`Slot` distinguishes “not computed yet” from a computed result. In particular, +`Unset` means the guarantee pass has not run, while `Set(None)` means it ran but +could not establish a guarantee. Collapsing those states would hide missing passes. +Schema is computed from the operator and its children rather than stored as another +copy that a rewrite could leave stale. +```rust +pub enum Slot { Unset, Set(T) } ``` +**Why return new nodes?** Candidate plans can share nodes. Mutating a node's timing +for one candidate could change another candidate that needs a different lifecycle. +Immutable nodes keep those plans independent: `with_*` and the attribute passes +return new nodes. This requires allocation and explicit preservation of sharing +within each resulting plan. + **Preserve sharing while filling attributes.** Guarantee derivation and timing assignment return new nodes because the IR is immutable. If two queries share a `Scan`, processing their roots independently could create two replacement scans and @@ -121,9 +185,10 @@ pub fn apply_lifecycle_timings( - **Recomputation:** `derive_guarantees` fills guarantee slots from local evidence and child guarantees. `apply_lifecycle_timings` overwrites timing slots from the assignment. Run rewrites before these passes. -- **Equality:** timing participates in `PartialEq` and `Hash`; derived guarantees do - not. Equal subtrees derive equal guarantees under the same accuracy model and - evidence. `ResultGuarantee` contains `f64` and has no `Hash`. +- **Equality:** timing participates in `PartialEq` and `Hash` because computations + assigned to different phases cannot be merged into one execution. Derived + guarantees do not: they describe the computation rather than identify it. Equal + subtrees derive equal guarantees under the same accuracy model and evidence. ```text Operator @@ -140,7 +205,9 @@ Operator ### 1.2 `NonASAPOp` `NonASAPOp` contains the operator variants split from `QueryExpr` in the -[decoupling proposal](decoupling_op_and_expr.md#2-types). +[decoupling proposal](decoupling_op_and_expr.md#2-types). Keeping scalar expressions +in operator fields makes the graph's edges represent table or state dependencies; +scalar expressions remain data interpreted against the owning operator's schema. ```rust pub enum NonASAPOp { @@ -166,7 +233,10 @@ Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted a ### 1.3 `ASAPOp` `ASAPOp` contains today's summary variants from `SummaryExpr` and the -summary-specific variants of `ValueOperation`. +summary-specific variants of `ValueOperation`. These remain distinct operations +because building state and estimating a value have different output types and +execution constraints. For example, one `SummaryAgg` can feed several estimates +without duplicating the summary state. ```rust pub enum ASAPOp { @@ -203,8 +273,10 @@ Legacy types map to the new IR as follows: ### 1.4 Child field -Non-ASAP operators now sit on the same level as ASAP operators, so their children must -be `Rc` to allow free placement: +Children use `Rc` for two reasons: `Operator` allows either category as +an input, and `Rc` lets multiple consumers reference the same node. Keeping children +as `Rc` would still prevent a relational operator from reading a summary; +storing children by value would represent separate copies instead of shared work: ```rust // After the decoupling doc // After this proposal From e9c2015454aea7c34c8094e0b94383917f7dcd5c Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:08:27 +0000 Subject: [PATCH 14/46] docs: distinguish shared operator types from shared DAG nodes --- .../design_docs/proposals/operator-sharing.md | 27 ++++++++++++------- 1 file changed, 18 insertions(+), 9 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 606826511..e10255cc3 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -129,11 +129,16 @@ Immutable nodes keep those plans independent: `with_*` and the attribute passes return new nodes. This requires allocation and explicit preservation of sharing within each resulting plan. -**Preserve sharing while filling attributes.** Guarantee derivation and timing -assignment return new nodes because the IR is immutable. If two queries share a -`Scan`, processing their roots independently could create two replacement scans and -lose that sharing. The design therefore requires each input node to map to one -output node within a pass, across every root in the workload: +**Preserve shared node instances while filling attributes.** Here, sharing means +that multiple parents—within one query or across queries in the same workload—refer +to the same node instance, such as one `Scan`. This differs from the title's use of +“sharing”: pre-ASAP and post-ASAP use the same operator *types*. The memo does not +make those two planning stages share node instances. + +Guarantee derivation and timing assignment return new nodes because the IR is +immutable. Processing two roots independently could turn their shared `Scan` into +two separate output scans. Each pass must instead preserve that common input as +one common output across all workload roots: ```text Before the pass After the pass @@ -153,10 +158,14 @@ Reusing it after a rewrite or for another assignment could return stale nodes. Today's `GlobalSelection.assembled_nodes` already uses this approach during assembly (`replacement.rs`); this proposal extends it to the attribute passes. -`DerivationMemo` below names that temporary table. Its name, storage and exposure in -these illustrative signatures are implementation details, not a new plan attribute -or a required public API. The design requirement is to preserve sharing and detect -timing conflicts. +`DerivationMemo` below names this per-pass table. It maps the input `Scan` to a +new `Scan′`; both output queries reference that same `Scan′`. The input and output +scans are distinct objects. Guarantee derivation and timing assignment each use +their own table. + +The table's name, storage and API are implementation details. The design requires +preserving shared node instances within the resulting workload DAG and rejecting +conflicting timings for any one of those nodes. ```rust /// Reuses each input node's output across all roots in one attribute pass. From 43a8d793e3c8793ff2f67991f47e1cc5fcfc5897 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:11:58 +0000 Subject: [PATCH 15/46] docs: focus operator proposals on design and rationale --- .../proposals/decoupling_op_and_expr.md | 150 ++- .../design_docs/proposals/operator-sharing.md | 855 +++++------------- 2 files changed, 271 insertions(+), 734 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index f593ea8bf..d48af6c7f 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -1,114 +1,86 @@ # Decoupling Operators From Scalar Expressions -> Status: proposal, not implemented. Audience: planner developers and designers. -> Companion to [Operator sharing](operator-sharing.md). Code references use `main` -> at `5a32b8b`. +> Status: proposal, not implemented. Audience: planner designers and architects. +> Companion to [Operator sharing](operator-sharing.md). -Split `QueryExpr` into `NonASAPOp` for table-producing operators and `ScalarExpr` -for expressions evaluated within an operator. This makes invalid combinations, -such as a literal used as a filter's table input, unrepresentable. +## 1. Problem and goal -``` -Today Proposed -Filter { pred: Rc, Filter { pred: Predicate(ScalarExpr), - child: Rc } child: Rc } -``` +The current query representation mixes table-producing operators and scalar +expressions. Their roles are distinguished by where they occur, so an expression +can be placed where a table input is expected and fail only when the plan is checked. -## 1. Problem +Separate these concepts so the plan model expresses which combinations are valid. +This also makes clear what the planner can replace or share as a computation. -``` -SELECT l_quantity * 2 AS q2 FROM lineitem WHERE l_quantity > 10 +Consider: -Project ← operator: outputs a table - cols: [Column(4) * Literal(2)] ← scalar expression: outputs one value per input row - child: Filter ← operator - pred: Column(4) > Literal(10) ← scalar expression - child: Scan lineitem ← operator +```sql +SELECT l_quantity * 2 AS q2 +FROM lineitem +WHERE l_quantity > 10 ``` -Operators produce tables and can be replaced or shared by the planner. Scalars -compute values within the schema chosen by their operator; `Column(4)` has no meaning without -that context. - -Today both are `QueryExpr` variants, separated only by convention: - -- `Filter { child: Literal(2) }` compiles, then fails with `ScalarHasNoRowSchema`. -- CSE and canonicalization skip scalars using repeated special-case branches. - -## 2. Types - -```rust -pub enum NonASAPOp { - Scan { .. }, Filter { pred: Predicate, child: Rc> }, Project { cols: Vec>, child }, - Aggregate { .. }, Join { .. }, SetOp { .. }, Concat { .. }, Dedup { .. }, Sort { .. }, Limit { .. }, BinaryOp { .. }, - SQLWindowFunc { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, - ScalarBridge(ScalarExpr), // formerly PromqlScalarBridge (the `2` in PromQL `v * 2`) - EvalTimestamp, // PromQL time() -} -pub enum ScalarExpr { - Column(C), Literal(ScalarValue), Compare { .. }, BoolAnd(..), BoolOr(..), Not(..), IsNull(..), IsNotNull(..), - Cast { .. }, InList { .. }, FunctionCall { .. }, Arithmetic { .. }, Case { .. }, CurrentTimestamp, -} -pub struct Predicate(pub ScalarExpr); -pub struct ProjectItem { pub alias: Option, pub expr: ScalarExpr } +```text +Scan lineitem → Filter → Project + │ │ + predicate expression + quantity quantity * 2 + > 10 ``` -Operators own their scalar fields by value. CSE hashes these fields as data; scalars -are not shared DAG nodes. Recursive scalar fields still require indirection such as -`Box` or `Vec`; the type sketches omit those details. +The scan, filter and projection produce tables. The predicate and multiplication +compute values within the schema selected by their owning operators. -```text -NonASAPOp -├─ children: Rc> (Rc> after operator sharing) -└─ scalar fields: Predicate / ProjectItem / SortKey / ScalarBridge / ... - └─ ScalarExpr - └─ children: ScalarExpr only, never an operator -``` +## 2. Design and rationale -`QueryExpr` is removed. `NonASAPOp` is named for its role in the companion proposal. -Scalar fields include predicates (`Filter`, `Join`, `HAVING`), project items, sort -keys, window-function arguments and relabel values. +| Concept | Meaning | Role in the plan | +|---|---|---| +| Operator | Produces a table, or summary state in the companion design | A graph node whose input computations may be replaced or shared | +| Scalar expression | Computes a value in a particular schema context | Part of an operator's predicate, projection, sort key or other expression | -Classify ambiguous variants by their position and schema: +Operator inputs must be other operators. Scalar expressions may contain other +scalar expressions, but do not contain operator subplans in this design. -- `ScalarBridge`, `EvalTimestamp`, `PromqlScalarFromVector` and - `PromqlVectorFromScalar` remain operators because they have row schemas. -- `CurrentTimestamp` (SQL `NOW()`) is a scalar. Its only operator-position use today - is a unit test. +A column reference has meaning only in its schema context. Two identical-looking +expressions in different operators may refer to different inputs. Keeping scalar +expressions attached to their operators preserves that context; this proposal does +not introduce independent shared scalar computations. -The split prevents scalars in operator positions and operators in scalar positions, -eliminating `ScalarHasNoRowSchema`. +The distinction depends on semantics, not on whether a result looks scalar. For +example, a PromQL conversion between a scalar and a vector participates in the +operator graph because it has query-level output semantics. SQL `NOW()` is an +expression evaluated within its owning operator. -## 3. Changes +## 3. Semantic requirements -| Location | Change | -|---|---| -| frontend expression lowering (`df_expr_to_unresolved`, PromQL `walk`) | scalar positions build `ScalarExpr`, operator positions `NonASAPOp` | -| `resolve`, `column_resolution.rs` | `resolve` takes `NonASAPOp`; `resolve_expr` takes `ScalarExpr`. Schema selection and leaf-schema inference are unchanged. | -| `canonicalize`, `pre_asap/cse.rs` | the "scalar: nothing to do" arms go; scalars are hashed as plain data | -| `scalar_signature.rs`, `infer_expr_type` | take `ScalarExpr` | -| `QueryExpr::output_schema` | becomes `NonASAPOp::output_schema`; the scalar arms and `ScalarHasNoRowSchema` go | +Column resolution keeps its existing meaning: -Scalars resolve against the schema chosen by their operator: usually the child's -output, the aggregate output for `HAVING`, both inputs for a join predicate, or the -scan schema for a scan predicate. Leaf schemas still use column references across -the whole tree. +- Most expressions use the input operator's output schema. +- A join predicate uses both input schemas. +- An aggregate's `HAVING` expression uses the aggregate output schema. +- A scan predicate uses the scanned data's schema. -## 4. Implementation and tests +Splitting the representation must preserve evaluation behavior, inferred output +types and source-language semantics. A scalar expression cannot serve as a table +input, and a table-producing operator cannot appear where a scalar is expected. -This is [stage 1 of operator sharing](operator-sharing.md#8-stages-and-tests). -Children stay `Rc` until stage 2 widens them to `Rc`. +The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) +extends the operator graph with summary operations. The scalar/operator distinction +continues to hold before and after that optimization: replacing a projection's input +with a summary estimate does not turn its scalar expressions into graph nodes. -The wire format stays unchanged: preserve the externally tagged variant names, -including `PromqlScalarBridge` via `#[serde(rename)]`. +## 4. Acceptance and scope -Update construction syntax in existing tests. Rewrite or remove tests that put -scalars in operator positions: the `CurrentTimestamp`, `ScalarHasNoRowSchema`, and -literal-as-`fallback` cases. +The example query must retain its result and output schema. Planning can replace or +share its table-producing computations while interpreting the filter predicate and +projection expression in the correct contexts. Invalid scalar/table combinations +must be excluded by the plan model. -## 5. Limits +This separation alone does not require a change to the external plan format. +The companion proposal addresses the separate decision to expose every operator in +the exported graph. -Scalars currently contain no operators: filter `IN (SELECT …)` and `EXISTS` lower -to semi-joins; other subquery-valued expressions are rejected -(`frontend-sql/src/sql/expr.rs`). Supporting scalar subqueries later would require -mutually recursive types and traversal into scalars by CSE and the planner. +Scalar subqueries are outside this design. Filter `IN (SELECT …)` and `EXISTS` can +be represented as joins, but a general scalar subquery introduces a dependency on +another operator graph. Supporting that requires a separate design for its scope, +dependencies and participation in optimization. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index e10255cc3..16ac27c02 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -1,703 +1,268 @@ # Sharing Operators Between Pre-ASAP IR and Post-ASAP IR -> Status: proposal, not implemented. Audience: planner developers and designers. -> Addresses [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468) and builds on -> [Decoupling operators from scalar expressions](decoupling_op_and_expr.md). -> Code references use `main` at `5a32b8b` (after #472, #478 and #510). +> Status: proposal, not implemented. Audience: planner designers and architects. +> Addresses [#468](https://github.com/ProjectASAP/ASAPPlanner/issues/468). +> Companion: [Decoupling operators from scalar expressions](decoupling_op_and_expr.md). -Use one operator language before and after ASAP optimization. Relational operators -and summary operators can be children of each other, without `KeepPreAsap` wrappers -or duplicate relational variants. +## Goal and problem -`Operator` has two categories: `NonASAP(NonASAPOp)` and `ASAP(ASAPOp)`. This separates -relational and summary operations while giving both the same child type, -`Rc`. Frontend inputs must contain only `NonASAP` nodes (§3). +Use one operator model before and after ASAP optimization, so ordinary query +operations and summary operations can form one visible computation graph. -Guarantees are derived from the DAG. Execution timings come from an explicit -lifecycle assignment and are validated before export (§2). +Today, the post-ASAP representation wraps relational subplans and duplicates some +relational operators outside those wrappers. This causes three problems: -``` -Today Proposed -ValueOperation(Project) ← a copy NonASAP(Project) - SummaryEstimate ASAP(SummaryEstimate) - SummaryAgg(Kll) ASAP(SummaryAgg(Kll)) - KeepPreAsap(Scan lineitem) ← a black box NonASAP(Scan lineitem) -``` - -| Part | Sections | -|---|---| -| I. New IR | §1 Types, §2 Schema, guarantee, and timing | -| II. Changes, in data-flow order | §3 Entry → §4 Planner → §5 Timing → §6 Export → §7 Other consumers | -| III. Implementation | §8 Stages and tests, §9 Out of scope, §10 Open questions | +- A projection above a summary needs a different representation from a projection + below it, although both perform the same operation. +- An exact aggregate cannot directly share a scan hidden inside a summary's input. +- An operator without a post-ASAP counterpart cannot naturally contain summary-based + children. ---- +The design removes these representation barriers. It makes composition and sharing +possible; whether a particular rewrite or shared computation is valid still depends +on query semantics, accuracy and execution timing. -# I. New IR - -## 1. Types +## 1. Operator model ### 1.1 Unified `Operator` type -**Why one operator type?** Today a relational subtree inside `KeepPreAsap` is -opaque to the surrounding summary plan. To put a `Project` above a summary, the -planner needs a second project representation; to share the enclosed `Scan` with -another computation, it must cross that wrapper. Both problems come from giving -relational and summary plans different node types. - -Use `Operator` as the common node type, with `Rc` children. Then -`Project → SummaryEstimate → SummaryAgg → Scan` is one traversable DAG. Assembly, -sharing and export can follow every edge using the same interface. Each relational -operator keeps one definition wherever it appears in the plan. - -**Why keep two categories?** Relational operators describe ordinary query semantics; -ASAP operators introduce summary state and its accuracy and lifecycle rules. Keeping -`NonASAPOp` and `ASAPOp` separate lets schema and validation code handle those rules -by category while common traversals work on `Operator`. A single flat enum could -also express the DAG; the two-level enum keeps that semantic distinction explicit, -at the cost of another match when dispatching to a variant. It does not prove that -a whole subtree is non-ASAP: children can contain either category, so optimizer -entry still validates frontend DAGs (§3). - -**Why a common interface?** Traversals need children, schema, guarantee and timing -for every node. These are methods on `Operator`; each category implements its own -rules. Data needed by only one operator, such as `Aggregate.measures` or -`SummaryAgg.family`, stays on that variant. `map_children` lets assembly replace -inputs while retaining the parent operator, including operators with no special -summary-planning logic. - -The methods below expose a node's inputs and attributes. They are illustrative -signatures, not a finalized API. `C` describes how column references are represented -(e.g. `ColumnId`); `C: ColState` requires that representation to support the column -operations used by the IR. - -| Method | Meaning and purpose | Example | -|---|---|---| -| `children()` | Read the node's immediate input operators. Generic traversals use this to walk the DAG without matching every operator kind. It does not return scalar expressions or all descendants. | A `Filter` has one input; a `Join` has two; a `Scan` has none. | -| `map_children(f)` | Return a copy of this operator with `f` applied to each immediate input, keeping its other fields. Assembly uses it to connect selected child plans under an existing parent. It does not recursively rewrite the DAG by itself. | Keep a `Project` and its expressions, but replace its input aggregate with a selected summary-estimation subtree. | -| `output_schema()` | Compute the output fields and their types, or report a schema error. Parents and export need this to interpret the node's result. | `SummaryAgg` outputs grouping fields and summary state; `SummaryEstimate` outputs grouping fields and estimated values. | -| `guarantee()` | Read the stored accuracy result without computing it. Selection and validation use the result after guarantee derivation. | `Unset`: not derived; `Set(Some(g))`: known guarantee, including exactness; `Set(None)`: no guarantee established. | -| `timing()` | Read the stored execution phase without choosing it. Export uses this after a lifecycle assignment has been applied. | `Unset`, `Set(IngestionTime)` or `Set(QueryTime)`. | -| `with_guarantee(g)` | Return a copy of this node with its guarantee slot set to `Set(g)`. The guarantee pass uses it to record its calculation. | `with_guarantee(None)` records that derivation found no guarantee; it does not leave the slot unset. | -| `with_timing(t)` | Return a copy of this node with its timing slot set to `Set(t)`. The timing pass uses it to record the lifecycle's choice. | `with_timing(QueryTime)` marks the copied node for query-time execution. | - -`&self` means the method reads the current node. `Self` means it returns a node of -this same type; `with_*` and `map_children` do not mutate the original or rewrite its -descendants. The passes are responsible for traversal and validation: setting a -slot alone does not prove that its value is valid. After replacing children, rerun -the attribute passes before export because the copied slots may be stale. - -`Rc` is a shared reference to a node. Thus `children()` returns borrowed references -to existing inputs, while the callback passed to `map_children` returns the shared -reference to use for each replacement input. - -```rust -pub enum Operator { - NonASAP(NonASAPOp), // today's relational and timeseries operators in `QueryExpr` (§1.2) - ASAP(ASAPOp), // summary operators (§1.3) -} - -impl Operator { // implemented for every operator - pub fn children(&self) -> Vec<&Rc>>; - pub fn map_children(&self, f: impl FnMut(&Rc>) -> Rc>) -> Self; - pub fn output_schema(&self) -> Result; // schema: §2.1 - pub fn guarantee(&self) -> &Slot>; // accuracy guarantee: §2.2 - pub fn timing(&self) -> &Slot; // execution timing: §2.3 - pub fn with_guarantee(&self, guarantee: Option) -> Self; - pub fn with_timing(&self, timing: ExecutionTiming) -> Self; -} -``` +Every computation is an operator node. Inputs are edges to other operator nodes, +including inputs from either of these two categories: -**Why defer attributes?** A node's complete accuracy guarantee depends on its -assembled inputs; its execution timing depends on the chosen lifecycle. Construction -therefore leaves both unresolved. Explicit passes fill them once that context is -available (§2), so constructors do not guess timings or duplicate error-composition -logic. Export reads the completed attributes and rejects an unfinished plan. - -`Slot` distinguishes “not computed yet” from a computed result. In particular, -`Unset` means the guarantee pass has not run, while `Set(None)` means it ran but -could not establish a guarantee. Collapsing those states would hide missing passes. -Schema is computed from the operator and its children rather than stored as another -copy that a rewrite could leave stale. - -```rust -pub enum Slot { Unset, Set(T) } -``` +| Category | Meaning | Examples | +|---|---|---| +| Ordinary query operators | Transform, combine or aggregate query data | Scan, filter, project, join, aggregate, union, time selection | +| ASAP operators | Build summary state or obtain results from it | Summary build, summary estimation, exact-state finalization, population maintenance and readout | -**Why return new nodes?** Candidate plans can share nodes. Mutating a node's timing -for one candidate could change another candidate that needs a different lifecycle. -Immutable nodes keep those plans independent: `with_*` and the attribute passes -return new nodes. This requires allocation and explicit preservation of sharing -within each resulting plan. +Both categories use the same graph model. A projection can consume a summary +estimate, and a summary can consume the result of a filter or join. There is no +separate relational subtree hidden inside a summary-plan node. -**Preserve shared node instances while filling attributes.** Here, sharing means -that multiple parents—within one query or across queries in the same workload—refer -to the same node instance, such as one `Scan`. This differs from the title's use of -“sharing”: pre-ASAP and post-ASAP use the same operator *types*. The memo does not -make those two planning stages share node instances. +The categories remain distinct because they have different semantic rules: +ordinary operators consume query values, while summary operations may produce or +consume state. A common graph lets planning reason about all dependencies; the +categories make state-specific accuracy and execution constraints explicit. A +frontend plan contains only ordinary query operators. Optimization may introduce +ASAP operators later. -Guarantee derivation and timing assignment return new nodes because the IR is -immutable. Processing two roots independently could turn their shared `Scan` into -two separate output scans. Each pass must instead preserve that common input as -one common output across all workload roots: +For example, arrows below show data flowing from producer to consumer: ```text -Before the pass After the pass -query A ─┐ query A′ ─┐ - ├─ shared Scan ├─ shared Scan′ -query B ─┘ query B′ ─┘ -``` - -A temporary lookup table can enforce this: record `input node → output node` on the -first visit and reuse the output on later visits. Key it by node identity, not -structural equality: the pass preserves existing sharing; it does not introduce -new sharing between separate computations. For timing, a repeated visit must also -agree with the timing already assigned, or the assignment is rejected (§4). - -Use a fresh table for each pass invocation, shared across its workload roots. -Reusing it after a rewrite or for another assignment could return stale nodes. -Today's `GlobalSelection.assembled_nodes` already uses this approach during assembly -(`replacement.rs`); this proposal extends it to the attribute passes. - -`DerivationMemo` below names this per-pass table. It maps the input `Scan` to a -new `Scan′`; both output queries reference that same `Scan′`. The input and output -scans are distinct objects. Guarantee derivation and timing assignment each use -their own table. - -The table's name, storage and API are implementation details. The design requires -preserving shared node instances within the resulting workload DAG and rejecting -conflicting timings for any one of those nodes. - -```rust -/// Reuses each input node's output across all roots in one attribute pass. -pub struct DerivationMemo { .. } - -/// Build the accuracy guarantee of one DAG root by derivation -pub fn derive_guarantees( - root: &Rc, - model: &dyn AccuracyModel, // accuracy model used for derivation - evidence: &dyn AccuracyEvidenceProvider, // evidence provider used for derivation - memo: &mut DerivationMemo, -) -> Result, AccuracyError>; -/// Write the timings of one lifecycle assignment into the `timing` slots of one DAG -/// root, top-down, then validate them (§5). Summary materialization chooses the -/// assignment (§2.3); nothing is inferred from operator kinds. A shared node reached -/// with two different timings is an error (§4). -pub fn apply_lifecycle_timings( - root: &Rc, - assignment: &LifecycleAssignment, // per-node timings, expanded from per-state lifecycle choices (§10) - memo: &mut DerivationMemo, -) -> Result, ExecutionDataStateError>; +Scan → Filter → Summary build → Summary estimation → Project ``` -- **Pass results:** lifecycle planning uses the guarantee-derived DAG; export uses - the timed DAG. Earlier pointers do not identify nodes in those results. -- **Recomputation:** `derive_guarantees` fills guarantee slots from local evidence - and child guarantees. `apply_lifecycle_timings` overwrites timing slots from the - assignment. Run rewrites before these passes. -- **Equality:** timing participates in `PartialEq` and `Hash` because computations - assigned to different phases cannot be merged into one execution. Derived - guarantees do not: they describe the computation rather than identify it. Equal - subtrees derive equal guarantees under the same accuracy model and evidence. +Each node describes its operation, inputs and output schema. An executable plan also +needs an accuracy assessment and an execution phase for each relevant computation. +These have different sources, described in §2; they are not all known when a logical +operator is first created. -```text -Operator -├─ NonASAP(NonASAPOp) -│ ├─ children: Rc> → back to Operator: NonASAP or ASAP -│ ├─ timing / guarantee -│ └─ scalar expressions: Predicate / ProjectItem / SortKey / ScalarBridge / ... -│ └─ ScalarExpr: never contains an Operator -└─ ASAP(ASAPOp) - ├─ children: Rc> → back to Operator: NonASAP or ASAP - └─ timing / guarantee -``` +### 1.2 Operators and scalar expressions -### 1.2 `NonASAPOp` - -`NonASAPOp` contains the operator variants split from `QueryExpr` in the -[decoupling proposal](decoupling_op_and_expr.md#2-types). Keeping scalar expressions -in operator fields makes the graph's edges represent table or state dependencies; -scalar expressions remain data interpreted against the owning operator's schema. - -```rust -pub enum NonASAPOp { - Scan { .. }, - Filter { pred: Predicate, child: Rc> }, - Project { cols: Vec>, child: Rc> }, - Aggregate { reduction, measures, having: Option>, child: Rc> }, - Join { kind, pred: Predicate, left: Rc>, right: Rc> }, - SetOp { kind, all, left: Rc>, right: Rc> }, - Concat { children: Vec>> }, - Sort { keys: Vec>, child: Rc> }, - Limit { n, offset, child: Rc> }, - BinaryOp { op, lhs, rhs }, - SQLWindowFunc { args: Vec>, order_by: Vec>, child: Rc>, .. }, - Dedup { .. }, TimeRange { .. }, TimeShift { .. }, Promql* { .. }, - ScalarBridge(ScalarExpr), // the `2` in PromQL `v * 2` - EvalTimestamp, // PromQL time() -} -``` +A filter is an operator because it transforms a table. Its predicate, such as +`latency > 100`, is a scalar expression evaluated in that table's schema. -Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. - -### 1.3 `ASAPOp` - -`ASAPOp` contains today's summary variants from `SummaryExpr` and the -summary-specific variants of `ValueOperation`. These remain distinct operations -because building state and estimating a value have different output types and -execution constraints. For example, one `SummaryAgg` can feed several estimates -without duplicating the summary state. - -```rust -pub enum ASAPOp { - SummaryAgg { child: Rc>, family: ASAPType, input, reduction, grouping, - exact_rule: Option }, - SummaryEstimate { child: Rc>, query: SketchQuery, - local_guarantee: Option }, - SummaryMerge { children: Vec>> }, - SummarySubtract { left: Rc>, right: Rc> }, - SummaryDelete { child: Rc>, key: C }, - SummaryJoin { outer: Rc>, inner: Rc>, key: C, family: ASAPType }, - FinalizeExactAccumulator { child: Rc> }, - MaintainPopulation { child: Rc>, population }, - ReadPopulation { child: Rc>, readout }, - Extension { child: Rc>, name: String }, -} -``` +Keep scalar expressions within their owning operators. Graph edges then represent +computation dependencies, while predicates, projection expressions and sort keys +describe how an operator processes its input. This prevents an expression from +being mistaken for a table-producing plan. The +[companion proposal](decoupling_op_and_expr.md) defines this distinction. -Every variant also carries the `timing` and `guarantee` slots (§1.1), omitted above. +### 1.3 Two meanings of sharing -`SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` -are currently built only in tests. Migrate their variants, but return `Unimplemented` -from schema, guarantee, timing and export operations. +**Sharing the operator model** means pre-ASAP and post-ASAP use the same definitions +for ordinary operators. It does not mean those two planning stages execute together +or must reference the same node instances. -Legacy types map to the new IR as follows: +**Sharing a computation** means multiple consumers within a workload use one +producer. For example, two queries may read one scan, or two estimates may use one +summary: -| Legacy types | Expressed as | -|---|---| -| `SummaryExpr::KeepPreAsap(q)` | `q` itself, an `NonASAP(..)` subtree | -| `ValueOperation::{Project, Filter, Sort, Limit}` | `NonASAPOp::{Project, Filter, Sort, Limit}` | -| `SummaryExpr::{BinaryOp, RelationalJoin}` | `NonASAPOp::{BinaryOp, Join}` | -| `ValueOperation::Exact(Aggregate)`, `ExactOperation` | `NonASAPOp::Aggregate` | -| `SummaryNode` | `Operator` itself: `schema` is computed, `timing` / `guarantee` are slots on every variant (§2) | - -### 1.4 Child field - -Children use `Rc` for two reasons: `Operator` allows either category as -an input, and `Rc` lets multiple consumers reference the same node. Keeping children -as `Rc` would still prevent a relational operator from reading a summary; -storing children by value would represent separate copies instead of shared work: - -```rust -// After the decoupling doc // After this proposal -Filter { pred: Predicate(ScalarExpr), Filter { pred: Predicate(ScalarExpr), - child: Rc } child: Rc } // NonASAP(..) or ASAP(SummaryEstimate ..) +```text + ┌→ p50 estimation → query A +Scan → KLL summary build ┤ + └→ p99 estimation → query B ``` -Now an original operator can also sit on ASAP operators, e.g. a `SetOp` sitting on two `SummaryEstimate` operators. +The shared producer must satisfy every consumer's input, window, accuracy and timing +requirements. Representing it once exposes reuse to planning and costing. Keeping +multiple query roots in one workload graph is therefore part of the design. -`Concat.children` is `Vec` today: branches are stored by value and have no `Rc` identity, so the planner (§4), which identifies targets by pointer, -cannot replace a branch — e.g. the branches of SQL `ROLLUP` or PromQL `histogram_quantiles`. It becomes `Vec>` (§8 stage 0). +Adding accuracy information or execution timing must preserve that sharing. It must +also leave alternative candidate plans independent: assigning a lifecycle to one +candidate must not change another candidate's choices. How nodes are stored or +reused is outside this design. -## 2. Schema, Guarantee, and Timing +## 2. Node properties and why they differ -| Attribute | Proposed rule | -|---|---| -| `schema` | Computed by `output_schema()` using one schema type for all nodes | -| `guarantee` | Binding records local evidence; `derive_guarantees()` fills the slots bottom-up | -| `timing` | Unset during assembly; `apply_lifecycle_timings()` fills the slots from an explicit assignment | - -A set guarantee of `None` means no guarantee is available. It is different from an -`Unset` slot, which means the pass has not run. - -### 2.1 Schema: fused into one type - -Pre-ASAP uses `Schema` with plain `DataType` columns; post-ASAP stores a separate -`SummarySchema` that also supports summary state. Use `Schema` for both: - -- Rename `Column` to `Field` and `Schema.columns` to `Schema.fields`. -- Widen `Field.dtype` to `FieldType`, covering plain values and ASAP state. -- Delete `SummarySchema` and `SummaryField`. - -`Schema` keeps the reserved column `PROMQL_SERIES_IDENTITY` (`"$promql_series_identity"`, -`pre_asap/schema.rs`) and `has_promql_series_identity()`: `maintained_population.rs` uses -it to decide whether a closed PromQL schema still identifies a series (§7). - -```rust -pub enum FieldType { DataType(DataType), ASAPType(ASAPType) } -pub enum ASAPType { // SummaryFamilyType without Plain - ExactAggregate(ExactKind, ExactParams), Sketch(SketchKind, GroupingStrategy), - Sample(SamplingKind, SamplingParams), Wavelet(WaveletKind, WaveletParams), StatModel(StatModelKind, StatModelParams), -} -pub struct Schema { pub fields: Vec, pub time_index, pub unique_keys, pub closed } -pub struct Field { pub name, pub dtype: FieldType, pub nullable, pub table: Option } -impl Field { - pub fn plain(name, DataType) -> Self; - pub fn plain_dtype(&self) -> Option<&DataType>; // None for a state column - pub fn expect_plain_dtype(&self) -> &DataType; // frontends, scalar type inference; panics on state -} -``` - -| Node | Today | After | +| Property | Meaning | How it is determined | |---|---|---| -| `NonASAPOp` | post-ASAP `KeepPreAsap`: `QueryExpr` schema lifted to `SummarySchema` and stored
post-ASAP `ValueOperation` / `BinaryOp` / `RelationalJoin` copies: stored at construction | using the same logic as `QueryExpr::output_schema()` | -| `SummaryAgg` | the replaced `Aggregate`'s output with the measure column retyped to `family` | grouping columns + one `ASAPType(family)` column | -| `SummaryEstimate` | the replaced operator's output schema | the child's grouping columns + the value columns of the `SketchQuery` | -| `FinalizeExactAccumulator` | the logical operator's output, lifted | the child's schema, `ASAPType(ExactAggregate ..)` columns changed into `DataType(..)` | -| `MaintainPopulation` / `ReadPopulation` | the source's schema / the replaced aggregate's output | the same rules, computed from the child and the `readout` | -| unused variants | one field typed `family` | unimplemented | - -### 2.2 Guarantee: always derived - -1. Binding records `SummaryEstimate.local_guarantee` (error over exact input) and - exact `SummaryAgg.exact_rule`, leaving guarantee slots unset. To size a sketch, it derives - the child's guarantee on a temporary copy, reads it, then drops the copy. -2. `derive_guarantees` fills slots bottom-up using the rules below. - -``` -Project p99 ±1% ← the child's - SummaryEstimate p99 ±1% ← local ±1%, composed with Scan t's: looks through the SummaryAgg - SummaryAgg(Kll) None ← state has no guarantee - Scan t exact ← no ASAP descendant -``` +| Output schema | What the node produces: field names, types and relevant identity/time information | From the operation and its inputs | +| Accuracy guarantee | What can be established about the result's accuracy | From local accuracy evidence and the guarantees of its inputs | +| Execution timing | Whether work runs at ingestion time or query time | From a lifecycle choice for the complete plan | -| Node | Guarantee rule | -|---|---| -| `SummaryEstimate` | Compose `local_guarantee` with the guarantee of the input below `SummaryAgg`; sketch state itself has no guarantee. The local guarantee is `None` if the family has no error model. | -| `SummaryAgg` | ExactAggregate: compose with the child under `exact_rule`; `ExactKind::Count` stays exact regardless of the child. Sketch state: `Set(None)`. | -| `NonASAPOp` | Compose child guarantees; exact when there is no ASAP descendant. | -| `FinalizeExactAccumulator` | Copy the child's guarantee. | -| `MaintainPopulation` / `ReadPopulation` | Exact. | -| Unused variants | Return `Unimplemented` (§1.3). | - -### 2.3 Timing: written from a lifecycle assignment - -Logical planning decides what to compute. Summary materialization chooses a lifecycle -for each summary state, including execution timing and retention. One logical DAG can -have several assignments; the planner selects among them using deployment-provided -costs, consistent with the [planning-stages proposal](https://github.com/ProjectASAP/ASAPPlanner/pull/509). - -Assembly leaves every timing slot `Unset`. `apply_lifecycle_timings` writes the -assignment with one memo across all roots and validates node and edge constraints -(§5). It rejects ingestion work that depends on query-time results, fixed-kind timing -violations, and conflicting timings on a shared node. Export and physical compilation -reject unset slots. +### 2.1 One schema model for values and state -``` - assembly after apply_lifecycle_timings -Project Unset QueryTime - SummaryEstimate Unset QueryTime - SummaryAgg Unset IngestionTime ← the state's lifecycle: maintained - Scan t Unset IngestionTime ← feeds a maintained state -``` +A common graph needs a common description of its edges. Schemas must distinguish +ordinary values from summary state, so a consumer can determine whether an input is +usable. -Exact compositions (`OperationPlacement::Read` / `Maintenance`), maintained -populations and #472's grouped `Rate`→`Sum` pair currently fix timing during -construction. Each becomes a logical candidate whose placement is a lifecycle choice -(§8, stage 5). +For example, a KLL build produces state; its p99 estimation produces a numeric value. +A numeric predicate can consume the estimate, but cannot treat the KLL state itself +as a number. Ordinary operators may carry state through only where their semantics +permit it; exact aggregate state must be finalized before use as an ordinary value. -The default assignment aims to preserve today's timings. It cannot do so when one -shared node requires both phases; resolving that conflict remains open (§10). +Schema information must preserve grouping fields, time information, uniqueness and +series identity where relevant. Sharing operators must not change SQL or PromQL +meaning. After a rewrite, schemas must describe the new inputs rather than the plan +that was replaced. -| Node | Timing constraint | -|---|---| -| `NonASAPOp` | Assigned timing must satisfy every consuming edge. | -| `SummaryAgg` | Assigned by the state's lifecycle. | -| `FinalizeExactAccumulator` | Either phase, subject to its inputs and consumers. | -| `SummaryEstimate` | Query time only. | -| `MaintainPopulation` / `ReadPopulation` | Ingestion time / query time only. | -| Unused variants | Return `Unimplemented` (§1.3). | +### 2.2 Accuracy follows the computation -Apply timing after assembly because it depends on the whole DAG and its lifecycle. +A summary's local error describes its behavior over an exact input. That alone is +not the guarantee of a larger query: its input may already be approximate, and +later operations may change the error. The planner must compose accuracy through +the actual computation graph. -### 2.4 Workflow +For example, a KLL estimate over exact input can carry the sketch's guarantee. If +that input is approximate, the estimate must also account for the upstream error. +The summary state itself is not a query answer and need not have a value-level +accuracy guarantee. -Today, construction and assembly compose guarantees, and export derives remaining -timings into a side table. The proposed sequence is: +Ordinary exact computations remain exact when their inputs and operation semantics +justify it. Exact accumulator finalization preserves the established guarantee. +Special cases, such as exact counting of the rows actually received, retain their +operation-specific rules. -```text -search / binding record local accuracy evidence; leave timing unset - check accuracy on a temporary derived copy (§4) -assembly build the selected roots, preserving shared nodes -derive_guarantees fill guarantees bottom-up, with one memo across roots -lifecycle choose per-state lifecycles using the derived DAG -apply_lifecycle_timings - write and validate timing, with one memo across roots -export read filled slots; reject any Unset slot -``` - ---- - -# II. Changes, in data-flow order - -## 3. Optimizer entry - -Frontends and `resolve` build `NonASAP` trees only and access children with -`expect_non_asap()`. `search_cse_workload_with`, which every `search_workload*` entry -reaches, panics on a root that `contains_asap()`: an ASAP node there is a caller bug. - -```rust -impl Operator { - pub fn contains_asap(&self) -> bool; - pub fn expect_non_asap(&self) -> &NonASAPOp; // an ASAP node here is a bug: panic -} -``` +Distinguish an assessment that has not happened from an assessment that found no +supported guarantee. Missing accuracy evidence cannot be treated as exactness or +as proof that a query's accuracy target is met. Candidate assessment and final-plan +validation must use the same accuracy semantics. -Library users do not call `search_workload*` themselves. Since #478 the external -boundary is `asap_aware_mapping::pass::optimize` (reached from `asap_planner::e2e_plan`), -and since #510 a pass always produces lifecycle-aware plans: +### 2.3 Timing is a planning choice -``` -asap_planner::e2e_plan - → asap_aware_mapping::pass::optimize - → MajorPass::optimize - → search_workload_with_targets - → global_selection_with_summary_maintenance_lifecycles - → assemble_selected_dag_with_summary_maintenance_lifecycles (one call per root) -``` +The same logical summary can be maintained as data arrives or computed when a query +needs it. Its position in the graph alone does not choose between these behaviors. +Logical planning therefore leaves timing undecided. Physical planning chooses +materialization and lifecycle behavior, then determines execution phases across the +complete graph. -```rust -pub struct QueryLifecyclePlan { pub entry_index: usize, pub plan: SummaryMaintenanceLifecyclePlan } -pub struct PlanOutput { pub plans: Vec } // one per workload entry -impl PlanOutput { pub fn dags(&self) -> Vec> } // each plan.root; becomes Rc -``` +The planner evaluates alternatives using deployment-provided cost and accuracy +models and capabilities. The deployment executes the selected plan, consistent with +the [planning-stages proposal](https://github.com/ProjectASAP/ASAPPlanner/pull/509). -`ParsedWorkload` (`asap_types::parsed_workload`) holds the roots as `Vec>` -and becomes `Vec>`. The entry check above stays in `search_cse_workload_with`; -the pass layer adds no check of its own. See -[`updated_interface_with_pluggable_optimization.md`](../architecture/updated_interface_with_pluggable_optimization.md). +A valid timing assignment must satisfy these constraints: -A compile-time alternative — an associated type on `ColState` with -`ColumnRef::ASAP = Never` — only protects frontend code before `resolve`: frontends -already return `ColumnId` trees, where `ASAP` is allowed. The entry check covers every -input (frontends after `resolve`, deserialized plans, test IR) with simpler types. +- Ingestion-time work cannot depend on a query-time result. +- Summary estimation runs at query time. Population maintenance and readout run at + ingestion time and query time, respectively. +- A shared computation has one execution phase compatible with all its consumers. + A single producer cannot simultaneously mean two separate executions. +- Stored state remains available for as long as its consumers need it. -## 4. Planner: search and assembly +This proposal covers materialization at summary-state boundaries. Materializing +arbitrary ordinary intermediate results requires a separate design. -```rust -pub enum Replacement { - Subtree(Rc), // formerly Summary(Rc) and Rewrite(Rc) - ExactComposition { .. }, // its plan becomes Rc -} -``` +## 3. Planning responsibilities -**Runtime support.** `runtime_support_evidence` dispatches on the replacement root: -`ASAP` uses `summary_support_evidence`; `NonASAP` returns `Some(true)`. Nested ASAP -nodes were checked when their own candidates were built. - -**Candidate construction.** Candidates remain bottom-up and contain concrete child -plans (`realize_child_with`, `prepare_compositions`). The old `keep_pre_asap` fallback -returns the original child subtree directly. - -**Accuracy checks.** Search derives each candidate's guarantee with a fresh memo, -checks its target, and stores the root guarantee on `ReplacementSubDAG` for selection. -It drops the derived copy and keeps the original shared candidate in `PlanSpace`. - -**Assembly** — one rule replaces `assemble_residual`: - -```rust -fn assemble(&self, t: &Rc) -> Rc { - memo by ptr; // shared children stay one Rc - let composed = matches!(self.chosen(t), // the chosen summary already folds the - Some(Subtree(r)) if r is ASAP(SummaryAgg { child, .. }) // inner aggregate: nothing is hidden - && !child.contains_asap() && !contains_aggregate(child)); - let chosen = if query_time_nested_sum(t) && !composed { None } // as today: keep the outer SUM so the - else { self.chosen(t) }; // inner target's own choice is assembled - match chosen { - Some(Subtree(r)) => r, // a complete plan, used as is - Some(ExactComposition{..}) => composition.plan, - None => t.map_children(|c| if is_target(c) { self.assemble(c) } else { c }), - // keep the node, assemble its children — assemble_residual does this for four operators only - } -} -``` +The common representation separates what a plan computes from how it executes: -`GlobalSelection::assemble_selected_dag` derives guarantees after assembly, -preserving shared nodes across roots as described in §1.1. Lifecycle planning reads -that derived DAG, then -`apply_lifecycle_timings` applies the chosen assignment. - -`assemble_selected_query` remains the query-result boundary (#472): it calls -`assemble_selected_dag`, then `finalize_query_candidate` to insert a -`FinalizeExactAccumulator` where needed. Binding and the `BinaryOp` / `Join` callers -keep this finalization step. The lifecycle assembly path still bypasses it (§9). - -Timing validation moves from candidate construction to lifecycle assignment: - -- Materialization must offer only valid assignments. Failure when applying a selected - assignment is a planning error. -- A shared node must have one timing across all roots. The pass rejects a conflict; - it does not split the node by phase. Plans needing both phases must already contain - distinct nodes (§10). - -`map_children` is `rebuild_children` from -`pre_asap/cse.rs`, dispatching to `NonASAPOp::map_children` / `ASAPOp::map_children`. -Deleted: `assemble_residual`, `keep_pre_asap` / `keep_pre_asap_rc`, and the -`KeepPreAsap` branch of `finalize_exact_accumulator`. Kept: `relink_summary`, -`assemble_selected_query` / `finalize_query_candidate`, and the `query_time_nested_sum` -special case with its #472 exception — the chosen candidate is a `SummaryAgg` whose -child is a `NonASAP` subtree without an `Aggregate` — where `contains_aggregate` takes an -`Operator` instead of a `QueryExpr`. - -| #468 problem | Resolution | +| Responsibility | Required result | |---|---| -| 1. A `Project` is a `QueryExpr` inside `KeepPreAsap` and a `ValueOperation` outside | one set of types | -| 2. Exact and summary operators cannot share a scan hidden in `KeepPreAsap` | Both can reference one `Scan` when their timings agree. Different phases require distinct nodes. Splitting a multi-measure aggregate remains out of scope (§9). | -| 3. `SetOp` and similar have no post-ASAP copy, so no summary below them | `SetOp` takes `None => t`; both children are assembled | +| Frontend translation | An ordinary query graph preserving source-language semantics | +| Logical optimization | Exact and summary-based alternatives, including legal shared computations | +| Candidate assessment | Accuracy and capability evidence for the actual candidate graph | +| Physical and lifecycle planning | Executable alternatives with materialization, retention and compatible timing | +| Plan selection | A valid plan chosen using workload-level costs and requirements | +| Export and execution | The selected graph with explicit dependencies and completed assessments | -## 5. Timing: validating the assignment +Ordinary operators remain in the graph when their inputs are replaced by summary +computations. For example, replacing an aggregate below a projection must not require +replacing the projection with a separate post-ASAP operator. -`apply_lifecycle_timings` replaces timing derivation with validation. It checks the -assigned timing against node kinds and consuming edges; the old -`ExecutionDataStateAssignment` side table is removed. +Changing a candidate's inputs can change its schema, accuracy and legal timing. +Those properties must be checked against the resulting graph. A plan with unresolved +execution timing or an unfinished accuracy assessment is not ready for export. -| Node | Produced state | -|---|---| -| `NonASAPOp` | `{assigned timing, Raw / …}`; checked against every consuming edge (`QUERY_ROWS` at the root) | -| `SummaryAgg` | `{assigned timing, SummaryState}` (§2.3) | -| other `ASAPOp` | today's rules, checked against the assigned timing | - -A data state is the `timing` slot plus a primitive (`Raw` / `SummaryState` / …) fixed by -the kind; only the timing is stored. - -The `KeepPreAsap` / `BinaryOp` / `ValueOperation` / `RelationalJoin` arms of today's -`validate_execution_data_states` merge into one `NonASAP` arm of that validation: - -- Check the edge to each child; an `ASAP` child is checked by the `ASAP` edge rules. - Ingestion work cannot depend on a query-time result. -- `check_plain_operands` stays: referenced columns must be `FieldType::DataType` - (`Project` / `Filter` / `Sort` / `Limit` may pass `ExactAggregate` columns through). - This rejects `Project(ASAP(SummaryAgg))`. -- `BinaryOp`'s ingestion-side constraints move into this arm. -- `AmbiguousKeepPreAsap` is deleted: a shared subtree assigned two timings is rejected by - `apply_lifecycle_timings` (§4). - -## 6. Export: one post-ASAP node per operator - -The relational cases of `fallback{expression: QueryExpr}`, `binary`, `value` and -`relational_join` become one payload: - -```rust -PostAsapOperatorPayload::Relational { - /// One non-ASAP operator: its kind, expressions and parameters, without its child - /// fields. The children are this node's incoming edges, in child-role order. - operator: NonASAPOpKind, -} -``` +This proposal changes the representation, not the search strategy. It does not +require a new enumeration algorithm or change when candidates commit to particular +child plans. -Export emits one node per operator, with children represented by edges. A shared -operator is exported once. Each relational node stores its kind, scalar expressions -and parameters in `NonASAPOpKind`; it embeds no child subtree or `DagInput` placeholder. - -Physical compilation lowers each node to one operator or a few helper operators. -Relational lowering reads inputs from edges, replacing the old whole-expression -`fallback` and the `binary`, `value::Project` and `relational_join` special cases. -ASAP nodes retain their summary payloads; `FinalizeExactAccumulator`, -`MaintainPopulation` and `ReadPopulation` retain `value{operation}`. - -- **Wire version 6:** replace relational payloads with per-operator `relational` - nodes and use `Schema` for output and intermediate schemas. Update downstream - readers together in stage 4 (§8). -- **Attributes:** export reads derived guarantees and assigned timings, rejecting - unset slots. Edge data state combines the producer's timing with its primitive (§5). -- **Compilation:** `compile_post_asap_dag` splits precompute and query DAGs by assigned - timing without repeating validation. -- **Phase boundaries:** assignments must satisfy §5. This proposal materializes only - ASAP state; assignments requiring materialization elsewhere are rejected. -- **`SummaryMerge`:** keep its wire payload; planner-side support remains open (§10). - -## 7. Other consumers - -| Location | Change | -|---|---| -| `post_asap/cse.rs` | delete; `share_common_subtrees` covers `ASAPOp` (derives `PartialEq` + serde) | -| `dag_export.rs` | delete `build_summary` / `build_summary_hybrid` / `summary_kind_tag`; one exporter with an `ASAP` arm; update the viewer's `node-style.js` and the pin test `viewer_categorizes_exactly_the_exported_node_kinds` | -| `summary_maintenance_cost/estimator.rs` (90 `SummaryExpr::` sites, 13 `KeepPreAsap`) | `KeepPreAsap` branches (`query_source_selections`, `retained_queries`) use the §6 `Relational` nodes; `exact_binary` / `value_operation` costs fold into it | -| `summary_maintenance_cost/evidence.rs` | `summary_operation_evidence` gets an `ASAP` arm; also fixes its missing `RelationalJoin` case | -| `physical_plan_cost_model.rs::estimate_candidate` | every `Relational` node goes through the per-operator lowering (§6) in `lower_query_physical_dag` | -| `summary_maintenance_lifecycle.rs` | `SummaryMaintenanceLifecyclePlan.root` becomes `Rc`; `selected_raw_recompute` becomes `!contains_asap(root)`; the `keep_pre_asap(target)` fallback in `assemble_selected_dag_with_summary_maintenance_lifecycles` becomes `target` | -| `pass/mod.rs`, `pass/major.rs` | `PlanOutput::dags()` returns `Vec>`; `MajorPass` otherwise unchanged (§3) | -| `types/parsed_workload.rs` | `ParsedWorkload` roots become `Rc` (§3) | -| `asap-planner` (`planner/src/lib.rs`, `tests/e2e_plan.rs`) | follows `PlanOutput` | -| `maintained_population.rs` | `KeepPreAsap(source)` becomes `source`; `population.matches_input` reads an `NonASAP` child directly and keeps the `has_promql_series_identity()` check (§2.1) | -| `replacement.rs::enumerate_candidate_dags`, `CandidateDagInventory` | walk `Rc` instead of `SummaryNode`; internal and test use only | -| `exact_composition.rs` | `ExactOperation::Aggregate` becomes a `NonASAP(Aggregate)`, built over each child candidate as `prepare_compositions` does today | -| `RelationalJoin.pruning` | never set to `Some` in production; delete. Candidate pruning can return as an `ASAPOp` variant | - ---- - -# III. Implementation +## 4. Example: one scan serving exact and approximate queries -## 8. Stages and tests +Consider an exact average and an approximate p99 over the same input interval. +Once a rewrite exposes the two computations, the common model can express: -`main` builds and passes all tests after every stage. +```text + ┌→ Exact average ─────────────────────→ query A +Scan ──┤ + └→ KLL summary build → p99 estimation → query B +``` -| Stage | Change | Main consumers | -|---|---|---| -| 0. Prepare | Give `Concat.children` `Rc` identity; rename `rebuild_children` to `map_children`; add `Column::plain`. | `asap-types` | -| 1. Split | Separate `NonASAPOp` and `ScalarExpr`; keep `Rc` children. | [Scalar consumers](decoupling_op_and_expr.md#3-changes) | -| 2. Unify operators | Add `Operator`, an empty `ASAPOp`, entry helpers, and unset attribute slots. Widen children to `Rc`. | All crates | -| 3. Unify schemas | Introduce `Field`, `FieldType`, `ASAPType` and the shared `Schema`. Keep wire types and serialized names unchanged. | Schema consumers | -| 4. Switch export | Fill `ASAPOp`, add attribute passes and entry validation, and export the new IR with wire version 6. | Types, devtools, viewer, backend, collector | -| 5. Migrate planning | Build candidates and assemble directly on `Rc`; move timing choices to lifecycle planning; migrate §7 consumers. | Mapping and planner crates | -| 6. Remove legacy code | Delete old IR types, duplicate operators, CSE and timing fallbacks; update architecture and developer docs. | All remaining consumers | - -Stage 4 is the only wire-breaking stage and must land with ASAPQuery-backend and -ASAPCollector updates. A temporary `flatten(&SummaryNode) -> Rc` adapter -copies existing guarantees and applies existing timings where valid. Wire schemas -switch from `SummarySchema` to `Schema`, and `Schema.fields` serializes as `fields` -instead of `columns`. - -Stage 5 deletes `flatten`. Binding leaves timing unset; lifecycle planning handles -the placement choices listed in §2.3. Paths without lifecycle selection use the -default assignment, subject to the shared-node conflict in §10. - -Stage 6 removes `SummaryExpr`, `SummaryNode`, duplicate `ValueOperation` variants, -`ExactOperation`, `post_asap/cse.rs`, `produced_data_state` defaults and -`validate_execution_data_states_at`. `PlanOutput::dags()` and `ParsedWorkload` finish -moving to `Operator`. Update `post-asap-ir.md`, `physical-plan-integration.md`, -`updated_interface_with_pluggable_optimization.md`, and developer/viewer docs. - -**Tests**: - -- One integration test per #468 problem: - 1. `WITH metric AS (SELECT avg(CASE WHEN l_quantity BETWEEN 1 AND 50 THEN 1.0 ELSE 0.0 END) AS in_range FROM lineitem) SELECT in_range, in_range = 1.0 AS ok FROM metric` — no post-ASAP-only node besides `ASAP`; all `Project`s are one variant. - 2. `SELECT avg(l_extendedprice), approx_percentile_cont(l_discount, 0.99) FROM lineitem` — the `avg` `Aggregate` and the KLL `SummaryAgg` share one `Scan` by `Rc::ptr_eq` under an assignment that runs both at query time (once a binding rule splits measures). - 3. `SELECT approx_distinct(l_partkey) FROM lineitem UNION ALL SELECT approx_distinct(l_suppkey) FROM lineitem` — each side of the `SetOp` has a `SummaryEstimate`. -- A shared `Scan` assigned two timings is rejected; assigned one timing, it stays one `Rc`. -- A node shared by two roots, assembled in two calls, is still one `Rc` after `derive_guarantees` and `apply_lifecycle_timings`. -- After both passes no slot is `Unset`; export rejects a tree with one, and a tree with no assignment applied. -- The default assignment preserves existing timings where valid; mixed-phase sharing is covered by the open question in §10. -- `apply_lifecycle_timings` rejects an assignment that puts ingestion work over a query-time result, and one that gives a `SummaryEstimate` ingestion time. -- A kept `NonASAP` node (e.g. a `SetOp`) reports the guarantee composed from its assembled children. -- Each unused branch (§1.3) returns `Unimplemented` from `output_schema`, `derive_guarantees`, `apply_lifecycle_timings` and export. -- `search_workload*` panics on a root containing `ASAP`. -- Rewrite the 110 `SummaryExpr::` assertions in `sql_to_post_asap.rs` / `promql_to_post_asap.rs` / `exact_composition.rs`. -- Wire 6 round trip of a DAG with `Relational` nodes both above and below `ASAP` nodes; a version-5 document is rejected. -- The 16 `execution_data_state.rs` tests keep their shapes; they apply an assignment and read the `timing` slot instead of `ExecutionDataStateAssignment`. - -## 9. Out of scope - -- The binding rule splitting a multi-measure `Aggregate` into exact + summary over one child. -- Candidate pruning as an `ASAPOp` variant. -- Accuracy through `SummaryMerge` / `Subtract` / `Delete` / `Join` (§2.2): an accuracy - descriptor on state, or composing error along the state chain at readout. -- Holes: letting a chosen plan's child be filled by that child target's own choice at - assembly, instead of fixing it when the candidate is built. A search-strategy change, - independent of the types here. -- Folding `ExactComposition` into `Subtree` (both of its forms become expressible); - deferred until stage 5 is stable. -- `assemble_selected_dag_with_summary_maintenance_lifecycles`, the only assembly `MajorPass` - uses, calls `assemble_selected_dag` rather than `assemble_selected_query` (§4), so pass - output carries no root `FinalizeExactAccumulator`. Not changed here. - -## 10. Open questions - -- **Assignment type:** reuse the type returned by #482's - `SummaryMaintenanceLifecyclePlan::execution_timed_dag()` to expand state lifecycles - into per-node timings, rather than define a competing type. -- **Summary merge:** does ASAPQuery insert `SummaryMerge` only in the exported DAG, - or through planner-side types? The latter remains unimplemented. -- **Mixed-phase sharing:** one shared `Scan` may feed an ingestion-time `SummaryAgg` - and a query-time `Aggregate`. The default assignment then conflicts. Decide whether - assembly creates separate subtrees or materialization gives the shared state one - lifecycle. The timing pass itself must not silently split it. +If both branches run at query time, they may share the scan. Computing accuracy and +assigning timing must retain that one producer for both consumers. + +If the KLL is maintained at ingestion time while the exact average reads raw data at +query time, the depicted scan cannot serve both phases as one execution. The plan +needs separate scans or a different lifecycle that makes sharing valid. The planner +must represent that choice explicitly; timing validation rejects the conflicting +shared plan. + +The representation enables the shared plan but does not supply the rewrite that +splits a multi-measure aggregate into these branches. That rewrite is outside this +proposal. How planning resolves the mixed-phase case remains open (§7). + +## 5. Export preserves the graph + +Export one node per operator and represent its input dependencies as edges. Export +a shared producer once, with edges to all its consumers. + +This keeps the graph visible to costing, physical compilation, execution and plan +inspection. Embedding a whole relational subtree in one exported node would hide +its internal sharing and recreate the original boundary problem. + +Export carries the resolved schemas, assessed guarantees and assigned execution +phases. Physical compilation may lower one logical operation to several physical +operations, but must preserve its dependencies and meaning. The execution layer +does not invent missing planning decisions. + +Changing the exported representation requires coordinated adoption by the planner +and downstream readers. Representation unification must preserve query semantics; +it does not by itself guarantee unchanged timings for plans with unresolved sharing +conflicts. + +## 6. Acceptance criteria + +The design is successful when: + +- A projection uses the same semantics above and below summary computations. +- An exact aggregate and a summary can share an input when all requirements agree. +- A union or another ordinary operator can consume summary estimates on its inputs. +- Shared producers remain shared when accuracy and timing are resolved, including + across different queries in the workload. +- Invalid value/state combinations, incompatible timing and unfinished plan + assessments are rejected before execution. +- Export preserves visible dependencies and shared producers. + +## 7. Scope and open questions + +This proposal defines a common operator model and its correctness constraints. It +does not define storage structures, public APIs, traversal algorithms, serialization +fields or a code migration sequence. + +The following decisions remain separate or unresolved: + +- **Mixed-phase sharing:** when consumers need different phases, should planning + separate the producer or find a common materialized lifecycle? Existing timing + behavior cannot be promised until this is resolved. +- **Summary composition:** where are merge operations introduced—during planner + optimization or downstream? Accuracy rules for merge, subtract, delete and + summary joins need a separate design. A representable operation is not a claim of + runtime support. +- **Lifecycle integration:** reuse the lifecycle concepts being developed in + [#482](https://github.com/ProjectASAP/ASAPPlanner/pull/482), without creating a second + competing source of timing decisions. +- **Additional optimization:** rules that split exact and approximate measures, + new pruning strategies and deferred child-plan choices are outside this proposal. +- **Finalization:** this representation change does not resolve the existing + difference between query-result assembly and lifecycle-plan assembly in adding + finalization of exact aggregate state. From 0fd2c3946ddf6b72cdb63bf29ebc01dd2da6bdb7 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:18:35 +0000 Subject: [PATCH 16/46] docs: illustrate operator representation barriers in the introduction --- .../design_docs/proposals/operator-sharing.md | 20 +++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 16ac27c02..8c4c8cec5 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -18,6 +18,26 @@ relational operators outside those wrappers. This causes three problems: - An operator without a post-ASAP counterpart cannot naturally contain summary-based children. +For example, consider a p99 latency query that projects its input columns, builds a +KLL summary, and projects the estimated result. The trees below read from the result +at the top to the data source at the bottom: + +```text +Today Proposed +Post-ASAP projection Project +└─ Summary estimation └─ Summary estimation + └─ KLL summary build └─ KLL summary build + └─ Wrapped relational subplan └─ Project + └─ Ordinary projection └─ Scan latency + └─ Scan latency +``` + +Today the two projections need separate representations, and the scan is hidden +inside the wrapped subplan. In the proposed graph, both projections use the same +operator definition and the scan is directly visible. An exact aggregate can also +consume that scan when sharing is valid, as shown in §4. Similarly, a union can +consume summary estimates without needing a separate post-ASAP union definition. + The design removes these representation barriers. It makes composition and sharing possible; whether a particular rewrite or shared computation is valid still depends on query semantics, accuracy and execution timing. From bb0ac8dc57b59a4f2bec965a5eda01b0a39b7f66 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:19:03 +0000 Subject: [PATCH 17/46] docs: retain NonASAP and ASAP in the operator type design --- .../design_docs/proposals/operator-sharing.md | 28 +++++++++++++------ 1 file changed, 20 insertions(+), 8 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 8c4c8cec5..e89555650 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -46,13 +46,24 @@ on query semantics, accuracy and execution timing. ### 1.1 Unified `Operator` type -Every computation is an operator node. Inputs are edges to other operator nodes, -including inputs from either of these two categories: +Every computation is represented by an `Operator` node. The type has two categories: -| Category | Meaning | Examples | +```rust +enum Operator { + NonASAP(NonASAPOp), + ASAP(ASAPOp), +} +``` + +| Category | Meaning | Example operations | |---|---|---| -| Ordinary query operators | Transform, combine or aggregate query data | Scan, filter, project, join, aggregate, union, time selection | -| ASAP operators | Build summary state or obtain results from it | Summary build, summary estimation, exact-state finalization, population maintenance and readout | +| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Filter`, `Project`, `Join`, `Aggregate`, `SetOp`, time selection | +| `ASAP(ASAPOp)` | Operations that build summary state or obtain results from it | `SummaryAgg`, `SummaryEstimate`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation` | + +`NonASAPOp` and `ASAPOp` describe the operation performed by a node. Inputs in both +categories connect to `Operator` nodes, so either category can consume the other +when their schema and execution constraints permit it. `NonASAP` classifies one +node; it does not require all of that node's descendants to be non-ASAP. Both categories use the same graph model. A projection can consume a summary estimate, and a summary can consume the result of a filter or join. There is no @@ -62,13 +73,14 @@ The categories remain distinct because they have different semantic rules: ordinary operators consume query values, while summary operations may produce or consume state. A common graph lets planning reason about all dependencies; the categories make state-specific accuracy and execution constraints explicit. A -frontend plan contains only ordinary query operators. Optimization may introduce -ASAP operators later. +frontend plan contains only `NonASAP` nodes. Optimization may introduce `ASAP` +nodes later. For example, arrows below show data flowing from producer to consumer: ```text -Scan → Filter → Summary build → Summary estimation → Project +NonASAP(Scan) → NonASAP(Filter) → ASAP(SummaryAgg) + → ASAP(SummaryEstimate) → NonASAP(Project) ``` Each node describes its operation, inputs and output schema. An executable plan also From 58fc2ce999cfaa3642eae9d19b81028141b166c1 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:21:06 +0000 Subject: [PATCH 18/46] docs: list all proposed NonASAP and ASAP operations --- docs/design_docs/proposals/operator-sharing.md | 17 ++++++++++++++--- 1 file changed, 14 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index e89555650..9b732edd4 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -55,10 +55,21 @@ enum Operator { } ``` -| Category | Meaning | Example operations | +The table lists all operator kinds in this proposal. Aggregate functions, join +kinds and scalar functions are choices within these operations, not additional +operator kinds. + +| Category | Meaning | All operations | |---|---|---| -| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Filter`, `Project`, `Join`, `Aggregate`, `SetOp`, time selection | -| `ASAP(ASAPOp)` | Operations that build summary state or obtain results from it | `SummaryAgg`, `SummaryEstimate`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation` | +| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `ScalarBridge`, `EvalTimestamp`, `PromqlVectorFromScalar`, `PromqlScalarFromVector`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | +| `ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | + +`ScalarBridge` is the proposed name for the existing `PromqlScalarBridge`. +`CurrentTimestamp` belongs to scalar expressions, so it is not in this operator list. + +`SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` +are reserved in the proposed planner model; listing them does not establish planner +or runtime support. Summary composition remains an open design question (§7). `NonASAPOp` and `ASAPOp` describe the operation performed by a node. Inputs in both categories connect to `Operator` nodes, so either category can consume the other From fb84379746f01096137aaaad1b9614a15c17f82e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:24:30 +0000 Subject: [PATCH 19/46] docs: specify NonASAP and ASAP operation data structures --- .../design_docs/proposals/operator-sharing.md | 120 +++++++++++++++++- 1 file changed, 117 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 9b732edd4..02157e33f 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -87,6 +87,120 @@ categories make state-specific accuracy and execution constraints explicit. A frontend plan contains only `NonASAP` nodes. Optimization may introduce `ASAP` nodes later. +**Proposed data structures.** The sketches below show the operation-specific data +carried by each category. They use Rust-like notation to describe the design, not +final API signatures. `NodeRef` means an edge to another `Operator` node; it does +not prescribe a pointer or storage type. Multiple edges may refer to one producer. +`ScalarExpr` means an expression within an operator, not another graph node. + +`NonASAPOp` retains the query semantics needed before and after optimization: + +```rust +enum NonASAPOp { + Scan { + source: Source, predicates: Vec, schema: Schema, + }, + Filter { child: NodeRef, predicate: ScalarExpr }, + Project { + child: NodeRef, columns: Vec, qualifier: Option, + }, + Aggregate { + child: NodeRef, reduction: Reduction, measures: Vec, + output_names: Vec, having: Option, + }, + Join { left: NodeRef, right: NodeRef, kind: JoinKind, predicate: ScalarExpr }, + SetOp { left: NodeRef, right: NodeRef, kind: SetOpKind, all: bool }, + Concat { + children: Vec, discriminator_unique_key: Option, + }, + Dedup { child: NodeRef, columns: Vec }, + Sort { child: NodeRef, keys: Vec, partition_by: GroupKeys }, + Limit { child: NodeRef, count: usize, offset: usize, partition_by: GroupKeys }, + BinaryOp { + left: NodeRef, right: NodeRef, operation: BinaryOpKind, + vector_match: Option, + }, + SQLWindowFunc { + child: NodeRef, function: WindowFunction, args: Vec, + partition_by: GroupKeys, order_by: Vec, + frame: WindowFrame, output_name: String, + }, + TimeRange { child: NodeRef, range: Duration }, + TimeShift { child: NodeRef, shift: TimeShiftSpec }, + ScalarBridge { expression: ScalarExpr }, + EvalTimestamp, + PromqlVectorFromScalar { child: NodeRef }, + PromqlScalarFromVector { child: NodeRef }, + PromqlRelabel { child: NodeRef, destination_label: String, value: ScalarExpr }, + PromqlInfoEnrich { child: NodeRef, selector: Vec }, + PromqlSeriesSample { child: NodeRef, by: GroupKeys, kind: SampleKind }, + PromqlSubquery { child: NodeRef, range: Duration, resolution: Option }, +} +``` + +The fields describe what an operator does to its inputs: + +- `child`, `left`, `right` and `children` are graph dependencies. They can lead to + either operator category, subject to the input's schema requirements. +- Predicates and named expressions describe row-level calculations. A named + expression contains a scalar expression and its optional output alias. +- `reduction` describes whether aggregation combines groups or operates per entity; + `measures` describes the requested aggregates. Grouping is distinct from ordering + or limiting within groups, represented by `partition_by`. +- Join/set kinds, vector matching, window frames and time selections preserve + source-language semantics. Output names, qualifiers and proven uniqueness also + survive optimization. The optional concatenation key records a discriminator + that distinguishes branches together with their within-branch key. + +`ASAPOp` describes state construction, state operations and readout separately: + +```rust +enum ASAPOp { + SummaryAgg { + child: NodeRef, family: SummaryFamily, input: SummaryUpdate, + reduction: Reduction, grouping: GroupingStrategy, + exact_rule: Option, + }, + SummaryEstimate { + child: NodeRef, query: SummaryQuery, + local_guarantee: Option, + }, + FinalizeExactAccumulator { child: NodeRef }, + MaintainPopulation { child: NodeRef, population: PopulationSpec }, + ReadPopulation { child: NodeRef, readout: PopulationReadout }, + + // Reserved operations; semantics and support require further design. + SummaryMerge { children: Vec }, + SummarySubtract { left: NodeRef, right: NodeRef }, + SummaryDelete { child: NodeRef, key: ColumnRef }, + SummaryJoin { + outer: NodeRef, inner: NodeRef, key: ColumnRef, family: SummaryFamily, + }, + Extension { child: NodeRef, name: String }, +} +``` + +The summary fields distinguish state construction, readout and accuracy evidence: + +| Field | Design meaning | +|---|---| +| `family` | The summary or exact accumulator chosen, including its family-specific parameters | +| `input` | The item identity and observation or weight supplied to a state update | +| `reduction` | Which input entities contribute to each logical result | +| `grouping` | Whether those groups use separate state instances or a supported shared structure | +| `query` / `readout` | The result requested from summary or maintained-population state | +| `local_guarantee` / `exact_rule` | Local accuracy evidence or composition semantics; neither is the final guarantee of the complete subtree | +| `population` | The population whose membership and values are maintained | + +For example, one KLL `SummaryAgg` can feed two `SummaryEstimate` nodes whose queries +request p50 and p99. The build operation and its state are shared; the requested +estimates differ. + +Schema, derived accuracy and execution timing describe every `Operator`, regardless +of category (§2). They are omitted from these operation-specific sketches. Timing +comes from lifecycle planning rather than a fixed field value implied by an operator +kind; the final accuracy assessment combines local evidence with the actual inputs. + For example, arrows below show data flowing from producer to consumer: ```text @@ -288,9 +402,9 @@ The design is successful when: ## 7. Scope and open questions -This proposal defines a common operator model and its correctness constraints. It -does not define storage structures, public APIs, traversal algorithms, serialization -fields or a code migration sequence. +This proposal defines a common operator model and its correctness constraints. It includes the operation-specific data needed to express those semantics, but +does not prescribe storage structures, public APIs, traversal algorithms, +serialization fields or a code migration sequence. The following decisions remain separate or unresolved: From 49770ee8935e2e7e5ce0aadc6d5e62013cd3d76e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:26:46 +0000 Subject: [PATCH 20/46] docs: distinguish operator unification from related design changes --- .../design_docs/proposals/operator-sharing.md | 35 +++++++++++++++++++ 1 file changed, 35 insertions(+) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 02157e33f..2c6ec5378 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -87,6 +87,41 @@ categories make state-specific accuracy and execution constraints explicit. A frontend plan contains only `NonASAP` nodes. Optimization may introduce `ASAP` nodes later. +**Relationship to the current code.** These structures are proposals, not copies +of the current definitions with two category labels added. The operations come from +the existing query and summary models, but the sketches combine several changes: + +| Change | Purpose and scope | +|---|---| +| Group ordinary operations under `NonASAP` and summary operations under `ASAP` | Organize the common operator model into two semantic categories. | +| Give both categories inputs that refer to `Operator` nodes | Enable composition and shared producers across the category boundary. Two category labels alone do not enable sharing. | +| Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | +| Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | +| Introduce `local_guarantee` and `exact_rule` on summary operations | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). These are not fields on today's corresponding summary operations. | +| Assign execution timing through lifecycle planning | Additional planning design (§2.3), replacing timing stored or inferred differently by today's operators. | + +Existing operation semantics and field names should be retained unless a change is +identified explicitly. The sketches use descriptive shorthand in several places; +those names do not propose additional types or renames: + +| Sketch notation | Current counterpart | +|---|---| +| `columns: Vec` | `cols: Vec` on a projection | +| `predicate` | `pred`, currently wrapped in `Predicate` | +| `AggregateMeasure` | `AggIntent` | +| `SummaryFamily` | The state-producing cases of `SummaryFamilyType` | +| `SummaryQuery` | `SketchQuery` | +| `AccuracyGuarantee` / `AccuracyCompositionRule` | `ResultGuarantee` / `CompositionOperator` | +| `PopulationSpec` | `MaintainedPopulation` | +| `child` on summary estimation or deletion | `summary_input` | + +Two differences need particular care. The proposed `Limit.partition_by` preserves a +capability of today's post-ASAP limit that the pre-ASAP limit does not carry; +unifying those definitions requires an explicit decision about its semantics. +The window-function sketch shows a resolved `frame`, while the current definition +allows an optional frame for compatibility. Neither difference should be treated +as an incidental rename or a silently approved behavior change. + **Proposed data structures.** The sketches below show the operation-specific data carried by each category. They use Rust-like notation to describe the design, not final API signatures. `NodeRef` means an edge to another `Operator` node; it does From c045963af223baf7b9adc97e77b1dac15d77d47b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:28:17 +0000 Subject: [PATCH 21/46] docs: clarify timing decisions and proposed accuracy fields --- .../design_docs/proposals/operator-sharing.md | 32 ++++++++++++++++++- 1 file changed, 31 insertions(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 2c6ec5378..fd0871240 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -98,7 +98,7 @@ the existing query and summary models, but the sketches combine several changes: | Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | | Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | | Introduce `local_guarantee` and `exact_rule` on summary operations | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). These are not fields on today's corresponding summary operations. | -| Assign execution timing through lifecycle planning | Additional planning design (§2.3), replacing timing stored or inferred differently by today's operators. | +| Let physical/lifecycle planning decide when each node executes | Additional planning design (§2.3): logical construction leaves timing undecided; planning chooses ingestion-time or query-time execution under the workload constraints. | Existing operation semantics and field names should be retained unless a change is identified explicitly. The sketches use descriptive shorthand in several places; @@ -315,6 +315,20 @@ not the guarantee of a larger query: its input may already be approximate, and later operations may change the error. The planner must compose accuracy through the actual computation graph. +**Existing concepts versus proposed fields.** The current code already calculates +local guarantees and uses accuracy-composition rules, but neither `local_guarantee` +nor `exact_rule` is a field on the corresponding summary operation today: + +| Proposed field | Meaning | Current code | What this proposal changes | +|---|---|---|---| +| `SummaryEstimate.local_guarantee` | The guarantee for this summary readout over exact input; it excludes upstream error. | `AccuracyModel.local_guarantee(...)` already computes it. It is used when composing the result guarantee, rather than retained on the estimation operation. | Retain that local evidence so accuracy can be derived from the assembled graph's actual inputs. | +| `SummaryAgg.exact_rule` | The rule for propagating input error through an exact aggregate. “Exact” describes the operation, not a promise that approximate inputs become exact. | Composition rules already exist. Exact-aggregate binding selects a rule from the aggregate's registered semantics; there is no `exact_rule` field. | Retain the selected rule on the operation so later derivation does not need to recover the original binding context. | + +Today, the composed result is stored as the summary node's `guarantee`. The proposed +fields retain inputs to that calculation; they do not replace the complete result +guarantee. Adding them is an accuracy-design change, not a requirement of dividing +operators into `NonASAP` and `ASAP`. + For example, a KLL estimate over exact input can carry the sketch's guarantee. If that input is approximate, the estimate must also account for the upstream error. The summary state itself is not a query answer and need not have a value-level @@ -332,6 +346,22 @@ validation must use the same accuracy semantics. ### 2.3 Timing is a planning choice +**Who decides when a node executes?** This concerns responsibility for choosing +an execution phase, not memory ownership or ownership of the data. + +| Current behavior | Proposed behavior | +|---|---| +| Some operations receive timing during construction; other timings are inferred from their inputs or consuming edges. | Logical construction describes what to compute. Physical/lifecycle planning explicitly chooses when it runs, using the complete workload and execution constraints. | + +For example, a KLL summary may be maintained at ingestion time or built at query +time. The presence of a summary-build node does not itself choose either option. +The planner compares those alternatives; the selected lifecycle determines timing +for the build and its dependencies. Operator-specific restrictions still apply, +such as summary estimation running at query time. + +This changes where the timing decision is made. It is a separate planning change, +not an automatic consequence of putting operators into two categories. + The same logical summary can be maintained as data arrives or computed when a query needs it. Its position in the graph alone does not choose between these behaviors. Logical planning therefore leaves timing undecided. Physical planning chooses From a2bf014eff318fe1bcdb42d7bf25c77b0f76d3e3 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:30:35 +0000 Subject: [PATCH 22/46] docs: align proposed operator fields with existing code names --- .../design_docs/proposals/operator-sharing.md | 148 +++++++++--------- 1 file changed, 73 insertions(+), 75 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index fd0871240..e49b8affb 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -61,10 +61,10 @@ operator kinds. | Category | Meaning | All operations | |---|---|---| -| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `ScalarBridge`, `EvalTimestamp`, `PromqlVectorFromScalar`, `PromqlScalarFromVector`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | +| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlScalarBridge`, `EvalTimestamp`, `PromqlVectorFromScalar`, `PromqlScalarFromVector`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | | `ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | -`ScalarBridge` is the proposed name for the existing `PromqlScalarBridge`. +`PromqlScalarBridge` keeps its existing name. `CurrentTimestamp` belongs to scalar expressions, so it is not in this operator list. `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` @@ -97,88 +97,84 @@ the existing query and summary models, but the sketches combine several changes: | Give both categories inputs that refer to `Operator` nodes | Enable composition and shared producers across the category boundary. Two category labels alone do not enable sharing. | | Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | | Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | -| Introduce `local_guarantee` and `exact_rule` on summary operations | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). These are not fields on today's corresponding summary operations. | +| Introduce `local_guarantee` and `exact_operation_rule` on summary operations | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). These are not fields on today's corresponding summary operations. | | Let physical/lifecycle planning decide when each node executes | Additional planning design (§2.3): logical construction leaves timing undecided; planning chooses ingestion-time or query-time execution under the workload constraints. | -Existing operation semantics and field names should be retained unless a change is -identified explicitly. The sketches use descriptive shorthand in several places; -those names do not propose additional types or renames: - -| Sketch notation | Current counterpart | -|---|---| -| `columns: Vec` | `cols: Vec` on a projection | -| `predicate` | `pred`, currently wrapped in `Predicate` | -| `AggregateMeasure` | `AggIntent` | -| `SummaryFamily` | The state-producing cases of `SummaryFamilyType` | -| `SummaryQuery` | `SketchQuery` | -| `AccuracyGuarantee` / `AccuracyCompositionRule` | `ResultGuarantee` / `CompositionOperator` | -| `PopulationSpec` | `MaintainedPopulation` | -| `child` on summary estimation or deletion | `summary_input` | - -Two differences need particular care. The proposed `Limit.partition_by` preserves a -capability of today's post-ASAP limit that the pre-ASAP limit does not carry; -unifying those definitions requires an explicit decision about its semantics. -The window-function sketch shows a resolved `frame`, while the current definition -allows an optional frame for compatibility. Neither difference should be treated -as an incidental rename or a silently approved behavior change. - -**Proposed data structures.** The sketches below show the operation-specific data -carried by each category. They use Rust-like notation to describe the design, not -final API signatures. `NodeRef` means an edge to another `Operator` node; it does -not prescribe a pointer or storage type. Multiple edges may refer to one producer. -`ScalarExpr` means an expression within an operator, not another graph node. +**Naming and compatibility.** The sketches retain current operation, field and +payload-type names. `NonASAPOp`, `ASAPOp` and the companion proposal's `ScalarExpr` +are the new structural concepts; ordinary payloads such as `Predicate`, +`ProjectItem`, `AggIntent`, `SketchQuery` and `SummaryFamilyType` keep their names. + +The following changes are explicit: + +- Operator inputs become references to the common `Operator`. The sketches use the + existing `Rc` notation for shared inputs and omit column-state generics for + readability; column IDs and scan schemas below show the resolved form. +- `Predicate`, `ProjectItem` and other scalar-bearing payloads keep their roles, but + contain `ScalarExpr` after the scalar/operator split. +- `BinaryOp` reuses the existing post-ASAP `BinaryOperator` payload. It carries the + binary operation, vector matching and checked-division requirements; the pre-ASAP + `op` and `vector_match` semantics must be preserved when mapped into it. +- `Limit.partition_by` comes from the existing post-ASAP limit. Applying that field + to the unified operator remains an explicit design choice, not new functionality + implied by a rename. `SQLWindowFunc.frame` retains its current optional form. +- `SummaryEstimate.local_guarantee` and `SummaryAgg.exact_operation_rule` are proposed + new fields, named after the existing accuracy-model concepts. Their types remain + `ResultGuarantee` and `CompositionOperator`; neither field exists on today's + corresponding summary operation (§2.2). + +**Proposed data structures.** These sketches describe operation-specific data using +current names. `Rc` represents a shared input edge; no new reference type +is introduced. Storage and traversal algorithms remain outside this design. `NonASAPOp` retains the query semantics needed before and after optimization: ```rust enum NonASAPOp { Scan { - source: Source, predicates: Vec, schema: Schema, + source: Source, predicates: Vec, schema: Schema, }, - Filter { child: NodeRef, predicate: ScalarExpr }, + Filter { child: Rc, pred: Predicate }, Project { - child: NodeRef, columns: Vec, qualifier: Option, + child: Rc, cols: Vec, qualifier: Option, }, Aggregate { - child: NodeRef, reduction: Reduction, measures: Vec, - output_names: Vec, having: Option, + child: Rc, reduction: Reduction, measures: Vec, + output_names: Vec, having: Option, }, - Join { left: NodeRef, right: NodeRef, kind: JoinKind, predicate: ScalarExpr }, - SetOp { left: NodeRef, right: NodeRef, kind: SetOpKind, all: bool }, + Join { left: Rc, right: Rc, kind: JoinKind, pred: Predicate }, + SetOp { left: Rc, right: Rc, kind: RelationalSetOpKind, all: bool }, Concat { - children: Vec, discriminator_unique_key: Option, - }, - Dedup { child: NodeRef, columns: Vec }, - Sort { child: NodeRef, keys: Vec, partition_by: GroupKeys }, - Limit { child: NodeRef, count: usize, offset: usize, partition_by: GroupKeys }, - BinaryOp { - left: NodeRef, right: NodeRef, operation: BinaryOpKind, - vector_match: Option, + children: Vec>, discriminator_unique_key: Option, }, + Dedup { child: Rc, cols: Vec }, + Sort { child: Rc, keys: Vec, partition_by: GroupKeys }, + Limit { child: Rc, n: usize, offset: usize, partition_by: GroupKeys }, + BinaryOp { lhs: Rc, rhs: Rc, operator: BinaryOperator }, SQLWindowFunc { - child: NodeRef, function: WindowFunction, args: Vec, + child: Rc, func: WindowFuncKind, args: Vec, partition_by: GroupKeys, order_by: Vec, - frame: WindowFrame, output_name: String, + frame: Option, output_name: String, }, - TimeRange { child: NodeRef, range: Duration }, - TimeShift { child: NodeRef, shift: TimeShiftSpec }, - ScalarBridge { expression: ScalarExpr }, + TimeRange { child: Rc, range: Duration }, + TimeShift { child: Rc, shift: TimeShift }, + PromqlScalarBridge(ScalarExpr), EvalTimestamp, - PromqlVectorFromScalar { child: NodeRef }, - PromqlScalarFromVector { child: NodeRef }, - PromqlRelabel { child: NodeRef, destination_label: String, value: ScalarExpr }, - PromqlInfoEnrich { child: NodeRef, selector: Vec }, - PromqlSeriesSample { child: NodeRef, by: GroupKeys, kind: SampleKind }, - PromqlSubquery { child: NodeRef, range: Duration, resolution: Option }, + PromqlVectorFromScalar(Rc), + PromqlScalarFromVector(Rc), + PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, + PromqlInfoEnrich { child: Rc, selector: Vec }, + PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, + PromqlSubquery { child: Rc, range: Duration, resolution: Option }, } ``` The fields describe what an operator does to its inputs: -- `child`, `left`, `right` and `children` are graph dependencies. They can lead to +- `child`, `left`, `right`, `lhs`, `rhs` and `children` are graph dependencies. They can lead to either operator category, subject to the input's schema requirements. -- Predicates and named expressions describe row-level calculations. A named - expression contains a scalar expression and its optional output alias. +- `Predicate` describes a row-level condition; `ProjectItem` contains a scalar + expression and its optional output alias. - `reduction` describes whether aggregation combines groups or operates per entity; `measures` describes the requested aggregates. Grouping is distinct from ordering or limiting within groups, represented by `partition_by`. @@ -187,31 +183,33 @@ The fields describe what an operator does to its inputs: survive optimization. The optional concatenation key records a discriminator that distinguishes branches together with their within-branch key. -`ASAPOp` describes state construction, state operations and readout separately: +`ASAPOp` describes state construction, state operations and readout separately. +`SummaryFamilyType` retains its current name; state-producing operations use its +summary or exact-accumulator cases, never its `Plain` case. ```rust enum ASAPOp { SummaryAgg { - child: NodeRef, family: SummaryFamily, input: SummaryUpdate, + child: Rc, family: SummaryFamilyType, input: SummaryUpdate, reduction: Reduction, grouping: GroupingStrategy, - exact_rule: Option, + exact_operation_rule: Option, // proposed new field }, SummaryEstimate { - child: NodeRef, query: SummaryQuery, - local_guarantee: Option, + summary_input: Rc, query: SketchQuery, + local_guarantee: Option, // proposed new field }, - FinalizeExactAccumulator { child: NodeRef }, - MaintainPopulation { child: NodeRef, population: PopulationSpec }, - ReadPopulation { child: NodeRef, readout: PopulationReadout }, + FinalizeExactAccumulator { child: Rc }, + MaintainPopulation { child: Rc, population: MaintainedPopulation }, + ReadPopulation { child: Rc, readout: PopulationReadout }, // Reserved operations; semantics and support require further design. - SummaryMerge { children: Vec }, - SummarySubtract { left: NodeRef, right: NodeRef }, - SummaryDelete { child: NodeRef, key: ColumnRef }, + SummaryMerge { children: Vec> }, + SummarySubtract { left: Rc, right: Rc }, + SummaryDelete { summary_input: Rc, key: ColumnRef }, SummaryJoin { - outer: NodeRef, inner: NodeRef, key: ColumnRef, family: SummaryFamily, + outer: Rc, inner: Rc, key: ColumnRef, family: SummaryFamilyType, }, - Extension { child: NodeRef, name: String }, + Extension { child: Rc, name: String }, } ``` @@ -224,7 +222,7 @@ The summary fields distinguish state construction, readout and accuracy evidence | `reduction` | Which input entities contribute to each logical result | | `grouping` | Whether those groups use separate state instances or a supported shared structure | | `query` / `readout` | The result requested from summary or maintained-population state | -| `local_guarantee` / `exact_rule` | Local accuracy evidence or composition semantics; neither is the final guarantee of the complete subtree | +| `local_guarantee` / `exact_operation_rule` | Local accuracy evidence or composition semantics; neither is the final guarantee of the complete subtree | | `population` | The population whose membership and values are maintained | For example, one KLL `SummaryAgg` can feed two `SummaryEstimate` nodes whose queries @@ -317,12 +315,12 @@ the actual computation graph. **Existing concepts versus proposed fields.** The current code already calculates local guarantees and uses accuracy-composition rules, but neither `local_guarantee` -nor `exact_rule` is a field on the corresponding summary operation today: +nor `exact_operation_rule` is a field on the corresponding summary operation today: | Proposed field | Meaning | Current code | What this proposal changes | |---|---|---|---| | `SummaryEstimate.local_guarantee` | The guarantee for this summary readout over exact input; it excludes upstream error. | `AccuracyModel.local_guarantee(...)` already computes it. It is used when composing the result guarantee, rather than retained on the estimation operation. | Retain that local evidence so accuracy can be derived from the assembled graph's actual inputs. | -| `SummaryAgg.exact_rule` | The rule for propagating input error through an exact aggregate. “Exact” describes the operation, not a promise that approximate inputs become exact. | Composition rules already exist. Exact-aggregate binding selects a rule from the aggregate's registered semantics; there is no `exact_rule` field. | Retain the selected rule on the operation so later derivation does not need to recover the original binding context. | +| `SummaryAgg.exact_operation_rule` | The rule for propagating input error through an exact aggregate. “Exact” describes the operation, not a promise that approximate inputs become exact. | Composition rules already exist. Exact-aggregate binding selects a rule from the aggregate's registered semantics; the accuracy model also exposes `exact_operation_rule(...)` for exact operations. Neither is a stored field on `SummaryAgg`. | Retain the selected rule on the operation so later derivation does not need to recover the original binding context. | Today, the composed result is stored as the summary node's `guarantee`. The proposed fields retain inputs to that calculation; they do not replace the complete result From 0fe132f7279d9845bb8bfecac8264f671c18170e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:31:37 +0000 Subject: [PATCH 23/46] docs: keep operator sharing focused on common definitions --- .../design_docs/proposals/operator-sharing.md | 38 ++++--------------- 1 file changed, 8 insertions(+), 30 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index e49b8affb..a72bfe98f 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -94,7 +94,7 @@ the existing query and summary models, but the sketches combine several changes: | Change | Purpose and scope | |---|---| | Group ordinary operations under `NonASAP` and summary operations under `ASAP` | Organize the common operator model into two semantic categories. | -| Give both categories inputs that refer to `Operator` nodes | Enable composition and shared producers across the category boundary. Two category labels alone do not enable sharing. | +| Give both categories inputs that refer to `Operator` nodes | Allow ordinary and summary operations to compose directly, without a wrapper hiding their dependencies. | | Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | | Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | | Introduce `local_guarantee` and `exact_operation_rule` on summary operations | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). These are not fields on today's corresponding summary operations. | @@ -225,10 +225,6 @@ The summary fields distinguish state construction, readout and accuracy evidence | `local_guarantee` / `exact_operation_rule` | Local accuracy evidence or composition semantics; neither is the final guarantee of the complete subtree | | `population` | The population whose membership and values are maintained | -For example, one KLL `SummaryAgg` can feed two `SummaryEstimate` nodes whose queries -request p50 and p99. The build operation and its state are shared; the requested -estimates differ. - Schema, derived accuracy and execution timing describe every `Operator`, regardless of category (§2). They are omitted from these operation-specific sketches. Timing comes from lifecycle planning rather than a fixed field value implied by an operator @@ -257,30 +253,12 @@ describe how an operator processes its input. This prevents an expression from being mistaken for a table-producing plan. The [companion proposal](decoupling_op_and_expr.md) defines this distinction. -### 1.3 Two meanings of sharing - -**Sharing the operator model** means pre-ASAP and post-ASAP use the same definitions -for ordinary operators. It does not mean those two planning stages execute together -or must reference the same node instances. - -**Sharing a computation** means multiple consumers within a workload use one -producer. For example, two queries may read one scan, or two estimates may use one -summary: - -```text - ┌→ p50 estimation → query A -Scan → KLL summary build ┤ - └→ p99 estimation → query B -``` - -The shared producer must satisfy every consumer's input, window, accuracy and timing -requirements. Representing it once exposes reuse to planning and costing. Keeping -multiple query roots in one workload graph is therefore part of the design. +### 1.3 Scope of operator sharing -Adding accuracy information or execution timing must preserve that sharing. It must -also leave alternative candidate plans independent: assigning a lifecycle to one -candidate must not change another candidate's choices. How nodes are stored or -reused is outside this design. +Here, sharing means pre-ASAP and post-ASAP use the same operator definitions. +A `Project`, for example, has one representation whether its input is an ordinary +aggregate or a summary estimate. This proposal removes the representation boundary; +it does not introduce rules for sharing computations across queries. ## 2. Node properties and why they differ @@ -457,8 +435,8 @@ The design is successful when: - A projection uses the same semantics above and below summary computations. - An exact aggregate and a summary can share an input when all requirements agree. - A union or another ordinary operator can consume summary estimates on its inputs. -- Shared producers remain shared when accuracy and timing are resolved, including - across different queries in the workload. +- Unifying the representation preserves existing graph dependencies, including + any shared inputs; it does not introduce new sharing rules. - Invalid value/state combinations, incompatible timing and unfinished plan assessments are rejected before execution. - Export preserves visible dependencies and shared producers. From a45fcbb4c2dcc2a5a49fbe098d57926e75da0617 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:33:46 +0000 Subject: [PATCH 24/46] docs: remove proposed exact operation rule field --- .../design_docs/proposals/operator-sharing.md | 27 +++++++++---------- 1 file changed, 13 insertions(+), 14 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index a72bfe98f..d3fe7c0ba 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -97,7 +97,7 @@ the existing query and summary models, but the sketches combine several changes: | Give both categories inputs that refer to `Operator` nodes | Allow ordinary and summary operations to compose directly, without a wrapper hiding their dependencies. | | Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | | Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | -| Introduce `local_guarantee` and `exact_operation_rule` on summary operations | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). These are not fields on today's corresponding summary operations. | +| Introduce `local_guarantee` on summary estimation | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). This field does not exist on today's summary estimation operation. | | Let physical/lifecycle planning decide when each node executes | Additional planning design (§2.3): logical construction leaves timing undecided; planning chooses ingestion-time or query-time execution under the workload constraints. | **Naming and compatibility.** The sketches retain current operation, field and @@ -118,10 +118,9 @@ The following changes are explicit: - `Limit.partition_by` comes from the existing post-ASAP limit. Applying that field to the unified operator remains an explicit design choice, not new functionality implied by a rename. `SQLWindowFunc.frame` retains its current optional form. -- `SummaryEstimate.local_guarantee` and `SummaryAgg.exact_operation_rule` are proposed - new fields, named after the existing accuracy-model concepts. Their types remain - `ResultGuarantee` and `CompositionOperator`; neither field exists on today's - corresponding summary operation (§2.2). +- `SummaryEstimate.local_guarantee` is a proposed new field, named after the existing + accuracy-model concept and using `ResultGuarantee`. It does not exist on today's + summary estimation operation (§2.2). **Proposed data structures.** These sketches describe operation-specific data using current names. `Rc` represents a shared input edge; no new reference type @@ -192,7 +191,6 @@ enum ASAPOp { SummaryAgg { child: Rc, family: SummaryFamilyType, input: SummaryUpdate, reduction: Reduction, grouping: GroupingStrategy, - exact_operation_rule: Option, // proposed new field }, SummaryEstimate { summary_input: Rc, query: SketchQuery, @@ -222,7 +220,7 @@ The summary fields distinguish state construction, readout and accuracy evidence | `reduction` | Which input entities contribute to each logical result | | `grouping` | Whether those groups use separate state instances or a supported shared structure | | `query` / `readout` | The result requested from summary or maintained-population state | -| `local_guarantee` / `exact_operation_rule` | Local accuracy evidence or composition semantics; neither is the final guarantee of the complete subtree | +| `local_guarantee` | Local accuracy evidence, not the final guarantee of the complete subtree | | `population` | The population whose membership and values are maintained | Schema, derived accuracy and execution timing describe every `Operator`, regardless @@ -291,19 +289,20 @@ not the guarantee of a larger query: its input may already be approximate, and later operations may change the error. The planner must compose accuracy through the actual computation graph. -**Existing concepts versus proposed fields.** The current code already calculates -local guarantees and uses accuracy-composition rules, but neither `local_guarantee` -nor `exact_operation_rule` is a field on the corresponding summary operation today: +**Existing concept versus proposed field.** The current code already calculates +local guarantees, but `local_guarantee` is not a field on summary estimation today: | Proposed field | Meaning | Current code | What this proposal changes | |---|---|---|---| | `SummaryEstimate.local_guarantee` | The guarantee for this summary readout over exact input; it excludes upstream error. | `AccuracyModel.local_guarantee(...)` already computes it. It is used when composing the result guarantee, rather than retained on the estimation operation. | Retain that local evidence so accuracy can be derived from the assembled graph's actual inputs. | -| `SummaryAgg.exact_operation_rule` | The rule for propagating input error through an exact aggregate. “Exact” describes the operation, not a promise that approximate inputs become exact. | Composition rules already exist. Exact-aggregate binding selects a rule from the aggregate's registered semantics; the accuracy model also exposes `exact_operation_rule(...)` for exact operations. Neither is a stored field on `SummaryAgg`. | Retain the selected rule on the operation so later derivation does not need to recover the original binding context. | Today, the composed result is stored as the summary node's `guarantee`. The proposed -fields retain inputs to that calculation; they do not replace the complete result -guarantee. Adding them is an accuracy-design change, not a requirement of dividing -operators into `NonASAP` and `ASAP`. +`local_guarantee` field retains one input to that calculation; it does not replace +the complete result guarantee. Adding it is an accuracy-design change, not a +requirement of dividing operators into `NonASAP` and `ASAP`. + +Exact aggregates continue to use the existing accuracy-composition rules. This +proposal does not add a field to store those rules on the operator. For example, a KLL estimate over exact input can carry the sketch's guarantee. If that input is approximate, the estimate must also account for the upstream error. From 244af7f7499bc2d97ae67fcc7fd9c3196cd0d2db Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:37:14 +0000 Subject: [PATCH 25/46] docs: limit operator proposal to representation changes --- .../design_docs/proposals/operator-sharing.md | 217 +++++------------- 1 file changed, 62 insertions(+), 155 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index d3fe7c0ba..71ef9414f 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -34,9 +34,8 @@ Post-ASAP projection Project Today the two projections need separate representations, and the scan is hidden inside the wrapped subplan. In the proposed graph, both projections use the same -operator definition and the scan is directly visible. An exact aggregate can also -consume that scan when sharing is valid, as shown in §4. Similarly, a union can -consume summary estimates without needing a separate post-ASAP union definition. +operator definition and the scan is directly visible. A union can likewise consume +summary estimates without needing a separate post-ASAP union definition. The design removes these representation barriers. It makes composition and sharing possible; whether a particular rewrite or shared computation is valid still depends @@ -69,7 +68,8 @@ operator kinds. `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` are reserved in the proposed planner model; listing them does not establish planner -or runtime support. Summary composition remains an open design question (§7). +or runtime support. Defining new summary-composition semantics is outside this +proposal (§6). `NonASAPOp` and `ASAPOp` describe the operation performed by a node. Inputs in both categories connect to `Operator` nodes, so either category can consume the other @@ -97,8 +97,7 @@ the existing query and summary models, but the sketches combine several changes: | Give both categories inputs that refer to `Operator` nodes | Allow ordinary and summary operations to compose directly, without a wrapper hiding their dependencies. | | Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | | Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | -| Introduce `local_guarantee` on summary estimation | Additional accuracy design: retain local evidence separately from the derived subtree guarantee (§2.2). This field does not exist on today's summary estimation operation. | -| Let physical/lifecycle planning decide when each node executes | Additional planning design (§2.3): logical construction leaves timing undecided; planning chooses ingestion-time or query-time execution under the workload constraints. | +| Represent execution timing chosen during physical planning | Follow the planning-stage design in [#509](https://github.com/ProjectASAP/ASAPPlanner/pull/509), rather than introduce a new timing policy here (§2.3). | **Naming and compatibility.** The sketches retain current operation, field and payload-type names. `NonASAPOp`, `ASAPOp` and the companion proposal's `ScalarExpr` @@ -118,9 +117,6 @@ The following changes are explicit: - `Limit.partition_by` comes from the existing post-ASAP limit. Applying that field to the unified operator remains an explicit design choice, not new functionality implied by a rename. `SQLWindowFunc.frame` retains its current optional form. -- `SummaryEstimate.local_guarantee` is a proposed new field, named after the existing - accuracy-model concept and using `ResultGuarantee`. It does not exist on today's - summary estimation operation (§2.2). **Proposed data structures.** These sketches describe operation-specific data using current names. `Rc` represents a shared input edge; no new reference type @@ -194,7 +190,6 @@ enum ASAPOp { }, SummaryEstimate { summary_input: Rc, query: SketchQuery, - local_guarantee: Option, // proposed new field }, FinalizeExactAccumulator { child: Rc }, MaintainPopulation { child: Rc, population: MaintainedPopulation }, @@ -211,7 +206,7 @@ enum ASAPOp { } ``` -The summary fields distinguish state construction, readout and accuracy evidence: +The summary fields distinguish state construction and readout: | Field | Design meaning | |---|---| @@ -220,7 +215,6 @@ The summary fields distinguish state construction, readout and accuracy evidence | `reduction` | Which input entities contribute to each logical result | | `grouping` | Whether those groups use separate state instances or a supported shared structure | | `query` / `readout` | The result requested from summary or maintained-population state | -| `local_guarantee` | Local accuracy evidence, not the final guarantee of the complete subtree | | `population` | The population whose membership and values are maintained | Schema, derived accuracy and execution timing describe every `Operator`, regardless @@ -282,133 +276,54 @@ series identity where relevant. Sharing operators must not change SQL or PromQL meaning. After a rewrite, schemas must describe the new inputs rather than the plan that was replaced. -### 2.2 Accuracy follows the computation +### 2.2 Preserve existing accuracy semantics -A summary's local error describes its behavior over an exact input. That alone is -not the guarantee of a larger query: its input may already be approximate, and -later operations may change the error. The planner must compose accuracy through -the actual computation graph. +The unified representation must preserve the existing accuracy model, composition +rules and result guarantees. An operation's guarantee must still account for its +actual inputs; unknown accuracy must not be treated as exactness. -**Existing concept versus proposed field.** The current code already calculates -local guarantees, but `local_guarantee` is not a field on summary estimation today: +This proposal adds no accuracy fields or new guarantee-calculation workflow. +Changing the operator representation must not change the accuracy meaning of the +same computation. -| Proposed field | Meaning | Current code | What this proposal changes | -|---|---|---|---| -| `SummaryEstimate.local_guarantee` | The guarantee for this summary readout over exact input; it excludes upstream error. | `AccuracyModel.local_guarantee(...)` already computes it. It is used when composing the result guarantee, rather than retained on the estimation operation. | Retain that local evidence so accuracy can be derived from the assembled graph's actual inputs. | +### 2.3 Timing follows the planning-stage design -Today, the composed result is stored as the summary node's `guarantee`. The proposed -`local_guarantee` field retains one input to that calculation; it does not replace -the complete result guarantee. Adding it is an accuracy-design change, not a -requirement of dividing operators into `NonASAP` and `ASAP`. +The [planning-stage design in #509](https://github.com/ProjectASAP/ASAPPlanner/pull/509) +separates logical decisions about what to compute from physical decisions about how +and when to compute it. This proposal follows that division. -Exact aggregates continue to use the existing accuracy-composition rules. This -proposal does not add a field to store those rules on the operator. +For example, a KLL summary build may execute at ingestion time or query time, +depending on the materialization choice. The unified operator representation must +carry the chosen execution timing without requiring separate operator definitions +for the two phases. -For example, a KLL estimate over exact input can carry the sketch's guarantee. If -that input is approximate, the estimate must also account for the upstream error. -The summary state itself is not a query answer and need not have a value-level -accuracy guarantee. - -Ordinary exact computations remain exact when their inputs and operation semantics -justify it. Exact accumulator finalization preserves the established guarantee. -Special cases, such as exact counting of the rows actually received, retain their -operation-specific rules. - -Distinguish an assessment that has not happened from an assessment that found no -supported guarantee. Missing accuracy evidence cannot be treated as exactness or -as proof that a query's accuracy target is met. Candidate assessment and final-plan -validation must use the same accuracy semantics. - -### 2.3 Timing is a planning choice - -**Who decides when a node executes?** This concerns responsibility for choosing -an execution phase, not memory ownership or ownership of the data. - -| Current behavior | Proposed behavior | -|---|---| -| Some operations receive timing during construction; other timings are inferred from their inputs or consuming edges. | Logical construction describes what to compute. Physical/lifecycle planning explicitly chooses when it runs, using the complete workload and execution constraints. | - -For example, a KLL summary may be maintained at ingestion time or built at query -time. The presence of a summary-build node does not itself choose either option. -The planner compares those alternatives; the selected lifecycle determines timing -for the build and its dependencies. Operator-specific restrictions still apply, -such as summary estimation running at query time. - -This changes where the timing decision is made. It is a separate planning change, -not an automatic consequence of putting operators into two categories. - -The same logical summary can be maintained as data arrives or computed when a query -needs it. Its position in the graph alone does not choose between these behaviors. -Logical planning therefore leaves timing undecided. Physical planning chooses -materialization and lifecycle behavior, then determines execution phases across the -complete graph. - -The planner evaluates alternatives using deployment-provided cost and accuracy -models and capabilities. The deployment executes the selected plan, consistent with -the [planning-stages proposal](https://github.com/ProjectASAP/ASAPPlanner/pull/509). - -A valid timing assignment must satisfy these constraints: - -- Ingestion-time work cannot depend on a query-time result. -- Summary estimation runs at query time. Population maintenance and readout run at - ingestion time and query time, respectively. -- A shared computation has one execution phase compatible with all its consumers. - A single producer cannot simultaneously mean two separate executions. -- Stored state remains available for as long as its consumers need it. - -This proposal covers materialization at summary-state boundaries. Materializing -arbitrary ordinary intermediate results requires a separate design. +The representation must preserve the resulting execution constraints: ingestion-time +work cannot depend on query-time results, and consumers must receive values or state +that are available when needed. Materialization choices, retention and plan selection +remain governed by #509; this document does not define another lifecycle policy. ## 3. Planning responsibilities -The common representation separates what a plan computes from how it executes: +These are the stages defined in +[#509](https://github.com/ProjectASAP/ASAPPlanner/pull/509), shown here only to explain +how they use the common operator model: -| Responsibility | Required result | +| Stage from #509 | Use of the unified representation | |---|---| -| Frontend translation | An ordinary query graph preserving source-language semantics | -| Logical optimization | Exact and summary-based alternatives, including legal shared computations | -| Candidate assessment | Accuracy and capability evidence for the actual candidate graph | -| Physical and lifecycle planning | Executable alternatives with materialization, retention and compatible timing | -| Plan selection | A valid plan chosen using workload-level costs and requirements | -| Export and execution | The selected graph with explicit dependencies and completed assessments | +| Frontends | Produce a graph containing only `NonASAP` operators, preserving source-language semantics. | +| Logical ASAP-aware optimization | Form candidate graphs containing ordinary and summary operators, with no wrappers hiding their dependencies. | +| Physical ASAP-aware optimization | Determine executable alternatives, including materialization and execution timing, for those candidate graphs. | +| Plan selection | Evaluate complete physical candidates using workload requirements and deployment-provided models and capabilities. | +| Deployment execution | Execute the selected graph, preserving its dependencies and assigned phases. | -Ordinary operators remain in the graph when their inputs are replaced by summary -computations. For example, replacing an aggregate below a projection must not require -replacing the projection with a separate post-ASAP operator. +Ordinary operators remain the same operations when their inputs are replaced by +summary computations. Replacing an aggregate below a projection, for example, does +not require a separate post-ASAP projection definition. -Changing a candidate's inputs can change its schema, accuracy and legal timing. -Those properties must be checked against the resulting graph. A plan with unresolved -execution timing or an unfinished accuracy assessment is not ready for export. +This document changes the representation used by these stages, not their search, +accuracy, costing or selection policies. -This proposal changes the representation, not the search strategy. It does not -require a new enumeration algorithm or change when candidates commit to particular -child plans. - -## 4. Example: one scan serving exact and approximate queries - -Consider an exact average and an approximate p99 over the same input interval. -Once a rewrite exposes the two computations, the common model can express: - -```text - ┌→ Exact average ─────────────────────→ query A -Scan ──┤ - └→ KLL summary build → p99 estimation → query B -``` - -If both branches run at query time, they may share the scan. Computing accuracy and -assigning timing must retain that one producer for both consumers. - -If the KLL is maintained at ingestion time while the exact average reads raw data at -query time, the depicted scan cannot serve both phases as one execution. The plan -needs separate scans or a different lifecycle that makes sharing valid. The planner -must represent that choice explicitly; timing validation rejects the conflicting -shared plan. - -The representation enables the shared plan but does not supply the rewrite that -splits a multi-measure aggregate into these branches. That rewrite is outside this -proposal. How planning resolves the mixed-phase case remains open (§7). - -## 5. Export preserves the graph +## 4. Export preserves the graph Export one node per operator and represent its input dependencies as edges. Export a shared producer once, with edges to all its consumers. @@ -423,43 +338,35 @@ operations, but must preserve its dependencies and meaning. The execution layer does not invent missing planning decisions. Changing the exported representation requires coordinated adoption by the planner -and downstream readers. Representation unification must preserve query semantics; -it does not by itself guarantee unchanged timings for plans with unresolved sharing -conflicts. +and downstream readers while preserving existing query semantics and the selected +plan's execution requirements. -## 6. Acceptance criteria +## 5. Acceptance criteria The design is successful when: - A projection uses the same semantics above and below summary computations. -- An exact aggregate and a summary can share an input when all requirements agree. - A union or another ordinary operator can consume summary estimates on its inputs. - Unifying the representation preserves existing graph dependencies, including any shared inputs; it does not introduce new sharing rules. -- Invalid value/state combinations, incompatible timing and unfinished plan - assessments are rejected before execution. +- Existing value/state, accuracy and execution constraints remain enforceable on + the unified representation. - Export preserves visible dependencies and shared producers. -## 7. Scope and open questions - -This proposal defines a common operator model and its correctness constraints. It includes the operation-specific data needed to express those semantics, but -does not prescribe storage structures, public APIs, traversal algorithms, -serialization fields or a code migration sequence. - -The following decisions remain separate or unresolved: - -- **Mixed-phase sharing:** when consumers need different phases, should planning - separate the producer or find a common materialized lifecycle? Existing timing - behavior cannot be promised until this is resolved. -- **Summary composition:** where are merge operations introduced—during planner - optimization or downstream? Accuracy rules for merge, subtract, delete and - summary joins need a separate design. A representable operation is not a claim of - runtime support. -- **Lifecycle integration:** reuse the lifecycle concepts being developed in - [#482](https://github.com/ProjectASAP/ASAPPlanner/pull/482), without creating a second - competing source of timing decisions. -- **Additional optimization:** rules that split exact and approximate measures, - new pruning strategies and deferred child-plan choices are outside this proposal. -- **Finalization:** this representation change does not resolve the existing - difference between query-result assembly and lifecycle-plan assembly in adding - finalization of exact aggregate state. +## 6. Scope and compatibility + +This proposal defines the common operator structure, its operation-specific data +and how dependencies remain visible through export. Existing names and semantics +are retained except for the structural changes identified in §1.1. + +The pre-ASAP and post-ASAP versions of some operations carry different information. +The unified `BinaryOp` must retain existing checked-division requirements, and +`Limit` must retain the existing ability to limit within groups. These compatibility +requirements belong in this design because removing duplicate operator definitions +must not remove existing behavior. + +New accuracy fields, accuracy-composition rules, computation-sharing algorithms and +lifecycle policies are outside this proposal. Planning responsibilities follow #509. +The scalar/operator separation is specified in the companion document. Storage, +traversal algorithms, serialization fields and a code migration sequence are also +outside this document. From 863d8e3da510ce6e820d17d1886d654b5502ea2a Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:41:33 +0000 Subject: [PATCH 26/46] docs: restore concrete scalar and operator split structures --- .../proposals/decoupling_op_and_expr.md | 156 +++++++++++++++--- .../design_docs/proposals/operator-sharing.md | 4 +- 2 files changed, 131 insertions(+), 29 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index d48af6c7f..aad6d4063 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -10,7 +10,6 @@ expressions. Their roles are distinguished by where they occur, so an expression can be placed where a table input is expected and fail only when the plan is checked. Separate these concepts so the plan model expresses which combinations are valid. -This also makes clear what the planner can replace or share as a computation. Consider: @@ -31,25 +30,134 @@ Scan lineitem → Filter → Project The scan, filter and projection produce tables. The predicate and multiplication compute values within the schema selected by their owning operators. -## 2. Design and rationale +## 2. Proposed data structures -| Concept | Meaning | Role in the plan | +Split `QueryExpr` into `NonASAPOp` and `ScalarExpr`. Keep existing variant names, +field names and semantics; change their types to express their roles: + +| Position | Current type | Proposed type | |---|---|---| -| Operator | Produces a table, or summary state in the companion design | A graph node whose input computations may be replaced or shared | -| Scalar expression | Computes a value in a particular schema context | Part of an operator's predicate, projection, sort key or other expression | +| Operator input | `QueryExpr` | `NonASAPOp` | +| Scalar expression | `QueryExpr` | `ScalarExpr` | + +The sketches retain existing `Rc` and collection shapes where possible; this +proposal does not redesign storage. `C` remains the existing column representation: +`ColumnRef` before resolution and `ColumnId` afterward. `C::ScanSchema` retains the +corresponding unresolved or resolved scan schema. + +### 2.1 Operator nodes + +`NonASAPOp` contains the current operator variants. Inputs reference other operators; +predicates and expression-bearing fields use the scalar structures below. + +```rust +enum NonASAPOp { + Scan { + source: Source, predicates: Vec>, schema: C::ScanSchema, + }, + Filter { pred: Predicate, child: Rc> }, + Project { + cols: Vec>, qualifier: Option, child: Rc>, + }, + Aggregate { + reduction: Reduction, measures: Vec>, output_names: Vec, + having: Option>, child: Rc>, + }, + Dedup { cols: Vec, child: Rc> }, + Concat { + children: Vec>, + discriminator_unique_key: Option>, + }, + Join { + kind: JoinKind, pred: Predicate, + left: Rc>, right: Rc>, + }, + SetOp { + kind: RelationalSetOpKind, all: bool, + left: Rc>, right: Rc>, + }, + Sort { keys: Vec>, partition_by: GroupKeys, child: Rc> }, + Limit { n: usize, offset: usize, child: Rc> }, + BinaryOp { + op: BinaryOpKind, lhs: Rc>, rhs: Rc>, + vector_match: Option, + }, + SQLWindowFunc { + func: WindowFuncKind, args: Vec>, partition_by: GroupKeys, + order_by: Vec>, frame: Option, + output_name: String, child: Rc>, + }, + TimeRange { range: Duration, child: Rc> }, + TimeShift { shift: TimeShift, child: Rc> }, + PromqlScalarBridge(Rc>), + EvalTimestamp, + PromqlVectorFromScalar(Rc>), + PromqlScalarFromVector(Rc>), + PromqlRelabel { dst: String, value: Rc>, child: Rc> }, + PromqlInfoEnrich { selector: Vec, child: Rc> }, + PromqlSeriesSample { by: GroupKeys, kind: SampleKind, child: Rc> }, + PromqlSubquery { + range: Duration, resolution: Option, child: Rc>, + }, +} +``` -Operator inputs must be other operators. Scalar expressions may contain other -scalar expressions, but do not contain operator subplans in this design. +This is the scalar/operator split alone. It retains the current pre-ASAP `BinaryOp`, +`Limit` and `Concat` shapes. The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) +separately widens operator inputs to the common `Operator` and reconciles differences +between pre-ASAP and post-ASAP operations. + +### 2.2 Scalar expressions and their owning fields + +`ScalarExpr` contains every current scalar variant, including `CurrentTimestamp`. +Its recursive inputs are scalar expressions only. + +```rust +enum ScalarExpr { + Column(C), + Literal(ScalarValue), + Compare { left: Rc>, op: CompareOpKind, right: Rc> }, + BoolAnd(Vec>), + BoolOr(Vec>), + Not(Rc>), + IsNull(Rc>), + IsNotNull(Rc>), + Cast { expr: Rc>, to: DataType, try_cast: bool }, + InList { expr: Rc>, list: Vec>, negated: bool }, + FunctionCall { name: String, args: Vec> }, + Arithmetic { + op: ArithmeticOpKind, left: Rc>, right: Rc>, + }, + Case { + operand: Option>>, + branches: Vec<(ScalarExpr, ScalarExpr)>, + else_expr: Option>>, + }, + CurrentTimestamp, +} + +struct Predicate(Rc>); + +struct ProjectItem { + alias: Option, + expr: ScalarExpr, +} + +struct SortKey { + expr: ScalarExpr, + ascending: bool, + nulls_first: bool, +} +``` -A column reference has meaning only in its schema context. Two identical-looking -expressions in different operators may refer to different inputs. Keeping scalar -expressions attached to their operators preserves that context; this proposal does -not introduce independent shared scalar computations. +`Predicate`, `ProjectItem` and `SortKey` keep their current names and roles. Only the +expression type changes. A column reference remains meaningful in the schema +selected by its owning operator; it is not an independent table input. -The distinction depends on semantics, not on whether a result looks scalar. For -example, a PromQL conversion between a scalar and a vector participates in the -operator graph because it has query-level output semantics. SQL `NOW()` is an -expression evaluated within its owning operator. +`PromqlScalarBridge`, `PromqlVectorFromScalar`, `PromqlScalarFromVector` and +`EvalTimestamp` remain operators because they participate in query-level evaluation. +`CurrentTimestamp` (SQL `NOW()`) belongs to scalar expressions. Classifying these by +their role preserves their existing semantics. ## 3. Semantic requirements @@ -64,23 +172,17 @@ Splitting the representation must preserve evaluation behavior, inferred output types and source-language semantics. A scalar expression cannot serve as a table input, and a table-producing operator cannot appear where a scalar is expected. -The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) -extends the operator graph with summary operations. The scalar/operator distinction -continues to hold before and after that optimization: replacing a projection's input -with a summary estimate does not turn its scalar expressions into graph nodes. - ## 4. Acceptance and scope -The example query must retain its result and output schema. Planning can replace or -share its table-producing computations while interpreting the filter predicate and -projection expression in the correct contexts. Invalid scalar/table combinations -must be excluded by the plan model. +The example query must retain its result and output schema. The filter predicate +and projection expression must resolve in the same contexts as before. Invalid +scalar/table combinations must be excluded by the plan model. This separation alone does not require a change to the external plan format. The companion proposal addresses the separate decision to expose every operator in the exported graph. Scalar subqueries are outside this design. Filter `IN (SELECT …)` and `EXISTS` can -be represented as joins, but a general scalar subquery introduces a dependency on -another operator graph. Supporting that requires a separate design for its scope, -dependencies and participation in optimization. +be represented as joins; a general scalar subquery would require expressions to +reference operator graphs and needs a separate design. This proposal adds no new +optimization, accuracy or execution-timing behavior. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 71ef9414f..2b8530973 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -153,11 +153,11 @@ enum NonASAPOp { }, TimeRange { child: Rc, range: Duration }, TimeShift { child: Rc, shift: TimeShift }, - PromqlScalarBridge(ScalarExpr), + PromqlScalarBridge(Rc), EvalTimestamp, PromqlVectorFromScalar(Rc), PromqlScalarFromVector(Rc), - PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, + PromqlRelabel { child: Rc, dst: String, value: Rc }, PromqlInfoEnrich { child: Rc, selector: Vec }, PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, PromqlSubquery { child: Rc, range: Duration, resolution: Option }, From 43600f9f54e82b07ff3a96d55f3877326722e960 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:42:38 +0000 Subject: [PATCH 27/46] Revert "docs: restore concrete scalar and operator split structures" This reverts commit 863d8e3da510ce6e820d17d1886d654b5502ea2a. --- .../proposals/decoupling_op_and_expr.md | 156 +++--------------- .../design_docs/proposals/operator-sharing.md | 4 +- 2 files changed, 29 insertions(+), 131 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index aad6d4063..d48af6c7f 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -10,6 +10,7 @@ expressions. Their roles are distinguished by where they occur, so an expression can be placed where a table input is expected and fail only when the plan is checked. Separate these concepts so the plan model expresses which combinations are valid. +This also makes clear what the planner can replace or share as a computation. Consider: @@ -30,134 +31,25 @@ Scan lineitem → Filter → Project The scan, filter and projection produce tables. The predicate and multiplication compute values within the schema selected by their owning operators. -## 2. Proposed data structures +## 2. Design and rationale -Split `QueryExpr` into `NonASAPOp` and `ScalarExpr`. Keep existing variant names, -field names and semantics; change their types to express their roles: - -| Position | Current type | Proposed type | +| Concept | Meaning | Role in the plan | |---|---|---| -| Operator input | `QueryExpr` | `NonASAPOp` | -| Scalar expression | `QueryExpr` | `ScalarExpr` | - -The sketches retain existing `Rc` and collection shapes where possible; this -proposal does not redesign storage. `C` remains the existing column representation: -`ColumnRef` before resolution and `ColumnId` afterward. `C::ScanSchema` retains the -corresponding unresolved or resolved scan schema. - -### 2.1 Operator nodes - -`NonASAPOp` contains the current operator variants. Inputs reference other operators; -predicates and expression-bearing fields use the scalar structures below. - -```rust -enum NonASAPOp { - Scan { - source: Source, predicates: Vec>, schema: C::ScanSchema, - }, - Filter { pred: Predicate, child: Rc> }, - Project { - cols: Vec>, qualifier: Option, child: Rc>, - }, - Aggregate { - reduction: Reduction, measures: Vec>, output_names: Vec, - having: Option>, child: Rc>, - }, - Dedup { cols: Vec, child: Rc> }, - Concat { - children: Vec>, - discriminator_unique_key: Option>, - }, - Join { - kind: JoinKind, pred: Predicate, - left: Rc>, right: Rc>, - }, - SetOp { - kind: RelationalSetOpKind, all: bool, - left: Rc>, right: Rc>, - }, - Sort { keys: Vec>, partition_by: GroupKeys, child: Rc> }, - Limit { n: usize, offset: usize, child: Rc> }, - BinaryOp { - op: BinaryOpKind, lhs: Rc>, rhs: Rc>, - vector_match: Option, - }, - SQLWindowFunc { - func: WindowFuncKind, args: Vec>, partition_by: GroupKeys, - order_by: Vec>, frame: Option, - output_name: String, child: Rc>, - }, - TimeRange { range: Duration, child: Rc> }, - TimeShift { shift: TimeShift, child: Rc> }, - PromqlScalarBridge(Rc>), - EvalTimestamp, - PromqlVectorFromScalar(Rc>), - PromqlScalarFromVector(Rc>), - PromqlRelabel { dst: String, value: Rc>, child: Rc> }, - PromqlInfoEnrich { selector: Vec, child: Rc> }, - PromqlSeriesSample { by: GroupKeys, kind: SampleKind, child: Rc> }, - PromqlSubquery { - range: Duration, resolution: Option, child: Rc>, - }, -} -``` +| Operator | Produces a table, or summary state in the companion design | A graph node whose input computations may be replaced or shared | +| Scalar expression | Computes a value in a particular schema context | Part of an operator's predicate, projection, sort key or other expression | -This is the scalar/operator split alone. It retains the current pre-ASAP `BinaryOp`, -`Limit` and `Concat` shapes. The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) -separately widens operator inputs to the common `Operator` and reconciles differences -between pre-ASAP and post-ASAP operations. - -### 2.2 Scalar expressions and their owning fields - -`ScalarExpr` contains every current scalar variant, including `CurrentTimestamp`. -Its recursive inputs are scalar expressions only. - -```rust -enum ScalarExpr { - Column(C), - Literal(ScalarValue), - Compare { left: Rc>, op: CompareOpKind, right: Rc> }, - BoolAnd(Vec>), - BoolOr(Vec>), - Not(Rc>), - IsNull(Rc>), - IsNotNull(Rc>), - Cast { expr: Rc>, to: DataType, try_cast: bool }, - InList { expr: Rc>, list: Vec>, negated: bool }, - FunctionCall { name: String, args: Vec> }, - Arithmetic { - op: ArithmeticOpKind, left: Rc>, right: Rc>, - }, - Case { - operand: Option>>, - branches: Vec<(ScalarExpr, ScalarExpr)>, - else_expr: Option>>, - }, - CurrentTimestamp, -} - -struct Predicate(Rc>); - -struct ProjectItem { - alias: Option, - expr: ScalarExpr, -} - -struct SortKey { - expr: ScalarExpr, - ascending: bool, - nulls_first: bool, -} -``` +Operator inputs must be other operators. Scalar expressions may contain other +scalar expressions, but do not contain operator subplans in this design. -`Predicate`, `ProjectItem` and `SortKey` keep their current names and roles. Only the -expression type changes. A column reference remains meaningful in the schema -selected by its owning operator; it is not an independent table input. +A column reference has meaning only in its schema context. Two identical-looking +expressions in different operators may refer to different inputs. Keeping scalar +expressions attached to their operators preserves that context; this proposal does +not introduce independent shared scalar computations. -`PromqlScalarBridge`, `PromqlVectorFromScalar`, `PromqlScalarFromVector` and -`EvalTimestamp` remain operators because they participate in query-level evaluation. -`CurrentTimestamp` (SQL `NOW()`) belongs to scalar expressions. Classifying these by -their role preserves their existing semantics. +The distinction depends on semantics, not on whether a result looks scalar. For +example, a PromQL conversion between a scalar and a vector participates in the +operator graph because it has query-level output semantics. SQL `NOW()` is an +expression evaluated within its owning operator. ## 3. Semantic requirements @@ -172,17 +64,23 @@ Splitting the representation must preserve evaluation behavior, inferred output types and source-language semantics. A scalar expression cannot serve as a table input, and a table-producing operator cannot appear where a scalar is expected. +The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) +extends the operator graph with summary operations. The scalar/operator distinction +continues to hold before and after that optimization: replacing a projection's input +with a summary estimate does not turn its scalar expressions into graph nodes. + ## 4. Acceptance and scope -The example query must retain its result and output schema. The filter predicate -and projection expression must resolve in the same contexts as before. Invalid -scalar/table combinations must be excluded by the plan model. +The example query must retain its result and output schema. Planning can replace or +share its table-producing computations while interpreting the filter predicate and +projection expression in the correct contexts. Invalid scalar/table combinations +must be excluded by the plan model. This separation alone does not require a change to the external plan format. The companion proposal addresses the separate decision to expose every operator in the exported graph. Scalar subqueries are outside this design. Filter `IN (SELECT …)` and `EXISTS` can -be represented as joins; a general scalar subquery would require expressions to -reference operator graphs and needs a separate design. This proposal adds no new -optimization, accuracy or execution-timing behavior. +be represented as joins, but a general scalar subquery introduces a dependency on +another operator graph. Supporting that requires a separate design for its scope, +dependencies and participation in optimization. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 2b8530973..71ef9414f 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -153,11 +153,11 @@ enum NonASAPOp { }, TimeRange { child: Rc, range: Duration }, TimeShift { child: Rc, shift: TimeShift }, - PromqlScalarBridge(Rc), + PromqlScalarBridge(ScalarExpr), EvalTimestamp, PromqlVectorFromScalar(Rc), PromqlScalarFromVector(Rc), - PromqlRelabel { child: Rc, dst: String, value: Rc }, + PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, PromqlInfoEnrich { child: Rc, selector: Vec }, PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, PromqlSubquery { child: Rc, range: Duration, resolution: Option }, From e3d95709bf67c5a987a6d80f7218e7c69a5ec384 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:43:37 +0000 Subject: [PATCH 28/46] Revert "Revert "docs: restore concrete scalar and operator split structures"" This reverts commit 43600f9f54e82b07ff3a96d55f3877326722e960. --- .../proposals/decoupling_op_and_expr.md | 156 +++++++++++++++--- .../design_docs/proposals/operator-sharing.md | 4 +- 2 files changed, 131 insertions(+), 29 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index d48af6c7f..aad6d4063 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -10,7 +10,6 @@ expressions. Their roles are distinguished by where they occur, so an expression can be placed where a table input is expected and fail only when the plan is checked. Separate these concepts so the plan model expresses which combinations are valid. -This also makes clear what the planner can replace or share as a computation. Consider: @@ -31,25 +30,134 @@ Scan lineitem → Filter → Project The scan, filter and projection produce tables. The predicate and multiplication compute values within the schema selected by their owning operators. -## 2. Design and rationale +## 2. Proposed data structures -| Concept | Meaning | Role in the plan | +Split `QueryExpr` into `NonASAPOp` and `ScalarExpr`. Keep existing variant names, +field names and semantics; change their types to express their roles: + +| Position | Current type | Proposed type | |---|---|---| -| Operator | Produces a table, or summary state in the companion design | A graph node whose input computations may be replaced or shared | -| Scalar expression | Computes a value in a particular schema context | Part of an operator's predicate, projection, sort key or other expression | +| Operator input | `QueryExpr` | `NonASAPOp` | +| Scalar expression | `QueryExpr` | `ScalarExpr` | + +The sketches retain existing `Rc` and collection shapes where possible; this +proposal does not redesign storage. `C` remains the existing column representation: +`ColumnRef` before resolution and `ColumnId` afterward. `C::ScanSchema` retains the +corresponding unresolved or resolved scan schema. + +### 2.1 Operator nodes + +`NonASAPOp` contains the current operator variants. Inputs reference other operators; +predicates and expression-bearing fields use the scalar structures below. + +```rust +enum NonASAPOp { + Scan { + source: Source, predicates: Vec>, schema: C::ScanSchema, + }, + Filter { pred: Predicate, child: Rc> }, + Project { + cols: Vec>, qualifier: Option, child: Rc>, + }, + Aggregate { + reduction: Reduction, measures: Vec>, output_names: Vec, + having: Option>, child: Rc>, + }, + Dedup { cols: Vec, child: Rc> }, + Concat { + children: Vec>, + discriminator_unique_key: Option>, + }, + Join { + kind: JoinKind, pred: Predicate, + left: Rc>, right: Rc>, + }, + SetOp { + kind: RelationalSetOpKind, all: bool, + left: Rc>, right: Rc>, + }, + Sort { keys: Vec>, partition_by: GroupKeys, child: Rc> }, + Limit { n: usize, offset: usize, child: Rc> }, + BinaryOp { + op: BinaryOpKind, lhs: Rc>, rhs: Rc>, + vector_match: Option, + }, + SQLWindowFunc { + func: WindowFuncKind, args: Vec>, partition_by: GroupKeys, + order_by: Vec>, frame: Option, + output_name: String, child: Rc>, + }, + TimeRange { range: Duration, child: Rc> }, + TimeShift { shift: TimeShift, child: Rc> }, + PromqlScalarBridge(Rc>), + EvalTimestamp, + PromqlVectorFromScalar(Rc>), + PromqlScalarFromVector(Rc>), + PromqlRelabel { dst: String, value: Rc>, child: Rc> }, + PromqlInfoEnrich { selector: Vec, child: Rc> }, + PromqlSeriesSample { by: GroupKeys, kind: SampleKind, child: Rc> }, + PromqlSubquery { + range: Duration, resolution: Option, child: Rc>, + }, +} +``` -Operator inputs must be other operators. Scalar expressions may contain other -scalar expressions, but do not contain operator subplans in this design. +This is the scalar/operator split alone. It retains the current pre-ASAP `BinaryOp`, +`Limit` and `Concat` shapes. The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) +separately widens operator inputs to the common `Operator` and reconciles differences +between pre-ASAP and post-ASAP operations. + +### 2.2 Scalar expressions and their owning fields + +`ScalarExpr` contains every current scalar variant, including `CurrentTimestamp`. +Its recursive inputs are scalar expressions only. + +```rust +enum ScalarExpr { + Column(C), + Literal(ScalarValue), + Compare { left: Rc>, op: CompareOpKind, right: Rc> }, + BoolAnd(Vec>), + BoolOr(Vec>), + Not(Rc>), + IsNull(Rc>), + IsNotNull(Rc>), + Cast { expr: Rc>, to: DataType, try_cast: bool }, + InList { expr: Rc>, list: Vec>, negated: bool }, + FunctionCall { name: String, args: Vec> }, + Arithmetic { + op: ArithmeticOpKind, left: Rc>, right: Rc>, + }, + Case { + operand: Option>>, + branches: Vec<(ScalarExpr, ScalarExpr)>, + else_expr: Option>>, + }, + CurrentTimestamp, +} + +struct Predicate(Rc>); + +struct ProjectItem { + alias: Option, + expr: ScalarExpr, +} + +struct SortKey { + expr: ScalarExpr, + ascending: bool, + nulls_first: bool, +} +``` -A column reference has meaning only in its schema context. Two identical-looking -expressions in different operators may refer to different inputs. Keeping scalar -expressions attached to their operators preserves that context; this proposal does -not introduce independent shared scalar computations. +`Predicate`, `ProjectItem` and `SortKey` keep their current names and roles. Only the +expression type changes. A column reference remains meaningful in the schema +selected by its owning operator; it is not an independent table input. -The distinction depends on semantics, not on whether a result looks scalar. For -example, a PromQL conversion between a scalar and a vector participates in the -operator graph because it has query-level output semantics. SQL `NOW()` is an -expression evaluated within its owning operator. +`PromqlScalarBridge`, `PromqlVectorFromScalar`, `PromqlScalarFromVector` and +`EvalTimestamp` remain operators because they participate in query-level evaluation. +`CurrentTimestamp` (SQL `NOW()`) belongs to scalar expressions. Classifying these by +their role preserves their existing semantics. ## 3. Semantic requirements @@ -64,23 +172,17 @@ Splitting the representation must preserve evaluation behavior, inferred output types and source-language semantics. A scalar expression cannot serve as a table input, and a table-producing operator cannot appear where a scalar is expected. -The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) -extends the operator graph with summary operations. The scalar/operator distinction -continues to hold before and after that optimization: replacing a projection's input -with a summary estimate does not turn its scalar expressions into graph nodes. - ## 4. Acceptance and scope -The example query must retain its result and output schema. Planning can replace or -share its table-producing computations while interpreting the filter predicate and -projection expression in the correct contexts. Invalid scalar/table combinations -must be excluded by the plan model. +The example query must retain its result and output schema. The filter predicate +and projection expression must resolve in the same contexts as before. Invalid +scalar/table combinations must be excluded by the plan model. This separation alone does not require a change to the external plan format. The companion proposal addresses the separate decision to expose every operator in the exported graph. Scalar subqueries are outside this design. Filter `IN (SELECT …)` and `EXISTS` can -be represented as joins, but a general scalar subquery introduces a dependency on -another operator graph. Supporting that requires a separate design for its scope, -dependencies and participation in optimization. +be represented as joins; a general scalar subquery would require expressions to +reference operator graphs and needs a separate design. This proposal adds no new +optimization, accuracy or execution-timing behavior. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 71ef9414f..2b8530973 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -153,11 +153,11 @@ enum NonASAPOp { }, TimeRange { child: Rc, range: Duration }, TimeShift { child: Rc, shift: TimeShift }, - PromqlScalarBridge(ScalarExpr), + PromqlScalarBridge(Rc), EvalTimestamp, PromqlVectorFromScalar(Rc), PromqlScalarFromVector(Rc), - PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, + PromqlRelabel { child: Rc, dst: String, value: Rc }, PromqlInfoEnrich { child: Rc, selector: Vec }, PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, PromqlSubquery { child: Rc, range: Duration, resolution: Option }, From fcfbe88de30ca0e1a0e6f609c387d42664067a83 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:44:33 +0000 Subject: [PATCH 29/46] docs: explain operators versus scalar expressions before the type split --- .../proposals/decoupling_op_and_expr.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index aad6d4063..1e24f1797 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -32,6 +32,23 @@ compute values within the schema selected by their owning operators. ## 2. Proposed data structures +An **operator** describes a query computation with its own output schema. It reads +from a source or other operators and produces a relation or time-series result. +For example, `Scan` produces input rows, `Filter` selects rows, and `Project` +produces a new set of columns. Operators form the nodes and input dependencies of +the query plan. + +A **scalar expression** describes a value calculated within the context of an +operator. For example, `l_quantity > 10` produces a Boolean used by a filter, and +`l_quantity * 2` produces a value for a projected column. An expression has a value +type, but no table output schema or independent data source. Its column references +are interpreted using the schema chosen by the owning operator. + +In the example above, `Filter.child` is the `Scan` operator that supplies rows; +`Filter.pred` is the expression `l_quantity > 10` evaluated on those rows. The +predicate cannot replace the scan as the filter's input, and the scan cannot be +used as its Boolean condition. This is the distinction the two types enforce. + Split `QueryExpr` into `NonASAPOp` and `ScalarExpr`. Keep existing variant names, field names and semantics; change their types to express their roles: From 6f13a8cb12263c27bfbc2885639be503b0ae08b0 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:46:34 +0000 Subject: [PATCH 30/46] docs: place operator field structures with NonASAPOp --- .../proposals/decoupling_op_and_expr.md | 43 +++++++++++-------- 1 file changed, 25 insertions(+), 18 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index 1e24f1797..4d8667ee5 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -119,12 +119,36 @@ enum NonASAPOp { } ``` +The following structures describe scalar-bearing fields of `NonASAPOp`: +`Predicate` is used by filters, joins and `HAVING`; `ProjectItem` by projections; +and `SortKey` by sorting and window functions. Their expressions use `ScalarExpr`, +defined in §2.2. + +```rust +struct Predicate(Rc>); + +struct ProjectItem { + alias: Option, + expr: ScalarExpr, +} + +struct SortKey { + expr: ScalarExpr, + ascending: bool, + nulls_first: bool, +} +``` + +`Predicate`, `ProjectItem` and `SortKey` keep their current names and roles. Only the +expression type changes. A column reference remains meaningful in the schema +selected by its owning operator; it is not an independent table input. + This is the scalar/operator split alone. It retains the current pre-ASAP `BinaryOp`, `Limit` and `Concat` shapes. The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) separately widens operator inputs to the common `Operator` and reconciles differences between pre-ASAP and post-ASAP operations. -### 2.2 Scalar expressions and their owning fields +### 2.2 Scalar expressions `ScalarExpr` contains every current scalar variant, including `CurrentTimestamp`. Its recursive inputs are scalar expressions only. @@ -152,25 +176,8 @@ enum ScalarExpr { }, CurrentTimestamp, } - -struct Predicate(Rc>); - -struct ProjectItem { - alias: Option, - expr: ScalarExpr, -} - -struct SortKey { - expr: ScalarExpr, - ascending: bool, - nulls_first: bool, -} ``` -`Predicate`, `ProjectItem` and `SortKey` keep their current names and roles. Only the -expression type changes. A column reference remains meaningful in the schema -selected by its owning operator; it is not an independent table input. - `PromqlScalarBridge`, `PromqlVectorFromScalar`, `PromqlScalarFromVector` and `EvalTimestamp` remain operators because they participate in query-level evaluation. `CurrentTimestamp` (SQL `NOW()`) belongs to scalar expressions. Classifying these by From 0064b8a7274d341c2ede6bed70ec1ddb2f815aed Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:49:27 +0000 Subject: [PATCH 31/46] docs: make operator and scalar ownership consistent --- .../proposals/decoupling_op_and_expr.md | 96 +++++++++++++------ .../design_docs/proposals/operator-sharing.md | 10 +- 2 files changed, 73 insertions(+), 33 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index 4d8667ee5..2c6348c1b 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -57,10 +57,23 @@ field names and semantics; change their types to express their roles: | Operator input | `QueryExpr` | `NonASAPOp` | | Scalar expression | `QueryExpr` | `ScalarExpr` | -The sketches retain existing `Rc` and collection shapes where possible; this -proposal does not redesign storage. `C` remains the existing column representation: -`ColumnRef` before resolution and `ColumnId` afterward. `C::ScanSchema` retains the -corresponding unresolved or resolved scan schema. +The split also makes ownership consistent with each structure's role: + +| Structure | Proposed rule | Reason | +|---|---|---| +| Operator inputs | Shared references, including every `Concat` child | All inputs are graph nodes; a multi-input operator should not embed copies while other operators reference nodes. | +| Scalar fields on operators | Owned `ScalarExpr` values, including predicates, relabel values and scalar bridges | These expressions belong to the operator's schema context and need no independent graph identity. | +| Scalar recursion | `Box` for individual recursive children; `Vec` for lists | Scalars form owned expression trees. `Box` gives recursive single-child fields finite size; `Vec` already provides that indirection for lists. | +| Column parameter | Keep `C: ColState` on operators; use plain `C` on scalar expressions and their wrappers | Operators need `C::ScanSchema`; scalar expressions only carry column references. | + +`C` continues to represent `ColumnRef` before resolution and `ColumnId` afterward. +No new column trait is introduced. Required cloning, comparison or serialization +bounds belong on the corresponding implementations, not on scalar data structures +through the unrelated scan-schema requirement. + +These are deliberate changes from the current mixed `Rc`/value representation. +They preserve operation names and evaluation semantics; they do not add scalar +sharing or a new optimization policy. ### 2.1 Operator nodes @@ -82,7 +95,7 @@ enum NonASAPOp { }, Dedup { cols: Vec, child: Rc> }, Concat { - children: Vec>, + children: Vec>>, discriminator_unique_key: Option>, }, Join { @@ -106,11 +119,11 @@ enum NonASAPOp { }, TimeRange { range: Duration, child: Rc> }, TimeShift { shift: TimeShift, child: Rc> }, - PromqlScalarBridge(Rc>), + PromqlScalarBridge(ScalarExpr), EvalTimestamp, PromqlVectorFromScalar(Rc>), PromqlScalarFromVector(Rc>), - PromqlRelabel { dst: String, value: Rc>, child: Rc> }, + PromqlRelabel { dst: String, value: ScalarExpr, child: Rc> }, PromqlInfoEnrich { selector: Vec, child: Rc> }, PromqlSeriesSample { by: GroupKeys, kind: SampleKind, child: Rc> }, PromqlSubquery { @@ -125,54 +138,58 @@ and `SortKey` by sorting and window functions. Their expressions use `ScalarExpr defined in §2.2. ```rust -struct Predicate(Rc>); +struct Predicate(ScalarExpr); -struct ProjectItem { +struct ProjectItem { alias: Option, expr: ScalarExpr, } -struct SortKey { +struct SortKey { expr: ScalarExpr, ascending: bool, nulls_first: bool, } ``` -`Predicate`, `ProjectItem` and `SortKey` keep their current names and roles. Only the -expression type changes. A column reference remains meaningful in the schema -selected by its owning operator; it is not an independent table input. +`Predicate`, `ProjectItem` and `SortKey` retain their names and roles and consistently +own their scalar expressions. `Predicate` marks a Boolean-expression position; +`ProjectItem` adds an alias, and `SortKey` adds ordering and null-placement rules. +These are different operator requirements, so the wrappers remain separate. -This is the scalar/operator split alone. It retains the current pre-ASAP `BinaryOp`, -`Limit` and `Concat` shapes. The [operator-sharing proposal](operator-sharing.md#11-unified-operator-type) -separately widens operator inputs to the common `Operator` and reconciles differences -between pre-ASAP and post-ASAP operations. +The pre-ASAP `BinaryOp` and `Limit` payloads remain unchanged. `Concat` now uses the +same shared-input representation as other operators. The +[operator-sharing proposal](operator-sharing.md#11-unified-operator-type) separately +widens those inputs to the common `Operator` and reconciles pre-ASAP/post-ASAP +operation differences. ### 2.2 Scalar expressions `ScalarExpr` contains every current scalar variant, including `CurrentTimestamp`. -Its recursive inputs are scalar expressions only. +Its recursive inputs are scalar expressions only. Individual recursive children +are boxed; variable-length children are owned lists. Neither form gives a scalar +expression shared DAG-node identity. ```rust -enum ScalarExpr { +enum ScalarExpr { Column(C), Literal(ScalarValue), - Compare { left: Rc>, op: CompareOpKind, right: Rc> }, + Compare { left: Box>, op: CompareOpKind, right: Box> }, BoolAnd(Vec>), BoolOr(Vec>), - Not(Rc>), - IsNull(Rc>), - IsNotNull(Rc>), - Cast { expr: Rc>, to: DataType, try_cast: bool }, - InList { expr: Rc>, list: Vec>, negated: bool }, + Not(Box>), + IsNull(Box>), + IsNotNull(Box>), + Cast { expr: Box>, to: DataType, try_cast: bool }, + InList { expr: Box>, list: Vec>, negated: bool }, FunctionCall { name: String, args: Vec> }, Arithmetic { - op: ArithmeticOpKind, left: Rc>, right: Rc>, + op: ArithmeticOpKind, left: Box>, right: Box>, }, Case { - operand: Option>>, + operand: Option>>, branches: Vec<(ScalarExpr, ScalarExpr)>, - else_expr: Option>>, + else_expr: Option>>, }, CurrentTimestamp, } @@ -192,6 +209,20 @@ Column resolution keeps its existing meaning: - An aggregate's `HAVING` expression uses the aggregate output schema. - A scan predicate uses the scanned data's schema. +The type split establishes the operator/scalar boundary; it does not by itself +prove that every scalar expression has the right value type. Two boundary rules +must also be validated: + +- **Predicate:** a resolved predicate must have Boolean type under the existing + expression-typing and coercion rules. Wrapping a numeric literal in `Predicate` + does not make it a valid condition. Nullable Boolean results retain the source + language's existing null semantics. +- **PromQL scalar bridge:** it has no input schema. Preserve its current role of + lifting a constant-folded PromQL scalar literal into an operator position; it + cannot accept free column references or arbitrary row-dependent expressions. + The existing scalar/vector conversion operators handle conversions involving + other operator results. + Splitting the representation must preserve evaluation behavior, inferred output types and source-language semantics. A scalar expression cannot serve as a table input, and a table-producing operator cannot appear where a scalar is expected. @@ -200,7 +231,14 @@ input, and a table-producing operator cannot appear where a scalar is expected. The example query must retain its result and output schema. The filter predicate and projection expression must resolve in the same contexts as before. Invalid -scalar/table combinations must be excluded by the plan model. +scalar/table combinations must be excluded by the plan model. In particular: + +- A `Concat` input has the same graph-node identity behavior as any other input. +- A scalar expression belongs to its owning operator; changing one operator's + expression cannot implicitly change another operator's expression. +- Non-Boolean resolved predicates and column-dependent scalar bridges are rejected. +- `CurrentTimestamp` remains scalar, and `EvalTimestamp` remains an operator with + its existing PromQL evaluation-time semantics. This separation alone does not require a change to the external plan format. The companion proposal addresses the separate decision to expose every operator in diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 2b8530973..0f6bfad19 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -109,8 +109,10 @@ The following changes are explicit: - Operator inputs become references to the common `Operator`. The sketches use the existing `Rc` notation for shared inputs and omit column-state generics for readability; column IDs and scan schemas below show the resolved form. -- `Predicate`, `ProjectItem` and other scalar-bearing payloads keep their roles, but - contain `ScalarExpr` after the scalar/operator split. +- `Predicate`, `ProjectItem` and other scalar-bearing payloads keep their roles and + own `ScalarExpr` values. The companion proposal defines owned scalar trees and + shared operator inputs, including `Concat`, as well as predicate and scalar-bridge + validation. This document uses those same structures. - `BinaryOp` reuses the existing post-ASAP `BinaryOperator` payload. It carries the binary operation, vector matching and checked-division requirements; the pre-ASAP `op` and `vector_match` semantics must be preserved when mapped into it. @@ -153,11 +155,11 @@ enum NonASAPOp { }, TimeRange { child: Rc, range: Duration }, TimeShift { child: Rc, shift: TimeShift }, - PromqlScalarBridge(Rc), + PromqlScalarBridge(ScalarExpr), EvalTimestamp, PromqlVectorFromScalar(Rc), PromqlScalarFromVector(Rc), - PromqlRelabel { child: Rc, dst: String, value: Rc }, + PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, PromqlInfoEnrich { child: Rc, selector: Vec }, PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, PromqlSubquery { child: Rc, range: Duration, resolution: Option }, From c793c83f9b9fbf67e433c6049b7b8638ded0a5b0 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:55:56 +0000 Subject: [PATCH 32/46] docs: clarify operator and scalar semantic boundaries --- .../proposals/decoupling_op_and_expr.md | 105 ++++++++++++++---- 1 file changed, 85 insertions(+), 20 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index 2c6348c1b..a1a7da66b 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -5,7 +5,7 @@ ## 1. Problem and goal -The current query representation mixes table-producing operators and scalar +The current query representation mixes query-plan operators and scalar expressions. Their roles are distinguished by where they occur, so an expression can be placed where a table input is expected and fail only when the plan is checked. @@ -32,8 +32,9 @@ compute values within the schema selected by their owning operators. ## 2. Proposed data structures -An **operator** describes a query computation with its own output schema. It reads -from a source or other operators and produces a relation or time-series result. +An **operator** describes a query-plan computation and its input dependencies. +Its result may be a relation, a time-series vector or a query-level scalar; the +plan must retain the result kind as well as its value types or output schema. For example, `Scan` produces input rows, `Filter` selects rows, and `Project` produces a new set of columns. Operators form the nodes and input dependencies of the query plan. @@ -202,6 +203,12 @@ their role preserves their existing semantics. ## 3. Semantic requirements +The split separates roles in the plan, not source-language result types. These +requirements define valid plans; retaining an existing variant does not establish +that its current implementation meets every requirement below. + +### 3.1 Expression context and result types + Column resolution keeps its existing meaning: - Most expressions use the input operator's output schema. @@ -209,23 +216,81 @@ Column resolution keeps its existing meaning: - An aggregate's `HAVING` expression uses the aggregate output schema. - A scan predicate uses the scanned data's schema. -The type split establishes the operator/scalar boundary; it does not by itself -prove that every scalar expression has the right value type. Two boundary rules -must also be validated: - -- **Predicate:** a resolved predicate must have Boolean type under the existing - expression-typing and coercion rules. Wrapping a numeric literal in `Predicate` - does not make it a valid condition. Nullable Boolean results retain the source - language's existing null semantics. -- **PromQL scalar bridge:** it has no input schema. Preserve its current role of - lifting a constant-folded PromQL scalar literal into an operator position; it - cannot accept free column references or arbitrary row-dependent expressions. - The existing scalar/vector conversion operators handle conversions involving - other operator results. - -Splitting the representation must preserve evaluation behavior, inferred output -types and source-language semantics. A scalar expression cannot serve as a table -input, and a table-producing operator cannot appear where a scalar is expected. +A resolved `Predicate` must have Boolean type under the expression-typing and +coercion rules. SQL conditions retain three-valued logic: `FALSE` and `NULL` do not +pass a filter. Wrapping a numeric value in `Predicate` does not make it Boolean. + +PromQL scalar, instant-vector and range-vector are query result kinds, distinct +from the internal `ScalarExpr` role. For example, `EvalTimestamp` and +`PromqlScalarFromVector` remain operators even though they produce scalar results. +Their consumers must validate the required result kind; a column schema alone +must not make a scalar interchangeable with a vector. + +`PromqlScalarBridge` has no input schema and retains its current restricted role: +lifting a constant-folded numeric literal into the plan. It cannot accept free +column references. Other scalar-valued queries use operator nodes, including the +existing scalar/vector conversions; PromQL scalar does not mean constant. + +### 3.2 PromQL binary and temporal operations + +`BinaryOp` operates on query results, including vector matching and label rules. +`ScalarExpr::Arithmetic` and `ScalarExpr::Compare` compute values within an +operator's context. Likewise, PromQL set operations remain distinct from SQL +`SetOp` and scalar Boolean expressions. + +PromQL comparison must distinguish filtering from the `bool` mode: `up > 0` +filters samples, whereas `up > bool 0` produces numeric `0` or `1` for matched, +valid samples. Scalar/scalar comparisons require `bool`. The current `BinaryOp` +payload and frontend conversion do not retain this mode. Supporting it requires +preserving the distinction; otherwise the frontend must reject it rather than +silently change the result. + +`TimeRange`, `PromqlSubquery` and `SQLWindowFunc` retain separate roles. A PromQL +subquery evaluates its input over a time grid and produces a range vector; a SQL +window computes values over partitions and frames of input rows. Subquery range, +resolution, `offset` and `@` must retain their meaning and scope. The current +subquery conversion retains only range, resolution and child. A design using the +existing `TimeShift` must specify how it shifts or fixes the subquery evaluation +time, including nested subqueries; unsupported modifiers must be rejected. + +These distinctions follow the PromQL references for +[result types and subqueries](https://prometheus.io/docs/prometheus/latest/querying/basics/) +and [binary operations](https://prometheus.io/docs/prometheus/latest/querying/operators/). + +### 3.3 SQL aggregate, window and subquery boundaries + +DataFusion 43's [`Expr`](https://github.com/apache/datafusion/blob/43.0.0/datafusion/expr/src/expr.rs) +includes aggregate, window and subquery expressions as well as ordinary scalar +expressions. It therefore does not map directly to this proposal's `ScalarExpr`. +Aggregate and window computations belong to `Aggregate` and `SQLWindowFunc`; +subsequent scalar expressions reference their output columns. + +Where an operator accepts only column references, an expression argument must be +computed by an input `Project` or explicitly rejected. For example, +`SUM(price * quantity)` can become a projection of the product followed by an +aggregate over that column. The current aggregate conversion rejects such +arguments; this example describes a valid extension, not existing support. +The same rule applies to expression-valued grouping and partition keys. + +Aggregate `FILTER`, `DISTINCT`, internal `ORDER BY`, and aggregate/window null +treatment must be preserved, translated equivalently or explicitly rejected. +They cannot be dropped merely because the current payload lacks a field. +General scalar subqueries remain outside this proposal's supported scope, as +specified in §4; the split does not claim full DataFusion SQL coverage. + +### 3.4 Evaluation behavior + +Owning or copying a `ScalarExpr` does not authorize changing how often it is +evaluated. Function resolution must retain the typing and evaluation properties +needed to preserve semantics, whether through the existing function catalog or +another established resolution mechanism. + +DataFusion distinguishes +[immutable, stable and volatile functions](https://github.com/apache/datafusion/blob/43.0.0/datafusion/expr-common/src/signature.rs). +`CurrentTimestamp` (`NOW()`) stays stable within a SQL query; separate calls to a +volatile function such as `random()` may differ. Expression copying, sharing or +movement must respect those properties. `EvalTimestamp` instead follows the +PromQL evaluation timestamp, including evaluation inside a subquery. ## 4. Acceptance and scope From 66531a81bcd164dedffe74da8fd517c1a3bd1873 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:10:23 +0000 Subject: [PATCH 33/46] docs: map operator and scalar design to current SQL and PromQL semantics --- .../proposals/decoupling_op_and_expr.md | 447 +++++++++++------- .../design_docs/proposals/operator-sharing.md | 52 +- 2 files changed, 295 insertions(+), 204 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index a1a7da66b..b1cc1378c 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -32,60 +32,37 @@ compute values within the schema selected by their owning operators. ## 2. Proposed data structures -An **operator** describes a query-plan computation and its input dependencies. -Its result may be a relation, a time-series vector or a query-level scalar; the -plan must retain the result kind as well as its value types or output schema. -For example, `Scan` produces input rows, `Filter` selects rows, and `Project` -produces a new set of columns. Operators form the nodes and input dependencies of -the query plan. - -A **scalar expression** describes a value calculated within the context of an -operator. For example, `l_quantity > 10` produces a Boolean used by a filter, and -`l_quantity * 2` produces a value for a projected column. An expression has a value -type, but no table output schema or independent data source. Its column references -are interpreted using the schema chosen by the owning operator. - -In the example above, `Filter.child` is the `Scan` operator that supplies rows; -`Filter.pred` is the expression `l_quantity > 10` evaluated on those rows. The -predicate cannot replace the scan as the filter's input, and the scan cannot be -used as its Boolean condition. This is the distinction the two types enforce. - -Split `QueryExpr` into `NonASAPOp` and `ScalarExpr`. Keep existing variant names, -field names and semantics; change their types to express their roles: - -| Position | Current type | Proposed type | -|---|---|---| -| Operator input | `QueryExpr` | `NonASAPOp` | -| Scalar expression | `QueryExpr` | `ScalarExpr` | +`NonASAPOp` describes a relation or vector computation: reading data, selecting +rows or samples, combining inputs, or reducing them. `ScalarExpr` computes one +value in a column and evaluation context. An operator supplies that context; +a scalar query root has no row columns. The distinction is about semantic role, +not whether the source language calls something an “expression”. -The split also makes ownership consistent with each structure's role: +Keep existing names where the semantics match. The proposal changes the boundary +where needed, rather than copying every current `QueryExpr` variant unchanged: -| Structure | Proposed rule | Reason | +| Current representation | Proposed representation | Reason | |---|---|---| -| Operator inputs | Shared references, including every `Concat` child | All inputs are graph nodes; a multi-input operator should not embed copies while other operators reference nodes. | -| Scalar fields on operators | Owned `ScalarExpr` values, including predicates, relabel values and scalar bridges | These expressions belong to the operator's schema context and need no independent graph identity. | -| Scalar recursion | `Box` for individual recursive children; `Vec` for lists | Scalars form owned expression trees. `Box` gives recursive single-child fields finite size; `Vec` already provides that indirection for lists. | -| Column parameter | Keep `C: ColState` on operators; use plain `C` on scalar expressions and their wrappers | Operators need `C::ScanSchema`; scalar expressions only carry column references. | - -`C` continues to represent `ColumnRef` before resolution and `ColumnId` afterward. -No new column trait is introduced. Required cloning, comparison or serialization -bounds belong on the corresponding implementations, not on scalar data structures -through the unrelated scan-schema requirement. - -These are deliberate changes from the current mixed `Rc`/value representation. -They preserve operation names and evaluation semantics; they do not add scalar -sharing or a new optimization policy. +| `PromqlScalarBridge(Literal(...))` | `ScalarExpr::Literal` | Constants need no operator node. | +| Operator `EvalTimestamp` | Scalar `EvalTimestamp` | `time()` reads evaluation context and returns one number. | +| Operator `PromqlScalarFromVector` | Scalar `PromqlScalarFromVector` referencing its input plan | `scalar(v)` has real cardinality/conversion semantics. | +| Scalar operands wrapped as `BinaryOp` inputs | `Arithmetic` / `Compare` inside `Project` or `Filter` | Vector/scalar operations need no constant-producing input node. | +| `BinaryOp` without comparison mode | Vector/vector `BinaryOp` with `return_bool` | Filtering and numeric comparison results differ. | +| Undifferentiated `TimeRange` | `TimeRange` with instant/range kind | Selecting the latest sample differs from selecting all samples in an interval. | ### 2.1 Operator nodes -`NonASAPOp` contains the current operator variants. Inputs reference other operators; -predicates and expression-bearing fields use the scalar structures below. +The sketches describe the proposed shape, not implemented Rust definitions. +`C` remains `ColumnRef` before resolution and `ColumnId` afterward; +`C::ScanSchema` retains its existing meaning. Operator references are shared, +including `Concat` inputs. Scalar fields own expression trees. ```rust enum NonASAPOp { Scan { source: Source, predicates: Vec>, schema: C::ScanSchema, }, + Values { rows: Vec>>, schema: C::ScanSchema }, Filter { pred: Predicate, child: Rc> }, Project { cols: Vec>, qualifier: Option, child: Rc>, @@ -108,22 +85,22 @@ enum NonASAPOp { left: Rc>, right: Rc>, }, Sort { keys: Vec>, partition_by: GroupKeys, child: Rc> }, - Limit { n: usize, offset: usize, child: Rc> }, + Limit { + n: Option, offset: usize, partition_by: GroupKeys, + child: Rc>, + }, BinaryOp { op: BinaryOpKind, lhs: Rc>, rhs: Rc>, - vector_match: Option, + vector_match: Option, return_bool: bool, }, SQLWindowFunc { func: WindowFuncKind, args: Vec>, partition_by: GroupKeys, order_by: Vec>, frame: Option, output_name: String, child: Rc>, }, - TimeRange { range: Duration, child: Rc> }, + TimeRange { range: Duration, kind: TimeRangeKind, child: Rc> }, TimeShift { shift: TimeShift, child: Rc> }, - PromqlScalarBridge(ScalarExpr), - EvalTimestamp, - PromqlVectorFromScalar(Rc>), - PromqlScalarFromVector(Rc>), + PromqlVectorFromScalar(ScalarExpr), PromqlRelabel { dst: String, value: ScalarExpr, child: Rc> }, PromqlInfoEnrich { selector: Vec, child: Rc> }, PromqlSeriesSample { by: GroupKeys, kind: SampleKind, child: Rc> }, @@ -131,51 +108,57 @@ enum NonASAPOp { range: Duration, resolution: Option, child: Rc>, }, } -``` -The following structures describe scalar-bearing fields of `NonASAPOp`: -`Predicate` is used by filters, joins and `HAVING`; `ProjectItem` by projections; -and `SortKey` by sorting and window functions. Their expressions use `ScalarExpr`, -defined in §2.2. +enum TimeRangeKind { Instant, Range } -```rust -struct Predicate(ScalarExpr); +struct Predicate(ScalarExpr); -struct ProjectItem { +struct ProjectItem { alias: Option, expr: ScalarExpr, } -struct SortKey { +struct SortKey { expr: ScalarExpr, ascending: bool, nulls_first: bool, } ``` -`Predicate`, `ProjectItem` and `SortKey` retain their names and roles and consistently -own their scalar expressions. `Predicate` marks a Boolean-expression position; -`ProjectItem` adds an alias, and `SortKey` adds ordering and null-placement rules. -These are different operator requirements, so the wrappers remain separate. +`Predicate`, `ProjectItem` and `SortKey` remain separate: they describe a condition, +a named output and an ordering requirement. `Values` is a relation constructor +for SQL `VALUES` and `SELECT` without `FROM`, not a wrapper for PromQL scalars. +One empty row provides the input for `SELECT 1`; zero rows represent an empty +relation. Its expressions have no input-column scope. -The pre-ASAP `BinaryOp` and `Limit` payloads remain unchanged. `Concat` now uses the -same shared-input representation as other operators. The -[operator-sharing proposal](operator-sharing.md#11-unified-operator-type) separately -widens those inputs to the common `Operator` and reconciles pre-ASAP/post-ASAP -operation differences. +`Limit.n = None` permits offset without a limit. `partition_by` adopts the existing +post-ASAP field for limits within groups, including constant-parameter PromQL +top/bottom selection. `BinaryOp` retains vector matching and gains `return_bool`, +valid only for comparisons. Scalar/vector cases lower as described in §3.3. + +`TimeRangeKind::Instant` uses `range` as the lookback horizon and selects the latest +eligible sample per series, respecting staleness. `Range` selects a sample window. +This fixes the current ambiguity between selectors; lookback is not the ingestion +interval. `TimeShift` around a selector or `PromqlSubquery` changes the evaluation +time of that whole input. Existing signed offsets and `AtModifier` anchors remain. ### 2.2 Scalar expressions -`ScalarExpr` contains every current scalar variant, including `CurrentTimestamp`. -Its recursive inputs are scalar expressions only. Individual recursive children -are boxed; variable-length children are owned lists. Neither form gives a scalar -expression shared DAG-node identity. +Scalar recursion uses owned `Box` and `Vec` children. The only plan references are +explicit operations that consume a query result to compute a value. Those edges +remain visible to plan traversal and costing; they cannot hide a separate plan. +Since these variants reference `NonASAPOp`, `ScalarExpr` and its wrappers retain +`C: ColState`; the earlier proposed removal of that bound no longer applies. ```rust -enum ScalarExpr { +enum ScalarExpr { Column(C), Literal(ScalarValue), - Compare { left: Box>, op: CompareOpKind, right: Box> }, + Negative { expr: Box>, semantics: ExprSemantics }, + Compare { + left: Box>, op: CompareOpKind, right: Box>, + semantics: ExprSemantics, + }, BoolAnd(Vec>), BoolOr(Vec>), Not(Box>), @@ -186,6 +169,7 @@ enum ScalarExpr { FunctionCall { name: String, args: Vec> }, Arithmetic { op: ArithmeticOpKind, left: Box>, right: Box>, + semantics: ExprSemantics, }, Case { operand: Option>>, @@ -193,123 +177,218 @@ enum ScalarExpr { else_expr: Option>>, }, CurrentTimestamp, + EvalTimestamp, + PromqlScalarFromVector(Rc>), + ScalarSubquery(Rc>), + Exists { subquery: Rc>, negated: bool }, + InSubquery { + expr: Box>, subquery: Rc>, negated: bool, + }, +} + +enum ExprSemantics { Sql, Promql } +``` + +`ExprSemantics` distinguishes numeric/comparison rules even when both languages +use `Float64`; a result type alone does not preserve NaN, ordering or error rules. +`Negative` preserves unary negation directly. `Compare` returns Boolean internally; +PromQL numeric comparison results use `Case` to produce `1.0` or `0.0`. +`FunctionCall.name` must resolve to an unambiguous function contract, including +argument/result types, null behavior and volatility. An arbitrary name is not +proof that a function is supported. + +The plan-reading variants have different contracts: + +| Scalar variant | Required input and result | +|---|---| +| `PromqlScalarFromVector` | An instant vector; one float sample becomes its value, otherwise NaN. | +| `ScalarSubquery` | A one-column SQL relation; zero rows gives typed NULL, one row gives its value, multiple rows is an error. | +| `Exists` | A SQL relation; returns a non-null Boolean based on whether it has any rows. | +| `InSubquery` | A one-column SQL relation; applies SQL membership and NULL rules, including for `NOT IN`. | + +These consume query results; they are not interchangeable bridge nodes. The SQL +variants cover uncorrelated subqueries here. Correlation needs outer-scope bindings +that this proposal does not define (§3.4). + +### 2.3 Query roots and composition without a bridge + +A query root selects one of the two representations; this tag adds no computation: + +```rust +enum QueryRoot { + Operator(Rc>), + Scalar(ScalarExpr), } ``` -`PromqlScalarBridge`, `PromqlVectorFromScalar`, `PromqlScalarFromVector` and -`EvalTimestamp` remain operators because they participate in query-level evaluation. -`CurrentTimestamp` (SQL `NOW()`) belongs to scalar expressions. Classifying these by -their role preserves their existing semantics. +SQL query results use the operator arm. PromQL numeric/string roots use the scalar +arm; vector roots use the operator arm. A scalar root cannot contain free columns. +`EvalTimestamp` reads the current PromQL evaluation time; `CurrentTimestamp` reads +SQL's statement time. Moving `EvalTimestamp` into `ScalarExpr` must not make it +constant across evaluation steps or nested subquery times. + +For float samples, the following plans need no `PromqlScalarBridge`: + +```text +2 → Scalar root: Literal(2.0) +time() → Scalar root: EvalTimestamp +up * 2 → Project(sample * 2.0, child = instant selection of up) +vector(time()) → PromqlVectorFromScalar(EvalTimestamp) +scalar(sum(up)) + 1 → Scalar root: Arithmetic( + PromqlScalarFromVector(Aggregate(...)), Literal(1.0)) +``` + +A PromQL projection retains the time and label fields required by the operation +and applies its metric-name rules; it does not project only the numeric sample. +For an open label schema, lowering must retain the complete series identity, +including unreferenced labels. If the input provides neither a complete label +schema nor a full identity value, this lowering is not valid. +`vector(s)` remains a real conversion to a one-element, label-free vector. +When combined with the [operator-sharing proposal](operator-sharing.md), all plan +references above target the common `Operator`, including references inside scalar +expressions and query roots. Scalar expression trees do not acquire shared-node +identity. ## 3. Semantic requirements -The split separates roles in the plan, not source-language result types. These -requirements define valid plans; retaining an existing variant does not establish -that its current implementation meets every requirement below. - -### 3.1 Expression context and result types - -Column resolution keeps its existing meaning: - -- Most expressions use the input operator's output schema. -- A join predicate uses both input schemas. -- An aggregate's `HAVING` expression uses the aggregate output schema. -- A scan predicate uses the scanned data's schema. - -A resolved `Predicate` must have Boolean type under the expression-typing and -coercion rules. SQL conditions retain three-valued logic: `FALSE` and `NULL` do not -pass a filter. Wrapping a numeric value in `Predicate` does not make it Boolean. - -PromQL scalar, instant-vector and range-vector are query result kinds, distinct -from the internal `ScalarExpr` role. For example, `EvalTimestamp` and -`PromqlScalarFromVector` remain operators even though they produce scalar results. -Their consumers must validate the required result kind; a column schema alone -must not make a scalar interchangeable with a vector. - -`PromqlScalarBridge` has no input schema and retains its current restricted role: -lifting a constant-folded numeric literal into the plan. It cannot accept free -column references. Other scalar-valued queries use operator nodes, including the -existing scalar/vector conversions; PromQL scalar does not mean constant. - -### 3.2 PromQL binary and temporal operations - -`BinaryOp` operates on query results, including vector matching and label rules. -`ScalarExpr::Arithmetic` and `ScalarExpr::Compare` compute values within an -operator's context. Likewise, PromQL set operations remain distinct from SQL -`SetOp` and scalar Boolean expressions. - -PromQL comparison must distinguish filtering from the `bool` mode: `up > 0` -filters samples, whereas `up > bool 0` produces numeric `0` or `1` for matched, -valid samples. Scalar/scalar comparisons require `bool`. The current `BinaryOp` -payload and frontend conversion do not retain this mode. Supporting it requires -preserving the distinction; otherwise the frontend must reject it rather than -silently change the result. - -`TimeRange`, `PromqlSubquery` and `SQLWindowFunc` retain separate roles. A PromQL -subquery evaluates its input over a time grid and produces a range vector; a SQL -window computes values over partitions and frames of input rows. Subquery range, -resolution, `offset` and `@` must retain their meaning and scope. The current -subquery conversion retains only range, resolution and child. A design using the -existing `TimeShift` must specify how it shifts or fixes the subquery evaluation -time, including nested subqueries; unsupported modifiers must be rejected. - -These distinctions follow the PromQL references for -[result types and subqueries](https://prometheus.io/docs/prometheus/latest/querying/basics/) -and [binary operations](https://prometheus.io/docs/prometheus/latest/querying/operators/). - -### 3.3 SQL aggregate, window and subquery boundaries - -DataFusion 43's [`Expr`](https://github.com/apache/datafusion/blob/43.0.0/datafusion/expr/src/expr.rs) -includes aggregate, window and subquery expressions as well as ordinary scalar -expressions. It therefore does not map directly to this proposal's `ScalarExpr`. -Aggregate and window computations belong to `Aggregate` and `SQLWindowFunc`; -subsequent scalar expressions reference their output columns. - -Where an operator accepts only column references, an expression argument must be -computed by an input `Project` or explicitly rejected. For example, -`SUM(price * quantity)` can become a projection of the product followed by an -aggregate over that column. The current aggregate conversion rejects such -arguments; this example describes a valid extension, not existing support. -The same rule applies to expression-valued grouping and partition keys. - -Aggregate `FILTER`, `DISTINCT`, internal `ORDER BY`, and aggregate/window null -treatment must be preserved, translated equivalently or explicitly rejected. -They cannot be dropped merely because the current payload lacks a field. -General scalar subqueries remain outside this proposal's supported scope, as -specified in §4; the split does not claim full DataFusion SQL coverage. - -### 3.4 Evaluation behavior - -Owning or copying a `ScalarExpr` does not authorize changing how often it is -evaluated. Function resolution must retain the typing and evaluation properties -needed to preserve semantics, whether through the existing function catalog or -another established resolution mechanism. - -DataFusion distinguishes -[immutable, stable and volatile functions](https://github.com/apache/datafusion/blob/43.0.0/datafusion/expr-common/src/signature.rs). -`CurrentTimestamp` (`NOW()`) stays stable within a SQL query; separate calls to a -volatile function such as `random()` may differ. Expression copying, sharing or -movement must respect those properties. `EvalTimestamp` instead follows the -PromQL evaluation timestamp, including evaluation inside a subquery. +### 3.1 Version and coverage contract + +This section maps source semantics to §2. **Direct** means the structure records +the operation; **Lowered** means an equivalent composition is specified. Neither +means implemented or runtime-verified. **Partial** limits coverage to individually +registered function contracts. **Gap** means §2 lacks a necessary payload, +type or binding rule; the frontend must reject that case until it is supplied. + +| Language | Version used for this proposal | Repository relationship | +|---|---|---| +| DataFusion SQL | **DataFusion 55.1.0**, its SQL query dialect and logical expressions | Target semantic baseline. The repository still uses 43.0.0; dependency migration is separate. This is not a claim about every SQL standard or dialect. | +| PromQL | **Prometheus 3.15.0**, with experimental features identified separately | Target semantic baseline. The repository pins ProjectASAP parser revision `9fede7eecca923c9882fe256484d00d37f8706cb`, declaring 3.8 compatibility; upgrading that parser is separate. | + +References are pinned to those versions: +[DataFusion SELECT][df-select], [logical plans][df-plan], [expressions][df-expr], +[aggregates][df-agg], [windows][df-window], [function volatility][df-volatility]; +Prometheus [query basics][prom-basics], [operators][prom-operators], +[functions][prom-functions], [AST][prom-ast] and [function signatures][prom-signatures]. +The [DataFusion release][df-release] and [Prometheus release][prom-release] identify +the targets; the parser's [compatibility declaration][parser-version] describes the +older repository dependency. This documentation change upgrades neither dependency. + +### 3.2 SQL semantic mapping + +| SQL construct / example | Representation in §2 | Coverage and semantic condition | +|---|---|---| +| `FROM t`, `WHERE x > 1`, `SELECT x * 2` | `Scan`, `Filter`, `Project`; scalar `Compare`, `Arithmetic` | **Direct.** Resolve columns against the input; predicates must be Boolean and only TRUE passes. | +| `VALUES (1), (2)`; `SELECT 1` | `Values`; `Project` over one empty row | **Direct.** SQL still returns a relation, unlike a PromQL scalar root. | +| Literals, columns, `CASE`, `CAST`, `TRY_CAST`, `IN (...)`, `IS NULL`, Boolean logic | Corresponding `ScalarExpr` variants | **Direct** within supported types. Preserve coercion, NULL propagation and conditional evaluation. Type gaps are listed below. | +| Unary minus, `BETWEEN`, `IS TRUE`, null-safe equality | `Negative`; compositions of `Compare`, `Case`, `IsNull`, Boolean expressions | **Direct/Lowered.** Evaluate reused nontrivial operands once, using a projected column where needed; do not duplicate volatile calls. | +| Scalar functions and `NOW()` | Resolved `FunctionCall`; `CurrentTimestamp` | **Direct** for registered contracts. Preserve stable/volatile evaluation behavior; unresolved names are unsupported. | +| Joins, cross joins, semi/anti joins | `Join` and its predicate | **Direct.** Predicate scope includes both inputs. Preserve outer-join null extension; ordinary anti-join is not nullable `NOT IN`. | +| `GROUP BY`, `SUM(x)`, `HAVING` | `Aggregate`, `Reduction`, existing `AggIntent`; scalar predicate on aggregate outputs | **Direct** for represented intents. `COUNT(*)` counts rows; `COUNT(x)` requires counting non-null values, not reusing row count unchanged. | +| `SUM(price * quantity)`; expression grouping keys | Input `Project`, then `Aggregate` referencing its output columns | **Lowered.** Current conversion rejects some expression arguments; the proposed representation can retain them. | +| `COUNT(x)` alongside other measures | Project a 0/1 non-null indicator, sum it, return 0 for an empty global group | **Lowered.** Filtering the whole input would incorrectly change the other measures. | +| `QUALIFY`, `DISTINCT ON` | `Filter` after window evaluation; ordered partitioned `Limit(n = Some(1))` | **Lowered.** Resolve aliases first; retain window/filter order and the selected row. | +| `ROLLUP`, `CUBE`, `GROUPING SETS` | One `Aggregate` per grouping set, projected missing keys and grouping discriminator, then `Concat` | **Lowered.** Retain duplicate grouping sets and distinguish omitted keys from input NULLs. | +| Wildcards, `UNION BY NAME`, pipe syntax for supported operations | Resolve columns, align with `Project`, compose the corresponding operators | **Lowered.** These syntax forms need no additional computational category. | +| `ROW_NUMBER()`, `LAG(x)`, `SUM(x) OVER (...)` | `SQLWindowFunc`, scalar args and sort keys; `Project` for expression partition keys | **Direct/Lowered** for existing `WindowFuncKind` and ROWS/RANGE frames. Window output is a column, not a scalar aggregate call. Missing features are listed below. | +| `DISTINCT`, `UNION`, `INTERSECT`, `EXCEPT`, their supported `ALL` forms | `Dedup`, `Concat` / `SetOp` | **Direct.** Preserve bag multiplicity and SQL duplicate/NULL equality rules. | +| `ORDER BY`, `LIMIT`, offset-only queries | `Sort`, `Limit` | **Direct.** Preserve direction and NULL placement; `n = None` means no fetch limit. | +| Uncorrelated scalar subquery, `EXISTS`, `IN` / `NOT IN (SELECT ...)` | Explicit scalar plan-reading variants | **Direct.** Preserve the cardinality and NULL contracts in §2.2, including when used in a SELECT list. | +| Derived tables and nonrecursive CTEs | Existing operator subgraphs; aliases resolved to output columns | **Lowered.** Naming alone needs no computation node. Reuse must not alter volatile evaluation. | + +### 3.3 PromQL semantic mapping + +Operator results must distinguish relations, instant vectors and range vectors; +`ScalarExpr` has a value type. Infer and validate these kinds from the source and +operation, rather than treating identical column schemas as interchangeable. +A range-query request evaluates its root at successive timestamps; it is not a +range-vector expression. The request permits scalar or instant-vector roots. + +| PromQL construct / example | Representation in §2 | Coverage and semantic condition | +|---|---|---| +| Numeric/string literals; parentheses; `time()` | Scalar `Literal`, nested expression, `EvalTimestamp` | **Direct.** No bridge node; free columns are invalid at a scalar root. | +| Scalar arithmetic, unary minus, `1 < bool 2` | `Arithmetic`, `Negative`, `Case(Compare(...), 1.0, 0.0)` | **Direct/Lowered** with PromQL numeric semantics. Scalar comparison without `bool` is invalid. | +| `up{job="api"}` | Time-series `Scan` with predicates, `TimeRange(Instant)` | **Direct.** Label matching treats absent labels as empty and regexes as anchored. Selection uses lookback and staleness, not SQL row filtering alone. | +| `up[5m]` | `TimeRange(Range)` over a time-series scan | **Direct.** Retain samples and their timestamps in the left-open, right-closed window. | +| `-up` | `Project` with `Negative(sample)` | **Lowered** for float samples; retain series identity and unary-operation naming rules. | +| `up * 2`, `2 / up`, `up * scalar(sum(other))` | `Project` with scalar arithmetic and, where needed, `PromqlScalarFromVector` | **Lowered** for float samples. Preserve operand order, time/label fields and metric-name removal. Scalar input is evaluated in the same query-time context. | +| `up > 0`; `up > bool 0` | `Filter`; or `Project` with `Case(Compare(...), 1.0, 0.0)` | **Lowered** for float samples. Filtering preserves surviving sample values; bool mode produces numbers and removes the metric name. Scalar-on-left comparisons still retain the vector's sample value when filtering. | +| `a / on(job) group_left b`; `a > bool b` | `BinaryOp` with `vector_match`, `return_bool` | **Direct.** Retain matching cardinality, labels and metric-name rules; unmatched elements disappear, not become false rows. | +| `a and b`, `a or b`, `a unless b` | `BinaryOpKind::Set` | **Direct.** Label-set matching, not SQL Boolean evaluation or SQL bag set operations. | +| `sum by(job)(up)`, `avg without(instance)(up)` | `Aggregate(Reduction::Reduce(...), AggIntent)` | **Direct** for represented intents; preserve PromQL label grouping and empty-input behavior. | +| `topk(3, up)`, `bottomk(3, up)` with optional grouping | `Sort` + partitioned `Limit` | **Lowered** for constant parameters and float samples, retaining selected series labels and specified NaN ordering. Existing heavy-hitter `AggIntent::TopK` is not a substitute for sample-value ranking. | +| `rate(x[5m])`, `sum_over_time(x[5m])` | Range input + `Aggregate(Reduction::PerEntity, corresponding AggIntent)` | **Direct** for existing contracts: counter resets, extrapolation and range reduction belong to the named intent, not ordinary SQL SUM. | +| `vector(s)`, `scalar(v)` | `PromqlVectorFromScalar`; scalar `PromqlScalarFromVector` | **Direct.** These change result kind/cardinality and cannot be removed as representation wrappers. | +| `expr[30m:1m] offset 5m`, selectors with `@` | `PromqlSubquery`; surrounding `TimeShift` | **Direct.** Apply the anchor/offset to the whole child evaluation; preserve nested grids, default resolution and query-level start/end anchors. | +| Per-sample math/date functions, including scalar parameters | `Project` with a resolved scalar `FunctionCall` | **Lowered** for registered float-sample contracts. Scalar parameters can be expressions; do not require constant-only `AggIntent::Math` payloads. | +| Request-context functions and duration expressions, e.g. `x[max_of(step(), 5s)]` | Context-reading `FunctionCall`; resolve duration arithmetic before constructing `TimeRange` / `TimeShift` | **Lowered** when request parameters and the relevant subquery context determine the duration. A stored duration is valid only for that binding. | +| Relabeling, absence, `info`, series sampling | `PromqlRelabel`, relevant `AggIntent`, `PromqlInfoEnrich`, `PromqlSeriesSample` | **Partial.** Registered contracts must preserve label construction, missing-series behavior and feature gates. Parameter/type gaps below remain. | + +### 3.4 Remaining semantic gaps + +These are limits of the proposed payloads or an as-yet unspecified lowering, not +reasons to mix all operators and expressions back into one enum. The tables above +cover the core structures; they do not assert full language conformance. + +| Semantics not fully covered | Why §2 cannot currently express it faithfully | Required extension or decision | +|---|---|---| +| General SQL aggregate `FILTER`, `DISTINCT`, internal ordering, null treatment and arbitrary aggregate UDFs | `AggIntent` is a fixed intent vocabulary with no general per-measure modifier payload. `COUNT(DISTINCT ...)` has `Cardinality`, but this does not cover every aggregate. | Define per-measure semantics or an equivalent lowering for each case. A filter on the whole aggregate input is not a general replacement. | +| All DataFusion window functions, GROUPS frames, explicit null treatment, window `FILTER` / `DISTINCT` | `WindowFuncKind` is a subset; `WindowFrameUnits` only has Rows/Range; the window payload lacks these modifier fields. | Extend these existing payloads for the requested features; ordinary scalar `FunctionCall` cannot supply window context. | +| Correlated SQL subqueries and recursive CTEs | Scalar plan references have no outer-scope binding, and an ordinary DAG has no recursive/fixpoint contract. | Define correlation scopes and recursive evaluation separately; only proven equivalent decorrelation is usable today. | +| SQL higher-order functions and lambda expressions | `FunctionCall.args` has no lambda parameter bindings or body scope; ordinary column references cannot stand for lambda variables. | Add explicit scalar lambda/binding structures before claiming coverage. | +| General SQL `ANY` / `ALL` subquery comparisons | `InSubquery` only records membership, not a comparison operator and quantifier. | Add a scalar `SetComparison` payload or a proven lowering retaining empty-set and NULL behavior. | +| Unbound SQL parameters / scalar variables | No placeholder or variable binding is represented. | Bind them to typed values before this IR, or define a binding payload; a free column is not a parameter. | +| Full DataFusion value/type fidelity | Existing `DataType`/`ScalarValue` cannot retain every decimal, unsigned width, timestamp unit/timezone or typed nested literal. Widening values can change result types and errors. | Extend the value/type model; casts cannot recover information already lost. The SQL mapping is restricted to faithfully represented types. | +| SQL pattern matching with explicit escape rules and all dialect-specific scalar operators | Current `CompareOpKind` pattern variants do not carry `ESCAPE`; function names alone do not define missing semantics. | Add the missing payload or a registered equivalent scalar contract. Simple LIKE does not prove ESCAPE support. | +| SQL `UNNEST` / general table functions | No operator expands a collection or invokes a table-valued function with its output cardinality/schema contract. | Add a relational operation, not a scalar function pretending to produce rows. | +| PromQL native-histogram samples and mixed-sample behavior | Existing value types have no native-histogram sample representation or annotation contract. Named histogram intents alone do not preserve those samples. | Extend sample types and define invalid-operation/annotation behavior; the float mappings above do not cover this case. | +| Dynamic PromQL parameters, e.g. `quantile(scalar(q), up)` | `AggIntent.q`, sampling parameters and `Limit.n` store constants rather than expression dependencies. Per-sample math can instead use `FunctionCall` as mapped above. | Permit scalar expressions in the relevant parameter positions. Constant folding only covers genuinely constant inputs. | +| PromQL experimental fill modifiers | `VectorMatch` has no left/right fill values, so it cannot distinguish dropping an unmatched series from supplying a numeric default. | Extend that payload for `fill`, `fill_left`, `fill_right`; preserve the `promql-binop-fill-modifiers` feature gate. | +| Extended PromQL range selectors and start-timestamp functions | `TimeRangeKind` has no anchored/smoothed selection mode, and the sample schema has no distinct start-timestamp metadata contract. | Define the temporal/sample metadata and feature gates before mapping these forms. A normal sample timestamp is not a start timestamp. | + +### 3.5 Context and evaluation invariants + +Scalar columns resolve against the owning input, both inputs for a join predicate, +and aggregate outputs for `HAVING`. Scan predicates use the source schema. SQL +subqueries in this proposal have their own scope and no implicit outer references. +Predicate wrapping does not bypass Boolean typing or SQL three-valued logic. + +Function resolution preserves volatility: SQL `NOW()` is stable within a statement; +repeated volatile calls need not agree. Expression ownership does not authorize +copying, sharing or moving evaluations. PromQL scalar plan reads and +`EvalTimestamp` use the active evaluation instant, including within subqueries. +A reused operator must not be evaluated once and then incorrectly reused across +different time contexts. ## 4. Acceptance and scope -The example query must retain its result and output schema. The filter predicate -and projection expression must resolve in the same contexts as before. Invalid -scalar/table combinations must be excluded by the plan model. In particular: - -- A `Concat` input has the same graph-node identity behavior as any other input. -- A scalar expression belongs to its owning operator; changing one operator's - expression cannot implicitly change another operator's expression. -- Non-Boolean resolved predicates and column-dependent scalar bridges are rejected. -- `CurrentTimestamp` remains scalar, and `EvalTimestamp` remains an operator with - its existing PromQL evaluation-time semantics. - -This separation alone does not require a change to the external plan format. -The companion proposal addresses the separate decision to expose every operator in -the exported graph. - -Scalar subqueries are outside this design. Filter `IN (SELECT …)` and `EXISTS` can -be represented as joins; a general scalar subquery would require expressions to -reference operator graphs and needs a separate design. This proposal adds no new -optimization, accuracy or execution-timing behavior. +This is a representation proposal, not a runtime implementation or a declaration +of complete SQL/PromQL support. Acceptance requires: + +- The SQL example in §1 retains its values, schema and predicate context. +- Every Direct/Lowered mapping in §3 has a valid typed representation; each Gap is + explicitly rejected until its payload or equivalent lowering is defined. +- The scalar roots and mixed scalar/vector examples in §2.3 need no + `PromqlScalarBridge` or equivalent constant-wrapper node. +- Operator dependencies inside scalar conversions/subqueries remain visible and + shared; scalar trees remain owned. Invalid result-kind combinations are rejected. +- SQL NULL/cardinality rules, PromQL labels and evaluation times survive conversion. + +Implementation will require frontend, validation and plan-format migration for +these explicit structural changes. DDL/DML, session commands, physical execution, +new optimization algorithms, accuracy and execution-timing policy are outside this +proposal. The companion document defines the common pre-/post-ASAP operator graph. + +[df-release]: https://github.com/apache/datafusion/releases/tag/55.1.0 +[prom-release]: https://github.com/prometheus/prometheus/releases/tag/v3.15.0 +[df-select]: https://github.com/apache/datafusion/blob/55.1.0/docs/source/user-guide/sql/select.md +[df-plan]: https://github.com/apache/datafusion/blob/55.1.0/datafusion/expr/src/logical_plan/plan.rs +[df-expr]: https://github.com/apache/datafusion/blob/55.1.0/datafusion/expr/src/expr.rs +[df-agg]: https://github.com/apache/datafusion/blob/55.1.0/docs/source/user-guide/sql/aggregate_functions.md +[df-window]: https://github.com/apache/datafusion/blob/55.1.0/docs/source/user-guide/sql/window_functions.md +[df-volatility]: https://github.com/apache/datafusion/blob/55.1.0/datafusion/expr-common/src/signature.rs +[prom-basics]: https://github.com/prometheus/prometheus/blob/v3.15.0/docs/querying/basics.md +[prom-operators]: https://github.com/prometheus/prometheus/blob/v3.15.0/docs/querying/operators.md +[prom-functions]: https://github.com/prometheus/prometheus/blob/v3.15.0/docs/querying/functions.md +[prom-ast]: https://github.com/prometheus/prometheus/blob/v3.15.0/promql/parser/ast.go +[prom-signatures]: https://github.com/prometheus/prometheus/blob/v3.15.0/promql/parser/functions.go +[parser-version]: https://github.com/ProjectASAP/promql-parser/blob/9fede7eecca923c9882fe256484d00d37f8706cb/README.md#promql-compliance diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 0f6bfad19..5b7940740 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -45,7 +45,9 @@ on query semantics, accuracy and execution timing. ### 1.1 Unified `Operator` type -Every computation is represented by an `Operator` node. The type has two categories: +Every relation or vector computation is represented by an `Operator` node. +Scalar computations use the companion proposal's `ScalarExpr`. The operator type +has two categories: ```rust enum Operator { @@ -60,11 +62,13 @@ operator kinds. | Category | Meaning | All operations | |---|---|---| -| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlScalarBridge`, `EvalTimestamp`, `PromqlVectorFromScalar`, `PromqlScalarFromVector`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | +| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Values`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlVectorFromScalar`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | | `ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | -`PromqlScalarBridge` keeps its existing name. -`CurrentTimestamp` belongs to scalar expressions, so it is not in this operator list. +`CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to scalar +expressions. Constants need no `PromqlScalarBridge`; query roots may directly hold +a scalar expression. The [companion proposal](decoupling_op_and_expr.md) defines +these boundaries and their SQL/PromQL semantic coverage. `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` are reserved in the proposed planner model; listing them does not establish planner @@ -111,14 +115,18 @@ The following changes are explicit: readability; column IDs and scan schemas below show the resolved form. - `Predicate`, `ProjectItem` and other scalar-bearing payloads keep their roles and own `ScalarExpr` values. The companion proposal defines owned scalar trees and - shared operator inputs, including `Concat`, as well as predicate and scalar-bridge - validation. This document uses those same structures. + shared operator inputs, including `Concat`. Explicit scalar conversions and SQL + subqueries also reference the common `Operator`; their dependencies remain visible. + This document uses those same structures. - `BinaryOp` reuses the existing post-ASAP `BinaryOperator` payload. It carries the binary operation, vector matching and checked-division requirements; the pre-ASAP `op` and `vector_match` semantics must be preserved when mapped into it. + The proposed `return_bool` field also retains PromQL comparison mode. - `Limit.partition_by` comes from the existing post-ASAP limit. Applying that field to the unified operator remains an explicit design choice, not new functionality - implied by a rename. `SQLWindowFunc.frame` retains its current optional form. + implied by a rename. Optional `Limit.n` adds offset-only queries, and + `TimeRange.kind` distinguishes instant selection from a range window, as specified + in the companion. `SQLWindowFunc.frame` retains its current optional form. **Proposed data structures.** These sketches describe operation-specific data using current names. `Rc` represents a shared input edge; no new reference type @@ -131,6 +139,7 @@ enum NonASAPOp { Scan { source: Source, predicates: Vec, schema: Schema, }, + Values { rows: Vec>, schema: Schema }, Filter { child: Rc, pred: Predicate }, Project { child: Rc, cols: Vec, qualifier: Option, @@ -146,19 +155,18 @@ enum NonASAPOp { }, Dedup { child: Rc, cols: Vec }, Sort { child: Rc, keys: Vec, partition_by: GroupKeys }, - Limit { child: Rc, n: usize, offset: usize, partition_by: GroupKeys }, - BinaryOp { lhs: Rc, rhs: Rc, operator: BinaryOperator }, + Limit { child: Rc, n: Option, offset: usize, partition_by: GroupKeys }, + BinaryOp { + lhs: Rc, rhs: Rc, operator: BinaryOperator, return_bool: bool, + }, SQLWindowFunc { child: Rc, func: WindowFuncKind, args: Vec, partition_by: GroupKeys, order_by: Vec, frame: Option, output_name: String, }, - TimeRange { child: Rc, range: Duration }, + TimeRange { child: Rc, range: Duration, kind: TimeRangeKind }, TimeShift { child: Rc, shift: TimeShift }, - PromqlScalarBridge(ScalarExpr), - EvalTimestamp, - PromqlVectorFromScalar(Rc), - PromqlScalarFromVector(Rc), + PromqlVectorFromScalar(ScalarExpr), PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, PromqlInfoEnrich { child: Rc, selector: Vec }, PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, @@ -169,7 +177,8 @@ enum NonASAPOp { The fields describe what an operator does to its inputs: - `child`, `left`, `right`, `lhs`, `rhs` and `children` are graph dependencies. They can lead to - either operator category, subject to the input's schema requirements. + either operator category, subject to the input's schema and result-kind requirements. + Dependencies referenced by scalar conversions/subqueries are edges in this same graph. - `Predicate` describes a row-level condition; `ProjectItem` contains a scalar expression and its optional output alias. - `reduction` describes whether aggregation combines groups or operates per entity; @@ -241,10 +250,11 @@ operator is first created. A filter is an operator because it transforms a table. Its predicate, such as `latency > 100`, is a scalar expression evaluated in that table's schema. -Keep scalar expressions within their owning operators. Graph edges then represent -computation dependencies, while predicates, projection expressions and sort keys -describe how an operator processes its input. This prevents an expression from -being mistaken for a table-producing plan. The +Scalar expressions belong to an operator field or a scalar query root. Predicates, +projection expressions and sort keys describe value computation in that context. +Explicit scalar conversions and subqueries may reference operators; those are +visible graph dependencies with defined cardinality rules. This prevents an +arbitrary expression from being mistaken for a table-producing plan. The [companion proposal](decoupling_op_and_expr.md) defines this distinction. ### 1.3 Scope of operator sharing @@ -328,7 +338,9 @@ accuracy, costing or selection policies. ## 4. Export preserves the graph Export one node per operator and represent its input dependencies as edges. Export -a shared producer once, with edges to all its consumers. +a shared producer once, with edges to all its consumers. Include dependencies +referenced by scalar conversions and subqueries. Preserve scalar query roots as +expressions and their operator dependencies; do not invent a bridge node for export. This keeps the graph visible to costing, physical compilation, execution and plan inspection. Embedding a whole relational subtree in one exported node would hide From 9763131fd58a5252e14b7018dc85831d0b0bbed8 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:14:48 +0000 Subject: [PATCH 34/46] docs: align unified graph with scalar query dependencies --- .../design_docs/proposals/operator-sharing.md | 92 ++++++++++++++----- 1 file changed, 69 insertions(+), 23 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 5b7940740..34cf9e322 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -45,15 +45,20 @@ on query semantics, accuracy and execution timing. ### 1.1 Unified `Operator` type -Every relation or vector computation is represented by an `Operator` node. -Scalar computations use the companion proposal's `ScalarExpr`. The operator type -has two categories: +Relation, vector and summary-state computations are represented by `Operator` +nodes. Scalar computations use the companion proposal's `ScalarExpr`. The operator +type has two categories: ```rust enum Operator { NonASAP(NonASAPOp), ASAP(ASAPOp), } + +enum QueryRoot { + Operator(Rc), + Scalar(ScalarExpr), +} ``` The table lists all operator kinds in this proposal. Aggregate functions, join @@ -68,7 +73,8 @@ operator kinds. `CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to scalar expressions. Constants need no `PromqlScalarBridge`; query roots may directly hold a scalar expression. The [companion proposal](decoupling_op_and_expr.md) defines -these boundaries and their SQL/PromQL semantic coverage. +these boundaries and their SQL/PromQL semantic coverage. A scalar-only frontend +query, such as `time()`, therefore needs no operator node. `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` are reserved in the proposed planner model; listing them does not establish planner @@ -77,8 +83,8 @@ proposal (§6). `NonASAPOp` and `ASAPOp` describe the operation performed by a node. Inputs in both categories connect to `Operator` nodes, so either category can consume the other -when their schema and execution constraints permit it. `NonASAP` classifies one -node; it does not require all of that node's descendants to be non-ASAP. +when their result kind, schema and execution constraints permit it. `NonASAP` +classifies one node; it does not require all of that node's descendants to be non-ASAP. Both categories use the same graph model. A projection can consume a summary estimate, and a summary can consume the result of a filter or join. There is no @@ -103,10 +109,11 @@ the existing query and summary models, but the sketches combine several changes: | Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | | Represent execution timing chosen during physical planning | Follow the planning-stage design in [#509](https://github.com/ProjectASAP/ASAPPlanner/pull/509), rather than introduce a new timing policy here (§2.3). | -**Naming and compatibility.** The sketches retain current operation, field and -payload-type names. `NonASAPOp`, `ASAPOp` and the companion proposal's `ScalarExpr` -are the new structural concepts; ordinary payloads such as `Predicate`, -`ProjectItem`, `AggIntent`, `SketchQuery` and `SummaryFamilyType` keep their names. +**Naming and compatibility.** Reuse current names where their semantics match. +The companion proposal defines deliberate additions and changes, including +`Values`, scalar query roots and typed scalar conversions. Ordinary payloads such +as `Predicate`, `ProjectItem`, `AggIntent`, `SketchQuery` and `SummaryFamilyType` +keep their names; retaining a name does not establish complete language coverage. The following changes are explicit: @@ -228,8 +235,8 @@ The summary fields distinguish state construction and readout: | `query` / `readout` | The result requested from summary or maintained-population state | | `population` | The population whose membership and values are maintained | -Schema, derived accuracy and execution timing describe every `Operator`, regardless -of category (§2). They are omitted from these operation-specific sketches. Timing +Result kind, schema, derived accuracy and execution timing describe `Operator` +nodes in both categories (§2). They are omitted from these operation-specific sketches. Timing comes from lifecycle planning rather than a fixed field value implied by an operator kind; the final accuracy assessment combines local evidence with the actual inputs. @@ -257,6 +264,22 @@ visible graph dependencies with defined cardinality rules. This prevents an arbitrary expression from being mistaken for a table-producing plan. The [companion proposal](decoupling_op_and_expr.md) defines this distinction. +Its pre-ASAP `Rc>` references become `Rc` in the unified +model, including `QueryRoot` and the inputs to `PromqlScalarFromVector`, +`ScalarSubquery`, `Exists` and `InSubquery`. Their cardinality, NULL and NaN rules +remain unchanged. + +For example, in `scalar(sum(up))`, `sum(up)` is an operator subgraph producing an +instant vector. The scalar expression `PromqlScalarFromVector` references its +result to obtain one number. A valid ASAP rewrite may replace that producer with +a summary readout, preserving the required vector and accuracy semantics; it cannot +substitute raw summary state. Ordinary expressions such as `price * 2` reference +columns and literals, not a query subgraph. + +These are **query subgraphs referenced by scalar expressions**. The reference is +an edge to a producer in the common graph, not a copy of the subgraph embedded in +the expression. Scalar ownership does not change the producer's graph identity. + ### 1.3 Scope of operator sharing Here, sharing means pre-ASAP and post-ASAP use the same operator definitions. @@ -268,15 +291,17 @@ it does not introduce rules for sharing computations across queries. | Property | Meaning | How it is determined | |---|---|---| -| Output schema | What the node produces: field names, types and relevant identity/time information | From the operation and its inputs | +| Result kind and output schema | Whether the node produces a relation, instant vector, range vector or state, and its fields, types and identity/time information | From the operation and its inputs | | Accuracy guarantee | What can be established about the result's accuracy | From local accuracy evidence and the guarantees of its inputs | | Execution timing | Whether work runs at ingestion time or query time | From a lifecycle choice for the complete plan | ### 2.1 One schema model for values and state -A common graph needs a common description of its edges. Schemas must distinguish -ordinary values from summary state, so a consumer can determine whether an input is -usable. +A common graph needs a common description of its edges. Validate result kind as +well as schema: matching numeric columns do not make a relation, an instant vector +and a range vector interchangeable. Summary state is also distinct from ordinary +values. The operation determines the applicable input/output contract; this does +not require adding the same result-kind field to every node. For example, a KLL build produces state; its p99 estimation produces a numeric value. A numeric predicate can consume the estimate, but cannot treat the KLL state itself @@ -286,13 +311,16 @@ permit it; exact aggregate state must be finalized before use as an ordinary val Schema information must preserve grouping fields, time information, uniqueness and series identity where relevant. Sharing operators must not change SQL or PromQL meaning. After a rewrite, schemas must describe the new inputs rather than the plan -that was replaced. +that was replaced. Scalar expressions instead have value types and evaluation +contexts. A scalar root does not need a fabricated relation schema. ### 2.2 Preserve existing accuracy semantics The unified representation must preserve the existing accuracy model, composition rules and result guarantees. An operation's guarantee must still account for its -actual inputs; unknown accuracy must not be treated as exactness. +actual inputs, including producers referenced by scalar expressions. Reading an +approximate result through `PromqlScalarFromVector` or a SQL scalar subquery does +not make it exact; unknown accuracy must not be treated as exactness. This proposal adds no accuracy fields or new guarantee-calculation workflow. Changing the operator representation must not change the accuracy meaning of the @@ -313,6 +341,12 @@ The representation must preserve the resulting execution constraints: ingestion- work cannot depend on query-time results, and consumers must receive values or state that are available when needed. Materialization choices, retention and plan selection remain governed by #509; this document does not define another lifecycle policy. +These constraints also apply to query subgraphs referenced by scalar expressions. + +PromQL evaluation timestamps and SQL statement time are separate from these +execution phases. `TimeShift`, subquery grids and `EvalTimestamp` retain their +source-language evaluation context. A shared node identity alone does not permit +reusing a result across different evaluation times. ## 3. Planning responsibilities @@ -322,7 +356,7 @@ how they use the common operator model: | Stage from #509 | Use of the unified representation | |---|---| -| Frontends | Produce a graph containing only `NonASAP` operators, preserving source-language semantics. | +| Frontends | Produce an operator or scalar query root; any referenced operator nodes are `NonASAP`. Preserve source-language semantics. | | Logical ASAP-aware optimization | Form candidate graphs containing ordinary and summary operators, with no wrappers hiding their dependencies. | | Physical ASAP-aware optimization | Determine executable alternatives, including materialization and execution timing, for those candidate graphs. | | Plan selection | Evaluate complete physical candidates using workload requirements and deployment-provided models and capabilities. | @@ -346,9 +380,9 @@ This keeps the graph visible to costing, physical compilation, execution and pla inspection. Embedding a whole relational subtree in one exported node would hide its internal sharing and recreate the original boundary problem. -Export carries the resolved schemas, assessed guarantees and assigned execution -phases. Physical compilation may lower one logical operation to several physical -operations, but must preserve its dependencies and meaning. The execution layer +Export preserves result kinds, resolved schemas, scalar value types and evaluation +context, together with the applicable guarantees and assigned execution phases. +Physical compilation may lower one logical operation to several physical operations, but must preserve its dependencies and meaning. The execution layer does not invent missing planning decisions. Changing the exported representation requires coordinated adoption by the planner @@ -365,6 +399,10 @@ The design is successful when: any shared inputs; it does not introduce new sharing rules. - Existing value/state, accuracy and execution constraints remain enforceable on the unified representation. +- Scalar query roots and conversions use the same representation before and after + optimization, with no bridge nodes or hidden subplans. +- Validation includes result kind and query subgraphs referenced by scalar + expressions when checking schema, accuracy and execution constraints. - Export preserves visible dependencies and shared producers. ## 6. Scope and compatibility @@ -377,7 +415,15 @@ The pre-ASAP and post-ASAP versions of some operations carry different informati The unified `BinaryOp` must retain existing checked-division requirements, and `Limit` must retain the existing ability to limit within groups. These compatibility requirements belong in this design because removing duplicate operator definitions -must not remove existing behavior. +must not remove existing behavior. If a rewrite moves arithmetic into a scalar +expression, it must preserve applicable checked-division guards and exact fallback; +`ExprSemantics` selects language rules and does not replace those proof conditions. + +The [companion semantic tables](decoupling_op_and_expr.md#3-semantic-requirements) +use DataFusion 55.1.0 and Prometheus 3.15.0 as design targets. Their gaps also apply +here: a common `Operator` type does not supply missing aggregate modifiers, value +types or function contracts. This proposal changes neither repository dependencies +nor the set of implemented language features. New accuracy fields, accuracy-composition rules, computation-sharing algorithms and lifecycle policies are outside this proposal. Planning responsibilities follow #509. From 774908d2caa94932aca1766af39e19d3233350db Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:16:22 +0000 Subject: [PATCH 35/46] docs: identify PromQL built-in scalar conversion in examples --- docs/design_docs/proposals/decoupling_op_and_expr.md | 4 ++++ docs/design_docs/proposals/operator-sharing.md | 6 ++++-- 2 files changed, 8 insertions(+), 2 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index b1cc1378c..2644508c0 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -238,6 +238,10 @@ scalar(sum(up)) + 1 → Scalar root: Arithmetic( PromqlScalarFromVector(Aggregate(...)), Literal(1.0)) ``` +Here, `scalar()` and `vector()` are Prometheus PromQL built-in conversion +functions explicitly present in the query, not wrappers inserted by this proposal. +`sum(up)` alone remains a valid instant-vector query. + A PromQL projection retains the time and label fields required by the operation and applies its metric-name rules; it does not project only the numeric sample. For an open label schema, lowering must retain the complete series identity, diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 34cf9e322..3de7e4ab2 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -269,8 +269,10 @@ model, including `QueryRoot` and the inputs to `PromqlScalarFromVector`, `ScalarSubquery`, `Exists` and `InSubquery`. Their cardinality, NULL and NaN rules remain unchanged. -For example, in `scalar(sum(up))`, `sum(up)` is an operator subgraph producing an -instant vector. The scalar expression `PromqlScalarFromVector` references its +In `scalar(sum(up))`, `scalar()` is Prometheus PromQL's built-in vector-to-scalar +function, explicitly written by the query author. This proposal does not insert +it automatically: `sum(up)` alone is a valid query returning an instant vector. +The scalar expression `PromqlScalarFromVector` represents that function and references its result to obtain one number. A valid ASAP rewrite may replace that producer with a summary readout, preserving the required vector and accuracy semantics; it cannot substitute raw summary state. Ordinary expressions such as `price * 2` reference From 2d94f9d4e666fb6a5a0f4895fc06fb5e48bc998c Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:20:38 +0000 Subject: [PATCH 36/46] docs: define common operator properties and schema interfaces --- .../design_docs/proposals/operator-sharing.md | 196 +++++++++++++++--- 1 file changed, 163 insertions(+), 33 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 3de7e4ab2..1b0f780ec 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -47,10 +47,10 @@ on query semantics, accuracy and execution timing. Relation, vector and summary-state computations are represented by `Operator` nodes. Scalar computations use the companion proposal's `ScalarExpr`. The operator -type has two categories: +node has common fields defined in §2 and an `expr` payload with two categories: ```rust -enum Operator { +enum OperatorExpr { NonASAP(NonASAPOp), ASAP(ASAPOp), } @@ -67,8 +67,8 @@ operator kinds. | Category | Meaning | All operations | |---|---|---| -| `NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Values`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlVectorFromScalar`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | -| `ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | +| `OperatorExpr::NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Values`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlVectorFromScalar`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | +| `OperatorExpr::ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | `CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to scalar expressions. Constants need no `PromqlScalarBridge`; query roots may directly hold @@ -235,10 +235,9 @@ The summary fields distinguish state construction and readout: | `query` / `readout` | The result requested from summary or maintained-population state | | `population` | The population whose membership and values are maintained | -Result kind, schema, derived accuracy and execution timing describe `Operator` -nodes in both categories (§2). They are omitted from these operation-specific sketches. Timing -comes from lifecycle planning rather than a fixed field value implied by an operator -kind; the final accuracy assessment combines local evidence with the actual inputs. +The payloads above describe operations and their inputs. The common `Operator` +fields and their derivation interfaces are defined once in §2; individual variants +do not repeat schema, accuracy or execution timing. For example, arrows below show data flowing from producer to consumer: @@ -291,30 +290,157 @@ it does not introduce rules for sharing computations across queries. ## 2. Node properties and why they differ -| Property | Meaning | How it is determined | +Both operation categories use this resolved node structure. It follows the current +`SummaryNode` separation between `expr`, `schema` and `guarantee`, generalized to +all operators. The former `Operator` category enum becomes `OperatorExpr` (§1.1) +so common fields do not have to be repeated in every variant. + +```rust +struct Operator { + expr: OperatorExpr, + result_kind: OperatorResultKind, + schema: Schema, + guarantee: Option, + timing: Option, +} + +// Existing enum; the node's Option represents an unassigned phase. +enum ExecutionTiming { + IngestionTime, + QueryTime, +} +``` + +| Field | Meaning | How it is determined | |---|---|---| -| Result kind and output schema | Whether the node produces a relation, instant vector, range vector or state, and its fields, types and identity/time information | From the operation and its inputs | -| Accuracy guarantee | What can be established about the result's accuracy | From local accuracy evidence and the guarantees of its inputs | -| Execution timing | Whether work runs at ingestion time or query time | From a lifecycle choice for the complete plan | +| `expr` | Operation category, parameters and dependencies | `OperatorExpr`, `NonASAPOp` and `ASAPOp` in §1 | +| `result_kind`, `schema` | The output category and fields, including identity/time metadata | Derived from `expr` and its actual inputs, then retained on the resolved node (§2.1) | +| `guarantee` | An established result-accuracy guarantee, when available | Existing `ResultGuarantee` and composition rules (§2.2); `None` never means exact | +| `timing` | The assigned ingestion/query execution phase | Physical planning under #509 (§2.3); `None` means not assigned | + +`ResultGuarantee` retains its existing definition. `OperatorExpr`, +`OperatorResultKind` and the common node layout are proposed; `Schema` is unified +as specified below. This is a resolved-plan interface: name resolution must finish +before producing these concrete `ColumnId`/`Schema` nodes. + +| Plan stage | Required property state | +|---|---| +| Resolved frontend / logical candidate | Valid `result_kind` and `schema`; `guarantee` only where established; `timing` may be `None`. | +| Executable physical candidate | Valid output metadata, accuracy acceptable under the existing requirements, and `Some(timing)` for every executable operator. | + +Changing an operation or dependency requires re-deriving its output metadata and +revalidating dependent guarantees and timing assignments. Derived fields must not +retain facts from the plan that was replaced. This defines consistency, not a new +caching or mutation mechanism. ### 2.1 One schema model for values and state -A common graph needs a common description of its edges. Validate result kind as -well as schema: matching numeric columns do not make a relation, an instant vector -and a range vector interchangeable. Summary state is also distinct from ordinary -values. The operation determines the applicable input/output contract; this does -not require adding the same result-kind field to every node. +Use one `Schema` for operator outputs before and after optimization. Reuse the +existing `SummaryFamilyType` to distinguish ordinary values from state, and retain +the current `Schema` metadata. The following is the proposed resolved interface; +it is not the current Rust definition. + +```rust +struct Column { + name: String, + dtype: T, + nullable: bool, + table: Option, +} + +struct Schema { + columns: Vec>, + time_index: Option, + unique_keys: Vec>, + closed: bool, +} + +// Existing variants and payload names, reused without renaming. +enum SummaryFamilyType { + Plain(DataType), + ExactAggregate(ExactKind, ExactParams), + Sketch(SketchKind, GroupingStrategy), + Sample(SamplingKind, SamplingParams), + Wavelet(WaveletKind, WaveletParams), + StatModel(StatModelKind, StatModelParams), +} + +// Proposed derived output classification, separate from column types. +enum OperatorResultKind { + Relation, + InstantVector, + RangeVector, + State, +} + +impl OperatorExpr { + fn output_schema(&self) -> Result; + fn output_kind(&self) -> Result; + fn validate_inputs(&self) -> Result<(), QueryExprError>; +} -For example, a KLL build produces state; its p99 estimation produces a numeric value. -A numeric predicate can consume the estimate, but cannot treat the KLL state itself -as a number. Ordinary operators may carry state through only where their semantics -permit it; exact aggregate state must be finalized before use as an ordinary value. +impl Operator { + fn validate(&self) -> Result<(), QueryExprError>; +} -Schema information must preserve grouping fields, time information, uniqueness and -series identity where relevant. Sharing operators must not change SQL or PromQL -meaning. After a rewrite, schemas must describe the new inputs rather than the plan -that was replaced. Scalar expressions instead have value types and evaluation -contexts. A scalar root does not need a fabricated relation schema. +impl ScalarExpr { + fn scalar_type(&self, input: &Schema) -> Result<(DataType, bool), QueryExprError>; +} +``` + +**Relationship to current types.** `Column` gains a type parameter so the common +operator schema can use `SummaryFamilyType`, while existing nested value types +such as `DataType::List` still use ordinary `Column`. The proposed common +`Schema` replaces the separate operator-edge roles of pre-ASAP `Schema` and +post-ASAP `SummarySchema` / `SummaryField`; it does not rename `DataType` or add a +second summary-family enum. A pre-ASAP value column becomes `Plain(dtype)`. +Frontend validation permits only ordinary value columns, preserving the current +pre-ASAP restriction even though the common schema can also express state. + +| Field | Meaning and requirement | +|---|---| +| `columns` | Ordered named fields. `Plain(DataType)` is a readable value; other variants retain the identity and parameters of summary or exact-accumulator state. | +| `Column.nullable`, `Column.table` | Preserve SQL nullability and qualified column resolution. | +| `time_index` | Identifies the time column when present; it does not by itself distinguish an instant vector from a range vector. | +| `unique_keys` | Proven column combinations identifying rows; an empty list asserts no known key. Recompute these proofs when a rewrite changes identity. | +| `closed` | Whether `columns` completely describes the output. An open PromQL schema must retain unlisted labels through the existing complete-series-identity contract. | + +`OperatorResultKind` is derived from the operation and its inputs and retained as +`Operator.result_kind`. `State` describes an output carrying unfinalized state; its +schema may also contain ordinary grouping keys. `SummaryEstimate`, +`FinalizeExactAccumulator` and other readouts derive the appropriate relation or +vector kind from their operation and input context. Matching numeric columns do +not make those kinds interchangeable. + +**Interface contracts.** `OperatorExpr::output_schema` derives the fields and +metadata for the actual inputs after a rewrite; `output_kind` derives the result +category. `validate_inputs` checks producer/consumer compatibility, including query +subgraphs referenced by scalar expressions. `Operator::validate` additionally +checks that retained output metadata agrees with that derivation and that any +guarantee or timing assignment is valid under the existing rules. For example, +`PromqlScalarFromVector` requires an instant vector, and a summary readout requires the compatible state family. These checks +must succeed before a plan is accepted; sharing an enum does not make every +producer/consumer combination legal. + +`scalar_type` keeps the existing method name and `(DataType, nullable)` result. +Its `input` is the applicable column scope: the child schema for a projection, +both input schemas for a join predicate, or aggregate outputs for `HAVING`. +Explicit subquery/conversion expressions validate their referenced producer using +the contracts above. Numeric expressions cannot consume state columns as numbers. +A scalar root is checked with an empty column scope and needs no fabricated +relation output schema. `QueryExprError` retains the existing error-type name; +result-kind and state-family mismatches require corresponding validation errors. + +For example, a KLL build outputs `State` with a +`Sketch(SketchKind, GroupingStrategy)` column identifying KLL and its parameters. +Its p99 readout outputs an ordinary `Plain(Float64)` column in the appropriate +relation/vector schema. A numeric predicate can use that readout, but not the KLL +state. Exact accumulator state similarly requires `FinalizeExactAccumulator`. +An ordinary operator may pass state through only where its input/output contract +permits it. A bare-column projection can preserve the column's `SummaryFamilyType` +directly during `output_schema` derivation; `scalar_type` applies when that column +is used as a scalar value and rejects state. Copying a state column does not turn +it into a readable scalar. ### 2.2 Preserve existing accuracy semantics @@ -324,9 +450,11 @@ actual inputs, including producers referenced by scalar expressions. Reading an approximate result through `PromqlScalarFromVector` or a SQL scalar subquery does not make it exact; unknown accuracy must not be treated as exactness. -This proposal adds no accuracy fields or new guarantee-calculation workflow. -Changing the operator representation must not change the accuracy meaning of the -same computation. +The common node reuses `guarantee: Option` from `SummaryNode`. +`Some` records an established guarantee; `None` covers an unassessed or unknown +result, or state whose accuracy is only established at readout. Exactness must be +explicitly established using the existing model. This proposal introduces no new +accuracy metric or guarantee-calculation workflow. ### 2.3 Timing follows the planning-stage design @@ -335,9 +463,11 @@ separates logical decisions about what to compute from physical decisions about and when to compute it. This proposal follows that division. For example, a KLL summary build may execute at ingestion time or query time, -depending on the materialization choice. The unified operator representation must -carry the chosen execution timing without requiring separate operator definitions -for the two phases. +depending on the materialization choice. `Operator.timing` records that assignment +as `Some(ExecutionTiming::IngestionTime)` or `Some(ExecutionTiming::QueryTime)`. +Logical nodes may retain `None`; physical-plan validation must reject unassigned +executable nodes. Timing is common node metadata rather than a separate payload +field on selected `ASAPOp` variants. The representation must preserve the resulting execution constraints: ingestion-time work cannot depend on query-time results, and consumers must receive values or state @@ -427,7 +557,7 @@ here: a common `Operator` type does not supply missing aggregate modifiers, valu types or function contracts. This proposal changes neither repository dependencies nor the set of implemented language features. -New accuracy fields, accuracy-composition rules, computation-sharing algorithms and +New accuracy metrics, accuracy-composition rules, computation-sharing algorithms and lifecycle policies are outside this proposal. Planning responsibilities follow #509. The scalar/operator separation is specified in the companion document. Storage, traversal algorithms, serialization fields and a code migration sequence are also From 30a805eeb033eb51470a4498d5c5ee91c571adbf Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:23:12 +0000 Subject: [PATCH 37/46] docs: distinguish operators, nodes and scalar query roots --- .../proposals/decoupling_op_and_expr.md | 8 +- .../design_docs/proposals/operator-sharing.md | 118 +++++++++--------- 2 files changed, 67 insertions(+), 59 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index 2644508c0..0813fd43e 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -212,12 +212,14 @@ that this proposal does not define (§3.4). ### 2.3 Query roots and composition without a bridge -A query root selects one of the two representations; this tag adds no computation: +A query root selects one of the two representations. The `Operator` variant +references the query's producer; `ScalarExpr` holds its scalar expression. This +enum adds no computation node: ```rust enum QueryRoot { Operator(Rc>), - Scalar(ScalarExpr), + ScalarExpr(ScalarExpr), } ``` @@ -249,7 +251,7 @@ including unreferenced labels. If the input provides neither a complete label schema nor a full identity value, this lowering is not valid. `vector(s)` remains a real conversion to a one-element, label-free vector. When combined with the [operator-sharing proposal](operator-sharing.md), all plan -references above target the common `Operator`, including references inside scalar +references above target the common `OperatorNode`, including references inside scalar expressions and query roots. Scalar expression trees do not acquire shared-node identity. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 1b0f780ec..e6b53ee3f 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -45,30 +45,36 @@ on query semantics, accuracy and execution timing. ### 1.1 Unified `Operator` type -Relation, vector and summary-state computations are represented by `Operator` -nodes. Scalar computations use the companion proposal's `ScalarExpr`. The operator -node has common fields defined in §2 and an `expr` payload with two categories: +The two computational categories are `Operator` and `ScalarExpr`. `Operator` +describes a relation, vector or summary-state operation and has two variants. +`OperatorNode` (§2) stores that operation together with its common properties. +`QueryRoot` selects either an operator node or a scalar expression as the query +entry point; it is an enum, not an additional computation node: ```rust -enum OperatorExpr { +enum Operator { NonASAP(NonASAPOp), ASAP(ASAPOp), } enum QueryRoot { - Operator(Rc), - Scalar(ScalarExpr), + Operator(Rc), + ScalarExpr(ScalarExpr), } ``` +`QueryRoot::Operator` is used for SQL query results and PromQL vector queries, +such as `sum(up)`. `QueryRoot::ScalarExpr` is used for PromQL scalar queries, such +as `2`, `time()` or `scalar(sum(up))`. Neither variant adds a wrapper operator. + The table lists all operator kinds in this proposal. Aggregate functions, join kinds and scalar functions are choices within these operations, not additional operator kinds. | Category | Meaning | All operations | |---|---|---| -| `OperatorExpr::NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Values`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlVectorFromScalar`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | -| `OperatorExpr::ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | +| `Operator::NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Values`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlVectorFromScalar`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | +| `Operator::ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | `CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to scalar expressions. Constants need no `PromqlScalarBridge`; query roots may directly hold @@ -82,7 +88,7 @@ or runtime support. Defining new summary-composition semantics is outside this proposal (§6). `NonASAPOp` and `ASAPOp` describe the operation performed by a node. Inputs in both -categories connect to `Operator` nodes, so either category can consume the other +categories connect to `OperatorNode` instances, so either category can consume the other when their result kind, schema and execution constraints permit it. `NonASAP` classifies one node; it does not require all of that node's descendants to be non-ASAP. @@ -104,7 +110,7 @@ the existing query and summary models, but the sketches combine several changes: | Change | Purpose and scope | |---|---| | Group ordinary operations under `NonASAP` and summary operations under `ASAP` | Organize the common operator model into two semantic categories. | -| Give both categories inputs that refer to `Operator` nodes | Allow ordinary and summary operations to compose directly, without a wrapper hiding their dependencies. | +| Give both categories inputs that refer to `OperatorNode` instances | Allow ordinary and summary operations to compose directly, without a wrapper hiding their dependencies. | | Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | | Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | | Represent execution timing chosen during physical planning | Follow the planning-stage design in [#509](https://github.com/ProjectASAP/ASAPPlanner/pull/509), rather than introduce a new timing policy here (§2.3). | @@ -117,13 +123,13 @@ keep their names; retaining a name does not establish complete language coverage The following changes are explicit: -- Operator inputs become references to the common `Operator`. The sketches use the +- Operator inputs become references to the common `OperatorNode`. The sketches use the existing `Rc` notation for shared inputs and omit column-state generics for readability; column IDs and scan schemas below show the resolved form. - `Predicate`, `ProjectItem` and other scalar-bearing payloads keep their roles and own `ScalarExpr` values. The companion proposal defines owned scalar trees and shared operator inputs, including `Concat`. Explicit scalar conversions and SQL - subqueries also reference the common `Operator`; their dependencies remain visible. + subqueries also reference the common `OperatorNode`; their dependencies remain visible. This document uses those same structures. - `BinaryOp` reuses the existing post-ASAP `BinaryOperator` payload. It carries the binary operation, vector matching and checked-division requirements; the pre-ASAP @@ -136,7 +142,7 @@ The following changes are explicit: in the companion. `SQLWindowFunc.frame` retains its current optional form. **Proposed data structures.** These sketches describe operation-specific data using -current names. `Rc` represents a shared input edge; no new reference type +current names. `Rc` represents a shared input edge; no new reference type is introduced. Storage and traversal algorithms remain outside this design. `NonASAPOp` retains the query semantics needed before and after optimization: @@ -147,37 +153,37 @@ enum NonASAPOp { source: Source, predicates: Vec, schema: Schema, }, Values { rows: Vec>, schema: Schema }, - Filter { child: Rc, pred: Predicate }, + Filter { child: Rc, pred: Predicate }, Project { - child: Rc, cols: Vec, qualifier: Option, + child: Rc, cols: Vec, qualifier: Option, }, Aggregate { - child: Rc, reduction: Reduction, measures: Vec, + child: Rc, reduction: Reduction, measures: Vec, output_names: Vec, having: Option, }, - Join { left: Rc, right: Rc, kind: JoinKind, pred: Predicate }, - SetOp { left: Rc, right: Rc, kind: RelationalSetOpKind, all: bool }, + Join { left: Rc, right: Rc, kind: JoinKind, pred: Predicate }, + SetOp { left: Rc, right: Rc, kind: RelationalSetOpKind, all: bool }, Concat { - children: Vec>, discriminator_unique_key: Option, + children: Vec>, discriminator_unique_key: Option, }, - Dedup { child: Rc, cols: Vec }, - Sort { child: Rc, keys: Vec, partition_by: GroupKeys }, - Limit { child: Rc, n: Option, offset: usize, partition_by: GroupKeys }, + Dedup { child: Rc, cols: Vec }, + Sort { child: Rc, keys: Vec, partition_by: GroupKeys }, + Limit { child: Rc, n: Option, offset: usize, partition_by: GroupKeys }, BinaryOp { - lhs: Rc, rhs: Rc, operator: BinaryOperator, return_bool: bool, + lhs: Rc, rhs: Rc, operator: BinaryOperator, return_bool: bool, }, SQLWindowFunc { - child: Rc, func: WindowFuncKind, args: Vec, + child: Rc, func: WindowFuncKind, args: Vec, partition_by: GroupKeys, order_by: Vec, frame: Option, output_name: String, }, - TimeRange { child: Rc, range: Duration, kind: TimeRangeKind }, - TimeShift { child: Rc, shift: TimeShift }, + TimeRange { child: Rc, range: Duration, kind: TimeRangeKind }, + TimeShift { child: Rc, shift: TimeShift }, PromqlVectorFromScalar(ScalarExpr), - PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, - PromqlInfoEnrich { child: Rc, selector: Vec }, - PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, - PromqlSubquery { child: Rc, range: Duration, resolution: Option }, + PromqlRelabel { child: Rc, dst: String, value: ScalarExpr }, + PromqlInfoEnrich { child: Rc, selector: Vec }, + PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, + PromqlSubquery { child: Rc, range: Duration, resolution: Option }, } ``` @@ -203,24 +209,24 @@ summary or exact-accumulator cases, never its `Plain` case. ```rust enum ASAPOp { SummaryAgg { - child: Rc, family: SummaryFamilyType, input: SummaryUpdate, + child: Rc, family: SummaryFamilyType, input: SummaryUpdate, reduction: Reduction, grouping: GroupingStrategy, }, SummaryEstimate { - summary_input: Rc, query: SketchQuery, + summary_input: Rc, query: SketchQuery, }, - FinalizeExactAccumulator { child: Rc }, - MaintainPopulation { child: Rc, population: MaintainedPopulation }, - ReadPopulation { child: Rc, readout: PopulationReadout }, + FinalizeExactAccumulator { child: Rc }, + MaintainPopulation { child: Rc, population: MaintainedPopulation }, + ReadPopulation { child: Rc, readout: PopulationReadout }, // Reserved operations; semantics and support require further design. - SummaryMerge { children: Vec> }, - SummarySubtract { left: Rc, right: Rc }, - SummaryDelete { summary_input: Rc, key: ColumnRef }, + SummaryMerge { children: Vec> }, + SummarySubtract { left: Rc, right: Rc }, + SummaryDelete { summary_input: Rc, key: ColumnRef }, SummaryJoin { - outer: Rc, inner: Rc, key: ColumnRef, family: SummaryFamilyType, + outer: Rc, inner: Rc, key: ColumnRef, family: SummaryFamilyType, }, - Extension { child: Rc, name: String }, + Extension { child: Rc, name: String }, } ``` @@ -235,7 +241,7 @@ The summary fields distinguish state construction and readout: | `query` / `readout` | The result requested from summary or maintained-population state | | `population` | The population whose membership and values are maintained | -The payloads above describe operations and their inputs. The common `Operator` +The payloads above describe operations and their inputs. The common `OperatorNode` fields and their derivation interfaces are defined once in §2; individual variants do not repeat schema, accuracy or execution timing. @@ -263,7 +269,7 @@ visible graph dependencies with defined cardinality rules. This prevents an arbitrary expression from being mistaken for a table-producing plan. The [companion proposal](decoupling_op_and_expr.md) defines this distinction. -Its pre-ASAP `Rc>` references become `Rc` in the unified +Its pre-ASAP `Rc>` references become `Rc` in the unified model, including `QueryRoot` and the inputs to `PromqlScalarFromVector`, `ScalarSubquery`, `Exists` and `InSubquery`. Their cardinality, NULL and NaN rules remain unchanged. @@ -291,13 +297,13 @@ it does not introduce rules for sharing computations across queries. ## 2. Node properties and why they differ Both operation categories use this resolved node structure. It follows the current -`SummaryNode` separation between `expr`, `schema` and `guarantee`, generalized to -all operators. The former `Operator` category enum becomes `OperatorExpr` (§1.1) -so common fields do not have to be repeated in every variant. +`SummaryNode` separation between an operation and its metadata, generalized to +all operators. The field is named `operator` because it holds `Operator` (§1.1), +not a scalar expression. Common fields are defined here once, not in each variant. ```rust -struct Operator { - expr: OperatorExpr, +struct OperatorNode { + operator: Operator, result_kind: OperatorResultKind, schema: Schema, guarantee: Option, @@ -313,12 +319,12 @@ enum ExecutionTiming { | Field | Meaning | How it is determined | |---|---|---| -| `expr` | Operation category, parameters and dependencies | `OperatorExpr`, `NonASAPOp` and `ASAPOp` in §1 | -| `result_kind`, `schema` | The output category and fields, including identity/time metadata | Derived from `expr` and its actual inputs, then retained on the resolved node (§2.1) | +| `operator` | Operation category, parameters and dependencies | `Operator`, `NonASAPOp` and `ASAPOp` in §1 | +| `result_kind`, `schema` | The output category and fields, including identity/time metadata | Derived from `operator` and its actual inputs, then retained on the resolved node (§2.1) | | `guarantee` | An established result-accuracy guarantee, when available | Existing `ResultGuarantee` and composition rules (§2.2); `None` never means exact | | `timing` | The assigned ingestion/query execution phase | Physical planning under #509 (§2.3); `None` means not assigned | -`ResultGuarantee` retains its existing definition. `OperatorExpr`, +`ResultGuarantee` retains its existing definition. `Operator`, `OperatorNode`, `OperatorResultKind` and the common node layout are proposed; `Schema` is unified as specified below. This is a resolved-plan interface: name resolution must finish before producing these concrete `ColumnId`/`Schema` nodes. @@ -373,13 +379,13 @@ enum OperatorResultKind { State, } -impl OperatorExpr { +impl Operator { fn output_schema(&self) -> Result; fn output_kind(&self) -> Result; fn validate_inputs(&self) -> Result<(), QueryExprError>; } -impl Operator { +impl OperatorNode { fn validate(&self) -> Result<(), QueryExprError>; } @@ -406,16 +412,16 @@ pre-ASAP restriction even though the common schema can also express state. | `closed` | Whether `columns` completely describes the output. An open PromQL schema must retain unlisted labels through the existing complete-series-identity contract. | `OperatorResultKind` is derived from the operation and its inputs and retained as -`Operator.result_kind`. `State` describes an output carrying unfinalized state; its +`OperatorNode.result_kind`. `State` describes an output carrying unfinalized state; its schema may also contain ordinary grouping keys. `SummaryEstimate`, `FinalizeExactAccumulator` and other readouts derive the appropriate relation or vector kind from their operation and input context. Matching numeric columns do not make those kinds interchangeable. -**Interface contracts.** `OperatorExpr::output_schema` derives the fields and +**Interface contracts.** `Operator::output_schema` derives the fields and metadata for the actual inputs after a rewrite; `output_kind` derives the result category. `validate_inputs` checks producer/consumer compatibility, including query -subgraphs referenced by scalar expressions. `Operator::validate` additionally +subgraphs referenced by scalar expressions. `OperatorNode::validate` additionally checks that retained output metadata agrees with that derivation and that any guarantee or timing assignment is valid under the existing rules. For example, `PromqlScalarFromVector` requires an instant vector, and a summary readout requires the compatible state family. These checks @@ -463,7 +469,7 @@ separates logical decisions about what to compute from physical decisions about and when to compute it. This proposal follows that division. For example, a KLL summary build may execute at ingestion time or query time, -depending on the materialization choice. `Operator.timing` records that assignment +depending on the materialization choice. `OperatorNode.timing` records that assignment as `Some(ExecutionTiming::IngestionTime)` or `Some(ExecutionTiming::QueryTime)`. Logical nodes may retain `None`; physical-plan validation must reject unassigned executable nodes. Timing is common node metadata rather than a separate payload From 6b84a099d62725ce9b7cec1339846ed5c4c46ccf Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:24:18 +0000 Subject: [PATCH 38/46] docs: introduce a compact overview of the proposed structure --- .../design_docs/proposals/operator-sharing.md | 77 ++++++++++++------- 1 file changed, 51 insertions(+), 26 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index e6b53ee3f..e22390bfe 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -49,19 +49,8 @@ The two computational categories are `Operator` and `ScalarExpr`. `Operator` describes a relation, vector or summary-state operation and has two variants. `OperatorNode` (§2) stores that operation together with its common properties. `QueryRoot` selects either an operator node or a scalar expression as the query -entry point; it is an enum, not an additional computation node: - -```rust -enum Operator { - NonASAP(NonASAPOp), - ASAP(ASAPOp), -} - -enum QueryRoot { - Operator(Rc), - ScalarExpr(ScalarExpr), -} -``` +entry point; it is an enum, not an additional computation node. The structural +overview below shows how these types fit together. `QueryRoot::Operator` is used for SQL query results and PromQL vector queries, such as `sum(up)`. `QueryRoot::ScalarExpr` is used for PromQL scalar queries, such @@ -141,9 +130,51 @@ The following changes are explicit: `TimeRange.kind` distinguishes instant selection from a range window, as specified in the companion. `SQLWindowFunc.frame` retains its current optional form. -**Proposed data structures.** These sketches describe operation-specific data using -current names. `Rc` represents a shared input edge; no new reference type -is introduced. Storage and traversal algorithms remain outside this design. +**Proposed data structures — overview.** The complete outer structure is below; +operation variants and schema internals are expanded afterward. These declarations +are shared by the detailed sections, not separate abbreviated types. + +```rust +// Query entry: an operator graph or an owned scalar expression. +enum QueryRoot { + Operator(Rc), + ScalarExpr(ScalarExpr), +} + +// A graph node combines its operation with common planning properties (§2). +struct OperatorNode { + operator: Operator, + result_kind: OperatorResultKind, + schema: Schema, + guarantee: Option, + timing: Option, +} + +// Operation payloads: each variant below defines its own inputs and parameters. +enum Operator { + NonASAP(NonASAPOp), + ASAP(ASAPOp), +} + +// NonASAPOp / ASAPOp: detailed below; their inputs are Rc. +// ScalarExpr: an owned value-expression tree, defined in the companion proposal. +// Schema / OperatorResultKind: defined in §2.1. +``` + +Read this structure from the query entry inward: + +- An operator root points to an `OperatorNode`. That node holds either a NonASAP + or ASAP operation and its result/schema, accuracy and execution-phase properties. +- Operations reference input nodes through `Rc`, forming the graph. + They also own scalar expressions where needed, such as a filter predicate or + projection value. A scalar expression is not another operator category. +- A scalar root owns its expression directly. An explicit conversion such as + PromQL `scalar(v)` may reference an operator node producing `v`; that dependency + is part of the same graph (§1.2). + +`Rc` retains the existing shared-reference notation. Storage and traversal +algorithms remain outside this design. The following sketches expand the two +operation payloads; §2 explains the common fields without declaring them again. `NonASAPOp` retains the query semantics needed before and after optimization: @@ -296,20 +327,14 @@ it does not introduce rules for sharing computations across queries. ## 2. Node properties and why they differ -Both operation categories use this resolved node structure. It follows the current +Both operation categories use the `OperatorNode` declared in the §1.1 overview. +That resolved node follows the current `SummaryNode` separation between an operation and its metadata, generalized to all operators. The field is named `operator` because it holds `Operator` (§1.1), -not a scalar expression. Common fields are defined here once, not in each variant. +not a scalar expression. The table below explains those common fields; individual +operation variants do not repeat them. ```rust -struct OperatorNode { - operator: Operator, - result_kind: OperatorResultKind, - schema: Schema, - guarantee: Option, - timing: Option, -} - // Existing enum; the node's Option represents an unassigned phase. enum ExecutionTiming { IngestionTime, From dba58f506df35381fbcf728c4b499d67259c068a Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:24:33 +0000 Subject: [PATCH 39/46] docs: use LogicalDAGRoot consistently with planning proposal --- docs/design_docs/proposals/decoupling_op_and_expr.md | 2 +- docs/design_docs/proposals/operator-sharing.md | 10 +++++----- 2 files changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index 0813fd43e..f585b6ff4 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -217,7 +217,7 @@ references the query's producer; `ScalarExpr` holds its scalar expression. This enum adds no computation node: ```rust -enum QueryRoot { +enum LogicalDAGRoot { Operator(Rc>), ScalarExpr(ScalarExpr), } diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index e22390bfe..40560b0f1 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -48,12 +48,12 @@ on query semantics, accuracy and execution timing. The two computational categories are `Operator` and `ScalarExpr`. `Operator` describes a relation, vector or summary-state operation and has two variants. `OperatorNode` (§2) stores that operation together with its common properties. -`QueryRoot` selects either an operator node or a scalar expression as the query +`LogicalDAGRoot` selects either an operator node or a scalar expression as the query entry point; it is an enum, not an additional computation node. The structural overview below shows how these types fit together. -`QueryRoot::Operator` is used for SQL query results and PromQL vector queries, -such as `sum(up)`. `QueryRoot::ScalarExpr` is used for PromQL scalar queries, such +`LogicalDAGRoot::Operator` is used for SQL query results and PromQL vector queries, +such as `sum(up)`. `LogicalDAGRoot::ScalarExpr` is used for PromQL scalar queries, such as `2`, `time()` or `scalar(sum(up))`. Neither variant adds a wrapper operator. The table lists all operator kinds in this proposal. Aggregate functions, join @@ -136,7 +136,7 @@ are shared by the detailed sections, not separate abbreviated types. ```rust // Query entry: an operator graph or an owned scalar expression. -enum QueryRoot { +enum LogicalDAGRoot { Operator(Rc), ScalarExpr(ScalarExpr), } @@ -301,7 +301,7 @@ arbitrary expression from being mistaken for a table-producing plan. The [companion proposal](decoupling_op_and_expr.md) defines this distinction. Its pre-ASAP `Rc>` references become `Rc` in the unified -model, including `QueryRoot` and the inputs to `PromqlScalarFromVector`, +model, including `LogicalDAGRoot` and the inputs to `PromqlScalarFromVector`, `ScalarSubquery`, `Exists` and `InSubquery`. Their cardinality, NULL and NaN rules remain unchanged. From 4af61cadb3776b418f254ec3f827aca2e3f13761 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:26:52 +0000 Subject: [PATCH 40/46] docs: remove the proposed query-root wrapper type --- .../proposals/decoupling_op_and_expr.md | 34 ++++++---------- .../design_docs/proposals/operator-sharing.md | 39 +++++++------------ 2 files changed, 25 insertions(+), 48 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index f585b6ff4..c11d82801 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -35,7 +35,7 @@ compute values within the schema selected by their owning operators. `NonASAPOp` describes a relation or vector computation: reading data, selecting rows or samples, combining inputs, or reducing them. `ScalarExpr` computes one value in a column and evaluation context. An operator supplies that context; -a scalar query root has no row columns. The distinction is about semantic role, +a standalone scalar query has no row columns. The distinction is about semantic role, not whether the source language calls something an “expression”. Keep existing names where the semantics match. The proposal changes the boundary @@ -210,21 +210,11 @@ These consume query results; they are not interchangeable bridge nodes. The SQL variants cover uncorrelated subqueries here. Correlation needs outer-scope bindings that this proposal does not define (§3.4). -### 2.3 Query roots and composition without a bridge +### 2.3 Composition without a bridge -A query root selects one of the two representations. The `Operator` variant -references the query's producer; `ScalarExpr` holds its scalar expression. This -enum adds no computation node: - -```rust -enum LogicalDAGRoot { - Operator(Rc>), - ScalarExpr(ScalarExpr), -} -``` - -SQL query results use the operator arm. PromQL numeric/string roots use the scalar -arm; vector roots use the operator arm. A scalar root cannot contain free columns. +Standalone scalar expressions have no input-column scope and cannot contain free +column references. This proposal defines their semantics without introducing a +separate query-entry data structure. `EvalTimestamp` reads the current PromQL evaluation time; `CurrentTimestamp` reads SQL's statement time. Moving `EvalTimestamp` into `ScalarExpr` must not make it constant across evaluation steps or nested subquery times. @@ -232,11 +222,11 @@ constant across evaluation steps or nested subquery times. For float samples, the following plans need no `PromqlScalarBridge`: ```text -2 → Scalar root: Literal(2.0) -time() → Scalar root: EvalTimestamp +2 → ScalarExpr: Literal(2.0) +time() → ScalarExpr: EvalTimestamp up * 2 → Project(sample * 2.0, child = instant selection of up) vector(time()) → PromqlVectorFromScalar(EvalTimestamp) -scalar(sum(up)) + 1 → Scalar root: Arithmetic( +scalar(sum(up)) + 1 → ScalarExpr: Arithmetic( PromqlScalarFromVector(Aggregate(...)), Literal(1.0)) ``` @@ -252,7 +242,7 @@ schema nor a full identity value, this lowering is not valid. `vector(s)` remains a real conversion to a one-element, label-free vector. When combined with the [operator-sharing proposal](operator-sharing.md), all plan references above target the common `OperatorNode`, including references inside scalar -expressions and query roots. Scalar expression trees do not acquire shared-node +expressions. Scalar expression trees do not acquire shared-node identity. ## 3. Semantic requirements @@ -284,7 +274,7 @@ older repository dependency. This documentation change upgrades neither dependen | SQL construct / example | Representation in §2 | Coverage and semantic condition | |---|---|---| | `FROM t`, `WHERE x > 1`, `SELECT x * 2` | `Scan`, `Filter`, `Project`; scalar `Compare`, `Arithmetic` | **Direct.** Resolve columns against the input; predicates must be Boolean and only TRUE passes. | -| `VALUES (1), (2)`; `SELECT 1` | `Values`; `Project` over one empty row | **Direct.** SQL still returns a relation, unlike a PromQL scalar root. | +| `VALUES (1), (2)`; `SELECT 1` | `Values`; `Project` over one empty row | **Direct.** SQL still returns a relation, unlike a PromQL scalar query. | | Literals, columns, `CASE`, `CAST`, `TRY_CAST`, `IN (...)`, `IS NULL`, Boolean logic | Corresponding `ScalarExpr` variants | **Direct** within supported types. Preserve coercion, NULL propagation and conditional evaluation. Type gaps are listed below. | | Unary minus, `BETWEEN`, `IS TRUE`, null-safe equality | `Negative`; compositions of `Compare`, `Case`, `IsNull`, Boolean expressions | **Direct/Lowered.** Evaluate reused nontrivial operands once, using a projected column where needed; do not duplicate volatile calls. | | Scalar functions and `NOW()` | Resolved `FunctionCall`; `CurrentTimestamp` | **Direct** for registered contracts. Preserve stable/volatile evaluation behavior; unresolved names are unsupported. | @@ -311,7 +301,7 @@ range-vector expression. The request permits scalar or instant-vector roots. | PromQL construct / example | Representation in §2 | Coverage and semantic condition | |---|---|---| -| Numeric/string literals; parentheses; `time()` | Scalar `Literal`, nested expression, `EvalTimestamp` | **Direct.** No bridge node; free columns are invalid at a scalar root. | +| Numeric/string literals; parentheses; `time()` | Scalar `Literal`, nested expression, `EvalTimestamp` | **Direct.** No bridge node; standalone scalar queries cannot reference free columns. | | Scalar arithmetic, unary minus, `1 < bool 2` | `Arithmetic`, `Negative`, `Case(Compare(...), 1.0, 0.0)` | **Direct/Lowered** with PromQL numeric semantics. Scalar comparison without `bool` is invalid. | | `up{job="api"}` | Time-series `Scan` with predicates, `TimeRange(Instant)` | **Direct.** Label matching treats absent labels as empty and regexes as anchored. Selection uses lookback and staleness, not SQL row filtering alone. | | `up[5m]` | `TimeRange(Range)` over a time-series scan | **Direct.** Retain samples and their timestamps in the left-open, right-closed window. | @@ -373,7 +363,7 @@ of complete SQL/PromQL support. Acceptance requires: - The SQL example in §1 retains its values, schema and predicate context. - Every Direct/Lowered mapping in §3 has a valid typed representation; each Gap is explicitly rejected until its payload or equivalent lowering is defined. -- The scalar roots and mixed scalar/vector examples in §2.3 need no +- The scalar queries and mixed scalar/vector examples in §2.3 need no `PromqlScalarBridge` or equivalent constant-wrapper node. - Operator dependencies inside scalar conversions/subqueries remain visible and shared; scalar trees remain owned. Invalid result-kind combinations are rejected. diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 40560b0f1..01014d730 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -48,13 +48,7 @@ on query semantics, accuracy and execution timing. The two computational categories are `Operator` and `ScalarExpr`. `Operator` describes a relation, vector or summary-state operation and has two variants. `OperatorNode` (§2) stores that operation together with its common properties. -`LogicalDAGRoot` selects either an operator node or a scalar expression as the query -entry point; it is an enum, not an additional computation node. The structural -overview below shows how these types fit together. - -`LogicalDAGRoot::Operator` is used for SQL query results and PromQL vector queries, -such as `sum(up)`. `LogicalDAGRoot::ScalarExpr` is used for PromQL scalar queries, such -as `2`, `time()` or `scalar(sum(up))`. Neither variant adds a wrapper operator. +The structural overview below shows how these types fit together. The table lists all operator kinds in this proposal. Aggregate functions, join kinds and scalar functions are choices within these operations, not additional @@ -66,8 +60,8 @@ operator kinds. | `Operator::ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | `CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to scalar -expressions. Constants need no `PromqlScalarBridge`; query roots may directly hold -a scalar expression. The [companion proposal](decoupling_op_and_expr.md) defines +expressions. Constants need no `PromqlScalarBridge`; scalar queries use +scalar expressions. The [companion proposal](decoupling_op_and_expr.md) defines these boundaries and their SQL/PromQL semantic coverage. A scalar-only frontend query, such as `time()`, therefore needs no operator node. @@ -106,7 +100,7 @@ the existing query and summary models, but the sketches combine several changes: **Naming and compatibility.** Reuse current names where their semantics match. The companion proposal defines deliberate additions and changes, including -`Values`, scalar query roots and typed scalar conversions. Ordinary payloads such +`Values` and typed scalar conversions. Ordinary payloads such as `Predicate`, `ProjectItem`, `AggIntent`, `SketchQuery` and `SummaryFamilyType` keep their names; retaining a name does not establish complete language coverage. @@ -135,12 +129,6 @@ operation variants and schema internals are expanded afterward. These declaratio are shared by the detailed sections, not separate abbreviated types. ```rust -// Query entry: an operator graph or an owned scalar expression. -enum LogicalDAGRoot { - Operator(Rc), - ScalarExpr(ScalarExpr), -} - // A graph node combines its operation with common planning properties (§2). struct OperatorNode { operator: Operator, @@ -161,14 +149,13 @@ enum Operator { // Schema / OperatorResultKind: defined in §2.1. ``` -Read this structure from the query entry inward: +The structure has three roles: -- An operator root points to an `OperatorNode`. That node holds either a NonASAP - or ASAP operation and its result/schema, accuracy and execution-phase properties. +- An `OperatorNode` holds either a NonASAP or ASAP operation and its result/schema, accuracy and execution-phase properties. - Operations reference input nodes through `Rc`, forming the graph. They also own scalar expressions where needed, such as a filter predicate or projection value. A scalar expression is not another operator category. -- A scalar root owns its expression directly. An explicit conversion such as +- A `ScalarExpr` describes value computation. An explicit conversion such as PromQL `scalar(v)` may reference an operator node producing `v`; that dependency is part of the same graph (§1.2). @@ -293,7 +280,7 @@ operator is first created. A filter is an operator because it transforms a table. Its predicate, such as `latency > 100`, is a scalar expression evaluated in that table's schema. -Scalar expressions belong to an operator field or a scalar query root. Predicates, +Scalar expressions belong to an operator field or a scalar query. Predicates, projection expressions and sort keys describe value computation in that context. Explicit scalar conversions and subqueries may reference operators; those are visible graph dependencies with defined cardinality rules. This prevents an @@ -301,7 +288,7 @@ arbitrary expression from being mistaken for a table-producing plan. The [companion proposal](decoupling_op_and_expr.md) defines this distinction. Its pre-ASAP `Rc>` references become `Rc` in the unified -model, including `LogicalDAGRoot` and the inputs to `PromqlScalarFromVector`, +model, including the inputs to `PromqlScalarFromVector`, `ScalarSubquery`, `Exists` and `InSubquery`. Their cardinality, NULL and NaN rules remain unchanged. @@ -458,7 +445,7 @@ Its `input` is the applicable column scope: the child schema for a projection, both input schemas for a join predicate, or aggregate outputs for `HAVING`. Explicit subquery/conversion expressions validate their referenced producer using the contracts above. Numeric expressions cannot consume state columns as numbers. -A scalar root is checked with an empty column scope and needs no fabricated +A standalone scalar expression is checked with an empty column scope and needs no fabricated relation output schema. `QueryExprError` retains the existing error-type name; result-kind and state-family mismatches require corresponding validation errors. @@ -519,7 +506,7 @@ how they use the common operator model: | Stage from #509 | Use of the unified representation | |---|---| -| Frontends | Produce an operator or scalar query root; any referenced operator nodes are `NonASAP`. Preserve source-language semantics. | +| Frontends | Represent queries using operators and scalar expressions; any referenced operator nodes are `NonASAP`. Preserve source-language semantics. | | Logical ASAP-aware optimization | Form candidate graphs containing ordinary and summary operators, with no wrappers hiding their dependencies. | | Physical ASAP-aware optimization | Determine executable alternatives, including materialization and execution timing, for those candidate graphs. | | Plan selection | Evaluate complete physical candidates using workload requirements and deployment-provided models and capabilities. | @@ -536,7 +523,7 @@ accuracy, costing or selection policies. Export one node per operator and represent its input dependencies as edges. Export a shared producer once, with edges to all its consumers. Include dependencies -referenced by scalar conversions and subqueries. Preserve scalar query roots as +referenced by scalar conversions and subqueries. Preserve scalar queries as expressions and their operator dependencies; do not invent a bridge node for export. This keeps the graph visible to costing, physical compilation, execution and plan @@ -562,7 +549,7 @@ The design is successful when: any shared inputs; it does not introduce new sharing rules. - Existing value/state, accuracy and execution constraints remain enforceable on the unified representation. -- Scalar query roots and conversions use the same representation before and after +- Scalar expressions and conversions use the same representation before and after optimization, with no bridge nodes or hidden subplans. - Validation includes result kind and query subgraphs referenced by scalar expressions when checking schema, accuracy and execution constraints. From 18d90db17c7ab616a3d8bd060d87a8017d8dba7b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:28:12 +0000 Subject: [PATCH 41/46] docs: illustrate logical DAG composition before and after ASAP --- .../proposals/decoupling_op_and_expr.md | 3 +- .../design_docs/proposals/operator-sharing.md | 101 +++++++++++++++--- 2 files changed, 90 insertions(+), 14 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index c11d82801..f027328fc 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -243,7 +243,8 @@ schema nor a full identity value, this lowering is not valid. When combined with the [operator-sharing proposal](operator-sharing.md), all plan references above target the common `OperatorNode`, including references inside scalar expressions. Scalar expression trees do not acquire shared-node -identity. +identity. The companion's [complete DAG example](operator-sharing.md#13-example-composing-a-logical-dag) +shows these expressions inside ordinary operators before and after an ASAP rewrite. ## 3. Semantic requirements diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 01014d730..decb1fe1d 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -263,18 +263,6 @@ The payloads above describe operations and their inputs. The common `OperatorNod fields and their derivation interfaces are defined once in §2; individual variants do not repeat schema, accuracy or execution timing. -For example, arrows below show data flowing from producer to consumer: - -```text -NonASAP(Scan) → NonASAP(Filter) → ASAP(SummaryAgg) - → ASAP(SummaryEstimate) → NonASAP(Project) -``` - -Each node describes its operation, inputs and output schema. An executable plan also -needs an accuracy assessment and an execution phase for each relevant computation. -These have different sources, described in §2; they are not all known when a logical -operator is first created. - ### 1.2 Operators and scalar expressions A filter is an operator because it transforms a table. Its predicate, such as @@ -305,7 +293,94 @@ These are **query subgraphs referenced by scalar expressions**. The reference is an edge to a producer in the common graph, not a copy of the subgraph embedded in the expression. Scalar ownership does not change the producer's graph identity. -### 1.3 Scope of operator sharing +### 1.3 Example: composing a logical DAG + +Consider this SQL query, with integer `bytes` and `status` columns: + +```sql +SELECT SUM(bytes) + 1 AS total_bytes +FROM requests +WHERE status = 200; +``` + +Before ASAP optimization, its logical DAG is composed as follows. Each box is an +`OperatorNode`; arrows point from a consumer to its input producer. The scalar +expressions shown beside nodes are owned fields, not additional DAG nodes. + +```text +p: OperatorNode + operator = Operator::NonASAP(NonASAPOp::Project) + cols[0].expr = ScalarExpr::Arithmetic(Column(sum_bytes), Add, Literal(1)) + │ child: Rc + ▼ +a: OperatorNode + operator = Operator::NonASAP(NonASAPOp::Aggregate) + measures = [AggIntent::Sum(bytes)] + │ child: Rc + ▼ +f: OperatorNode + operator = Operator::NonASAP(NonASAPOp::Filter) + pred = Predicate(ScalarExpr::Compare(Column(status), Eq, Literal(200))) + │ child: Rc + ▼ +s: OperatorNode + operator = Operator::NonASAP(NonASAPOp::Scan) + source = requests +``` + +This is abbreviated structural notation: `Column` and `Literal` above are +`ScalarExpr` variants; column names stand for resolved `ColumnId`s. The arithmetic +and comparison use `ExprSemantics::Sql`. The aggregate has no grouping keys and +names its output `sum_bytes`; the projection names its output `total_bytes`. + +An eligible ASAP rewrite can implement the sum using an exact accumulator. The +resulting logical DAG contains both operation categories: + +```text +p': NonASAP(Project) owns the same scalar expression: sum_bytes + 1 + │ child + ▼ +r: ASAP(FinalizeExactAccumulator) produces the ordinary sum_bytes value + │ child + ▼ +b: ASAP(SummaryAgg) produces exact SUM accumulator state + │ child + ▼ +f: NonASAP(Filter) owns the same predicate: status = 200 + │ child + ▼ +s: NonASAP(Scan) reads requests +``` + +The second diagram abbreviates the same nesting: `ASAP(SummaryAgg)` means an +`OperatorNode` whose `operator` is `Operator::ASAP(ASAPOp::SummaryAgg { ... })`. +Its family is `SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum)`; +its update reads `bytes`, and it uses the same ungrouped reduction. Finalization +must preserve SQL SUM's NULL and empty-input behavior. This example assumes the +existing capability and rewrite checks permit that exact implementation. + +| Part of the design | Role in this example | +|---|---| +| `OperatorNode` | Every graph node, holding its operation and common result/schema, guarantee and timing properties. | +| `Operator` | Selects the `NonASAP` or `ASAP` operation category in each node. | +| `NonASAPOp` | Scan, filter, aggregate and projection before optimization; scan, filter and projection still use these definitions afterward. | +| `ASAPOp` | Builds accumulator state and finalizes it after the rewrite. | +| `ScalarExpr` | Computes `status = 200` and `sum_bytes + 1` within the filter and projection; neither computation needs a bridge node. | +| `Rc` | Connects each consumer to its producer, including `Project.child` pointing to an ASAP finalization node. | + +The common properties also follow the new graph. Scan/filter/project and the +finalized sum have `Relation` results with ordinary `Plain(DataType)` columns. +The build has `State` result kind and an `ExactAggregate(...)` column; the +projection cannot consume that state directly. Guarantees are assessed under the +existing rules. At this logical stage, `timing` may remain `None`; physical planning +later assigns execution phases. The query result is produced by `p` before the +rewrite and `p'` afterward, without an additional query-root data structure. + +This illustrates the connection between the two proposals: scalar separation +makes predicates and value expressions explicit; operator unification lets those +same ordinary operations consume ASAP results through normal graph edges. + +### 1.4 Scope of operator sharing Here, sharing means pre-ASAP and post-ASAP use the same operator definitions. A `Project`, for example, has one representation whether its input is an ordinary From 94cbef002386e9b405786d7be42512d2cb9edfdb Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:29:44 +0000 Subject: [PATCH 42/46] docs: use descriptive node labels in DAG examples --- .../design_docs/proposals/operator-sharing.md | 29 +++++++++++-------- 1 file changed, 17 insertions(+), 12 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index decb1fe1d..e0809a696 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -303,27 +303,27 @@ FROM requests WHERE status = 200; ``` -Before ASAP optimization, its logical DAG is composed as follows. Each box is an +Before ASAP optimization, its logical DAG is composed as follows. Each named node is an `OperatorNode`; arrows point from a consumer to its input producer. The scalar expressions shown beside nodes are owned fields, not additional DAG nodes. ```text -p: OperatorNode +Project node: OperatorNode operator = Operator::NonASAP(NonASAPOp::Project) cols[0].expr = ScalarExpr::Arithmetic(Column(sum_bytes), Add, Literal(1)) │ child: Rc ▼ -a: OperatorNode +Aggregate node: OperatorNode operator = Operator::NonASAP(NonASAPOp::Aggregate) measures = [AggIntent::Sum(bytes)] │ child: Rc ▼ -f: OperatorNode +Filter node: OperatorNode operator = Operator::NonASAP(NonASAPOp::Filter) pred = Predicate(ScalarExpr::Compare(Column(status), Eq, Literal(200))) │ child: Rc ▼ -s: OperatorNode +Scan node: OperatorNode operator = Operator::NonASAP(NonASAPOp::Scan) source = requests ``` @@ -337,19 +337,24 @@ An eligible ASAP rewrite can implement the sum using an exact accumulator. The resulting logical DAG contains both operation categories: ```text -p': NonASAP(Project) owns the same scalar expression: sum_bytes + 1 +Project node: NonASAP(Project) + expression: sum_bytes + 1 │ child ▼ -r: ASAP(FinalizeExactAccumulator) produces the ordinary sum_bytes value +Finalize node: ASAP(FinalizeExactAccumulator) + output: ordinary sum_bytes value │ child ▼ -b: ASAP(SummaryAgg) produces exact SUM accumulator state +Summary build node: ASAP(SummaryAgg) + output: exact SUM accumulator state │ child ▼ -f: NonASAP(Filter) owns the same predicate: status = 200 +Filter node: NonASAP(Filter) + predicate: status = 200 │ child ▼ -s: NonASAP(Scan) reads requests +Scan node: NonASAP(Scan) + source: requests ``` The second diagram abbreviates the same nesting: `ASAP(SummaryAgg)` means an @@ -373,8 +378,8 @@ finalized sum have `Relation` results with ordinary `Plain(DataType)` columns. The build has `State` result kind and an `ExactAggregate(...)` column; the projection cannot consume that state directly. Guarantees are assessed under the existing rules. At this logical stage, `timing` may remain `None`; physical planning -later assigns execution phases. The query result is produced by `p` before the -rewrite and `p'` afterward, without an additional query-root data structure. +later assigns execution phases. The topmost Project node produces the query +result in both diagrams, without an additional query-root data structure. This illustrates the connection between the two proposals: scalar separation makes predicates and value expressions explicit; operator unification lets those From c81ba14a526a94ed066052a6eb4fbda3d8762f36 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:38:20 +0000 Subject: [PATCH 43/46] docs: consolidate operator and scalar proposal interfaces --- .../proposals/decoupling_op_and_expr.md | 129 +++------ .../design_docs/proposals/operator-sharing.md | 246 +++++++----------- 2 files changed, 137 insertions(+), 238 deletions(-) diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index f027328fc..58859db41 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -52,74 +52,27 @@ where needed, rather than copying every current `QueryExpr` variant unchanged: ### 2.1 Operator nodes -The sketches describe the proposed shape, not implemented Rust definitions. -`C` remains `ColumnRef` before resolution and `ColumnId` afterward; -`C::ScanSchema` retains its existing meaning. Operator references are shared, -including `Concat` inputs. Scalar fields own expression trees. +Use the canonical [`NonASAPOp` and `OperatorNode` definitions](operator-sharing.md#11-unified-operator-type) +from the sharing proposal. Both documents describe the same resolved model: +operator inputs and scalar query-result references use `Rc`. +`NonASAPOp` is the payload of an ordinary operator, not a second graph-node type. +`BinaryOp` likewise uses the single `BinaryOperator` payload specified there. -```rust -enum NonASAPOp { - Scan { - source: Source, predicates: Vec>, schema: C::ScanSchema, - }, - Values { rows: Vec>>, schema: C::ScanSchema }, - Filter { pred: Predicate, child: Rc> }, - Project { - cols: Vec>, qualifier: Option, child: Rc>, - }, - Aggregate { - reduction: Reduction, measures: Vec>, output_names: Vec, - having: Option>, child: Rc>, - }, - Dedup { cols: Vec, child: Rc> }, - Concat { - children: Vec>>, - discriminator_unique_key: Option>, - }, - Join { - kind: JoinKind, pred: Predicate, - left: Rc>, right: Rc>, - }, - SetOp { - kind: RelationalSetOpKind, all: bool, - left: Rc>, right: Rc>, - }, - Sort { keys: Vec>, partition_by: GroupKeys, child: Rc> }, - Limit { - n: Option, offset: usize, partition_by: GroupKeys, - child: Rc>, - }, - BinaryOp { - op: BinaryOpKind, lhs: Rc>, rhs: Rc>, - vector_match: Option, return_bool: bool, - }, - SQLWindowFunc { - func: WindowFuncKind, args: Vec>, partition_by: GroupKeys, - order_by: Vec>, frame: Option, - output_name: String, child: Rc>, - }, - TimeRange { range: Duration, kind: TimeRangeKind, child: Rc> }, - TimeShift { shift: TimeShift, child: Rc> }, - PromqlVectorFromScalar(ScalarExpr), - PromqlRelabel { dst: String, value: ScalarExpr, child: Rc> }, - PromqlInfoEnrich { selector: Vec, child: Rc> }, - PromqlSeriesSample { by: GroupKeys, kind: SampleKind, child: Rc> }, - PromqlSubquery { - range: Duration, resolution: Option, child: Rc>, - }, -} - -enum TimeRangeKind { Instant, Range } +Names are resolved to `ColumnId` before constructing these nodes. Parsing and +unresolved `ColumnRef` handling remain frontend concerns; no alternative generic +operator definition is proposed here. These wrappers belong to operator fields +and use the `ScalarExpr` defined in §2.2: -struct Predicate(ScalarExpr); +```rust +struct Predicate(ScalarExpr); -struct ProjectItem { +struct ProjectItem { alias: Option, - expr: ScalarExpr, + expr: ScalarExpr, } -struct SortKey { - expr: ScalarExpr, +struct SortKey { + expr: ScalarExpr, ascending: bool, nulls_first: bool, } @@ -147,42 +100,42 @@ time of that whole input. Existing signed offsets and `AtModifier` anchors remai Scalar recursion uses owned `Box` and `Vec` children. The only plan references are explicit operations that consume a query result to compute a value. Those edges remain visible to plan traversal and costing; they cannot hide a separate plan. -Since these variants reference `NonASAPOp`, `ScalarExpr` and its wrappers retain -`C: ColState`; the earlier proposed removal of that bound no longer applies. +The definitions below use resolved `ColumnId`s and the common `OperatorNode`; +there is no separate pre-ASAP scalar representation. ```rust -enum ScalarExpr { - Column(C), +enum ScalarExpr { + Column(ColumnId), Literal(ScalarValue), - Negative { expr: Box>, semantics: ExprSemantics }, + Negative { expr: Box, semantics: ExprSemantics }, Compare { - left: Box>, op: CompareOpKind, right: Box>, + left: Box, op: CompareOpKind, right: Box, semantics: ExprSemantics, }, - BoolAnd(Vec>), - BoolOr(Vec>), - Not(Box>), - IsNull(Box>), - IsNotNull(Box>), - Cast { expr: Box>, to: DataType, try_cast: bool }, - InList { expr: Box>, list: Vec>, negated: bool }, - FunctionCall { name: String, args: Vec> }, + BoolAnd(Vec), + BoolOr(Vec), + Not(Box), + IsNull(Box), + IsNotNull(Box), + Cast { expr: Box, to: DataType, try_cast: bool }, + InList { expr: Box, list: Vec, negated: bool }, + FunctionCall { name: String, args: Vec }, Arithmetic { - op: ArithmeticOpKind, left: Box>, right: Box>, + op: ArithmeticOpKind, left: Box, right: Box, semantics: ExprSemantics, }, Case { - operand: Option>>, - branches: Vec<(ScalarExpr, ScalarExpr)>, - else_expr: Option>>, + operand: Option>, + branches: Vec<(ScalarExpr, ScalarExpr)>, + else_expr: Option>, }, CurrentTimestamp, EvalTimestamp, - PromqlScalarFromVector(Rc>), - ScalarSubquery(Rc>), - Exists { subquery: Rc>, negated: bool }, + PromqlScalarFromVector(Rc), + ScalarSubquery(Rc), + Exists { subquery: Rc, negated: bool }, InSubquery { - expr: Box>, subquery: Rc>, negated: bool, + expr: Box, subquery: Rc, negated: bool, }, } @@ -240,9 +193,7 @@ For an open label schema, lowering must retain the complete series identity, including unreferenced labels. If the input provides neither a complete label schema nor a full identity value, this lowering is not valid. `vector(s)` remains a real conversion to a one-element, label-free vector. -When combined with the [operator-sharing proposal](operator-sharing.md), all plan -references above target the common `OperatorNode`, including references inside scalar -expressions. Scalar expression trees do not acquire shared-node +Scalar expression trees are owned, while their operator references preserve graph identity. The companion's [complete DAG example](operator-sharing.md#13-example-composing-a-logical-dag) shows these expressions inside ordinary operators before and after an ASAP rewrite. @@ -309,8 +260,8 @@ range-vector expression. The request permits scalar or instant-vector roots. | `-up` | `Project` with `Negative(sample)` | **Lowered** for float samples; retain series identity and unary-operation naming rules. | | `up * 2`, `2 / up`, `up * scalar(sum(other))` | `Project` with scalar arithmetic and, where needed, `PromqlScalarFromVector` | **Lowered** for float samples. Preserve operand order, time/label fields and metric-name removal. Scalar input is evaluated in the same query-time context. | | `up > 0`; `up > bool 0` | `Filter`; or `Project` with `Case(Compare(...), 1.0, 0.0)` | **Lowered** for float samples. Filtering preserves surviving sample values; bool mode produces numbers and removes the metric name. Scalar-on-left comparisons still retain the vector's sample value when filtering. | -| `a / on(job) group_left b`; `a > bool b` | `BinaryOp` with `vector_match`, `return_bool` | **Direct.** Retain matching cardinality, labels and metric-name rules; unmatched elements disappear, not become false rows. | -| `a and b`, `a or b`, `a unless b` | `BinaryOpKind::Set` | **Direct.** Label-set matching, not SQL Boolean evaluation or SQL bag set operations. | +| `a / on(job) group_left b`; `a > bool b` | `BinaryOp` with `operator.vector_match`, `return_bool` | **Direct.** Retain matching cardinality, labels and metric-name rules; unmatched elements disappear, not become false rows. | +| `a and b`, `a or b`, `a unless b` | `BinaryOp.operator.kind = BinaryOpKind::Set` | **Direct.** Label-set matching, not SQL Boolean evaluation or SQL bag set operations. | | `sum by(job)(up)`, `avg without(instance)(up)` | `Aggregate(Reduction::Reduce(...), AggIntent)` | **Direct** for represented intents; preserve PromQL label grouping and empty-input behavior. | | `topk(3, up)`, `bottomk(3, up)` with optional grouping | `Sort` + partitioned `Limit` | **Lowered** for constant parameters and float samples, retaining selected series labels and specified NaN ordering. Existing heavy-hitter `AggIntent::TopK` is not a substitute for sample-value ranking. | | `rate(x[5m])`, `sum_over_time(x[5m])` | Range input + `Aggregate(Reduction::PerEntity, corresponding AggIntent)` | **Direct** for existing contracts: counter resets, extrapolation and range reduction belong to the named intent, not ordinary SQL SUM. | diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index e0809a696..f63a8ce1e 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -45,84 +45,11 @@ on query semantics, accuracy and execution timing. ### 1.1 Unified `Operator` type -The two computational categories are `Operator` and `ScalarExpr`. `Operator` -describes a relation, vector or summary-state operation and has two variants. -`OperatorNode` (§2) stores that operation together with its common properties. -The structural overview below shows how these types fit together. - -The table lists all operator kinds in this proposal. Aggregate functions, join -kinds and scalar functions are choices within these operations, not additional -operator kinds. - -| Category | Meaning | All operations | -|---|---|---| -| `Operator::NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Values`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlVectorFromScalar`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | -| `Operator::ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | - -`CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to scalar -expressions. Constants need no `PromqlScalarBridge`; scalar queries use -scalar expressions. The [companion proposal](decoupling_op_and_expr.md) defines -these boundaries and their SQL/PromQL semantic coverage. A scalar-only frontend -query, such as `time()`, therefore needs no operator node. - -`SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` -are reserved in the proposed planner model; listing them does not establish planner -or runtime support. Defining new summary-composition semantics is outside this -proposal (§6). - -`NonASAPOp` and `ASAPOp` describe the operation performed by a node. Inputs in both -categories connect to `OperatorNode` instances, so either category can consume the other -when their result kind, schema and execution constraints permit it. `NonASAP` -classifies one node; it does not require all of that node's descendants to be non-ASAP. - -Both categories use the same graph model. A projection can consume a summary -estimate, and a summary can consume the result of a filter or join. There is no -separate relational subtree hidden inside a summary-plan node. - -The categories remain distinct because they have different semantic rules: -ordinary operators consume query values, while summary operations may produce or -consume state. A common graph lets planning reason about all dependencies; the -categories make state-specific accuracy and execution constraints explicit. A -frontend plan contains only `NonASAP` nodes. Optimization may introduce `ASAP` -nodes later. - -**Relationship to the current code.** These structures are proposals, not copies -of the current definitions with two category labels added. The operations come from -the existing query and summary models, but the sketches combine several changes: - -| Change | Purpose and scope | -|---|---| -| Group ordinary operations under `NonASAP` and summary operations under `ASAP` | Organize the common operator model into two semantic categories. | -| Give both categories inputs that refer to `OperatorNode` instances | Allow ordinary and summary operations to compose directly, without a wrapper hiding their dependencies. | -| Remove relational-subplan wrappers and duplicate relational operations | Make all dependencies visible and give each ordinary operation one definition. | -| Separate scalar expressions from operators | A related proposal, described in the [companion document](decoupling_op_and_expr.md); not a consequence of categorization alone. | -| Represent execution timing chosen during physical planning | Follow the planning-stage design in [#509](https://github.com/ProjectASAP/ASAPPlanner/pull/509), rather than introduce a new timing policy here (§2.3). | - -**Naming and compatibility.** Reuse current names where their semantics match. -The companion proposal defines deliberate additions and changes, including -`Values` and typed scalar conversions. Ordinary payloads such -as `Predicate`, `ProjectItem`, `AggIntent`, `SketchQuery` and `SummaryFamilyType` -keep their names; retaining a name does not establish complete language coverage. - -The following changes are explicit: - -- Operator inputs become references to the common `OperatorNode`. The sketches use the - existing `Rc` notation for shared inputs and omit column-state generics for - readability; column IDs and scan schemas below show the resolved form. -- `Predicate`, `ProjectItem` and other scalar-bearing payloads keep their roles and - own `ScalarExpr` values. The companion proposal defines owned scalar trees and - shared operator inputs, including `Concat`. Explicit scalar conversions and SQL - subqueries also reference the common `OperatorNode`; their dependencies remain visible. - This document uses those same structures. -- `BinaryOp` reuses the existing post-ASAP `BinaryOperator` payload. It carries the - binary operation, vector matching and checked-division requirements; the pre-ASAP - `op` and `vector_match` semantics must be preserved when mapped into it. - The proposed `return_bool` field also retains PromQL comparison mode. -- `Limit.partition_by` comes from the existing post-ASAP limit. Applying that field - to the unified operator remains an explicit design choice, not new functionality - implied by a rename. Optional `Limit.n` adds offset-only queries, and - `TimeRange.kind` distinguishes instant selection from a range window, as specified - in the companion. `SQLWindowFunc.frame` retains its current optional form. +`Operator` describes an ordinary or ASAP operation; `OperatorNode` combines it +with common planning properties. `ScalarExpr` describes value computation. The +following overview and payload definitions are the canonical resolved interfaces +used by both proposals. The companion document defines `ScalarExpr` and its +operator-field wrappers; it does not define a second operator model. **Proposed data structures — overview.** The complete outer structure is below; operation variants and schema internals are expanded afterward. These declarations @@ -149,19 +76,22 @@ enum Operator { // Schema / OperatorResultKind: defined in §2.1. ``` -The structure has three roles: +An operator owns its scalar expressions and references input nodes through +`Rc`. Either operation category can consume the other's outputs when +the input contract permits it. `NonASAP` describes one operation, not its entire +subgraph. Frontend graphs contain only NonASAP operations; ASAP optimization may +introduce state construction and readout. -- An `OperatorNode` holds either a NonASAP or ASAP operation and its result/schema, accuracy and execution-phase properties. -- Operations reference input nodes through `Rc`, forming the graph. - They also own scalar expressions where needed, such as a filter predicate or - projection value. A scalar expression is not another operator category. -- A `ScalarExpr` describes value computation. An explicit conversion such as - PromQL `scalar(v)` may reference an operator node producing `v`; that dependency - is part of the same graph (§1.2). +| Category | Meaning | All operations | +|---|---|---| +| `Operator::NonASAP(NonASAPOp)` | Ordinary query operations that transform, combine or aggregate data | `Scan`, `Values`, `Filter`, `Project`, `Aggregate`, `Join`, `SetOp`, `Concat`, `Dedup`, `Sort`, `Limit`, `BinaryOp`, `SQLWindowFunc`, `TimeRange`, `TimeShift`, `PromqlVectorFromScalar`, `PromqlRelabel`, `PromqlInfoEnrich`, `PromqlSeriesSample`, `PromqlSubquery` | +| `Operator::ASAP(ASAPOp)` | Operations on summary state and its results, including reserved operations | `SummaryAgg`, `SummaryEstimate`, `SummaryMerge`, `SummarySubtract`, `SummaryDelete`, `SummaryJoin`, `FinalizeExactAccumulator`, `MaintainPopulation`, `ReadPopulation`, `Extension` | -`Rc` retains the existing shared-reference notation. Storage and traversal -algorithms remain outside this design. The following sketches expand the two -operation payloads; §2 explains the common fields without declaring them again. +`CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to +`ScalarExpr`, defined in the [companion proposal](decoupling_op_and_expr.md#22-scalar-expressions). +A constant needs no bridge operator. The sketches use resolved `ColumnId`s and +`Schema`; name resolution precedes construction of these nodes. Compatibility +changes from current types are collected in §6. `NonASAPOp` retains the query semantics needed before and after optimization: @@ -203,13 +133,12 @@ enum NonASAPOp { PromqlSeriesSample { child: Rc, by: GroupKeys, kind: SampleKind }, PromqlSubquery { child: Rc, range: Duration, resolution: Option }, } + +enum TimeRangeKind { Instant, Range } ``` -The fields describe what an operator does to its inputs: +Ordinary payload fields have these roles: -- `child`, `left`, `right`, `lhs`, `rhs` and `children` are graph dependencies. They can lead to - either operator category, subject to the input's schema and result-kind requirements. - Dependencies referenced by scalar conversions/subqueries are edges in this same graph. - `Predicate` describes a row-level condition; `ProjectItem` contains a scalar expression and its optional output alias. - `reduction` describes whether aggregation combines groups or operates per entity; @@ -240,9 +169,9 @@ enum ASAPOp { // Reserved operations; semantics and support require further design. SummaryMerge { children: Vec> }, SummarySubtract { left: Rc, right: Rc }, - SummaryDelete { summary_input: Rc, key: ColumnRef }, + SummaryDelete { summary_input: Rc, key: ColumnId }, SummaryJoin { - outer: Rc, inner: Rc, key: ColumnRef, family: SummaryFamilyType, + outer: Rc, inner: Rc, key: ColumnId, family: SummaryFamilyType, }, Extension { child: Rc, name: String }, } @@ -259,9 +188,9 @@ The summary fields distinguish state construction and readout: | `query` / `readout` | The result requested from summary or maintained-population state | | `population` | The population whose membership and values are maintained | -The payloads above describe operations and their inputs. The common `OperatorNode` -fields and their derivation interfaces are defined once in §2; individual variants -do not repeat schema, accuracy or execution timing. +The common `OperatorNode` fields are declared in the overview above and explained +in §2. The operation variants do not repeat them. Reserved ASAP variants require +further semantic and capability design before use. ### 1.2 Operators and scalar expressions @@ -275,10 +204,9 @@ visible graph dependencies with defined cardinality rules. This prevents an arbitrary expression from being mistaken for a table-producing plan. The [companion proposal](decoupling_op_and_expr.md) defines this distinction. -Its pre-ASAP `Rc>` references become `Rc` in the unified -model, including the inputs to `PromqlScalarFromVector`, -`ScalarSubquery`, `Exists` and `InSubquery`. Their cardinality, NULL and NaN rules -remain unchanged. +The companion's `ScalarExpr` uses `Rc` for `PromqlScalarFromVector`, +`ScalarSubquery`, `Exists` and `InSubquery`, so those expressions already reference +this common graph before and after optimization. In `scalar(sum(up))`, `scalar()` is Prometheus PromQL's built-in vector-to-scalar function, explicitly written by the query author. This proposal does not insert @@ -289,9 +217,8 @@ a summary readout, preserving the required vector and accuracy semantics; it can substitute raw summary state. Ordinary expressions such as `price * 2` reference columns and literals, not a query subgraph. -These are **query subgraphs referenced by scalar expressions**. The reference is -an edge to a producer in the common graph, not a copy of the subgraph embedded in -the expression. Scalar ownership does not change the producer's graph identity. +These are **query subgraphs referenced by scalar expressions**, with the same +producer identity as any other operator dependency. ### 1.3 Example: composing a logical DAG @@ -373,13 +300,19 @@ existing capability and rewrite checks permit that exact implementation. | `ScalarExpr` | Computes `status = 200` and `sum_bytes + 1` within the filter and projection; neither computation needs a bridge node. | | `Rc` | Connects each consumer to its producer, including `Project.child` pointing to an ASAP finalization node. | -The common properties also follow the new graph. Scan/filter/project and the -finalized sum have `Relation` results with ordinary `Plain(DataType)` columns. -The build has `State` result kind and an `ExactAggregate(...)` column; the -projection cannot consume that state directly. Guarantees are assessed under the -existing rules. At this logical stage, `timing` may remain `None`; physical planning -later assigns execution phases. The topmost Project node produces the query -result in both diagrams, without an additional query-root data structure. +For this example, assume `bytes` is nullable `Int64`. The output metadata is: + +| Node | `result_kind` | Output columns (`name: dtype`, nullability) | +|---|---|---| +| Aggregate before optimization | `Relation` | `sum_bytes: Plain(Int64)`, nullable | +| Summary build after optimization | `State` | `sum_state: ExactAggregate(Sum, Sum)`, non-null accumulator state | +| Finalize after optimization | `Relation` | `sum_bytes: Plain(Int64)`, nullable | +| Project in either graph | `Relation` | `total_bytes: Plain(Int64)`, nullable | + +The empty accumulator finalizes to SQL NULL; the accumulator itself is state, not +a nullable numeric value. The projection consumes the finalized column. Guarantees +follow the existing assessment rules, while `timing` may remain `None` until +physical planning. The topmost Project node produces the query result. This illustrates the connection between the two proposals: scalar separation makes predicates and value expressions explicit; operator unification lets those @@ -478,7 +411,8 @@ impl Operator { } impl OperatorNode { - fn validate(&self) -> Result<(), QueryExprError>; + fn validate_structure(&self) -> Result<(), QueryExprError>; + fn validate_execution_timing(&self) -> Result<(), QueryExprError>; } impl ScalarExpr { @@ -510,15 +444,30 @@ schema may also contain ordinary grouping keys. `SummaryEstimate`, vector kind from their operation and input context. Matching numeric columns do not make those kinds interchangeable. -**Interface contracts.** `Operator::output_schema` derives the fields and -metadata for the actual inputs after a rewrite; `output_kind` derives the result -category. `validate_inputs` checks producer/consumer compatibility, including query -subgraphs referenced by scalar expressions. `OperatorNode::validate` additionally -checks that retained output metadata agrees with that derivation and that any -guarantee or timing assignment is valid under the existing rules. For example, -`PromqlScalarFromVector` requires an instant vector, and a summary readout requires the compatible state family. These checks -must succeed before a plan is accepted; sharing an enum does not make every -producer/consumer combination legal. +**Interface contracts.** `Operator::output_schema` and `output_kind` derive output +metadata from the payload and validated inputs. `validate_inputs` checks local +producer/consumer compatibility, such as vector inputs for `BinaryOp` or the +required state family for a summary readout. Scalar typing checks the input-kind +contract of `PromqlScalarFromVector` and other scalar plan reads. + +| Validation entry | Scope and stage | +|---|---| +| `OperatorNode::validate_structure()` | Walks the reachable operator graph, including scalar plan references; checks input contracts, scalar typing and agreement between retained and derived output metadata. Valid for logical and physical plans; permits `timing = None`. | +| `OperatorNode::validate_execution_timing()` | Includes structural validation, then requires assigned timing on every executable operator and checks phase dependencies. Used for executable physical candidates. | +| Existing planner assessment and selection (#509) | Establishes guarantees using the existing accuracy models and checks them against request requirements and deployment capabilities. Neither node method re-proves a guarantee or decides request feasibility. | + +The two node methods need only the graph and its annotations. Request requirements +and deployment models remain inputs to the existing planning/selection workflow, +not implicit globals of `validate_structure`. Passing the timing check alone does +not establish that a physical candidate satisfies the query's accuracy requirement. + +`Scan.schema` declares the source columns; `Values.schema` declares the constructed +row shape. `OperatorNode.schema` is the derived output for any operation. A scan's +predicates cannot change its declared output columns; a Values row must match the +declared arity, types and nullability. These leaf outputs retain the declaration's +column layout and time/identity information, with only justified metadata changes. +The declaration and derived output therefore have distinct roles, and structural +validation rejects disagreement rather than trusting two independent schemas. `scalar_type` keeps the existing method name and `(DataType, nullable)` result. Its `input` is the applicable column scope: the child schema for a projection, @@ -527,7 +476,8 @@ Explicit subquery/conversion expressions validate their referenced producer usin the contracts above. Numeric expressions cannot consume state columns as numbers. A standalone scalar expression is checked with an empty column scope and needs no fabricated relation output schema. `QueryExprError` retains the existing error-type name; -result-kind and state-family mismatches require corresponding validation errors. +result-kind, state-family, schema and execution-phase mismatches require +corresponding validation errors. For example, a KLL build outputs `State` with a `Sketch(SketchKind, GroupingStrategy)` column identifying KLL and its parameters. @@ -563,7 +513,7 @@ and when to compute it. This proposal follows that division. For example, a KLL summary build may execute at ingestion time or query time, depending on the materialization choice. `OperatorNode.timing` records that assignment as `Some(ExecutionTiming::IngestionTime)` or `Some(ExecutionTiming::QueryTime)`. -Logical nodes may retain `None`; physical-plan validation must reject unassigned +Logical nodes may retain `None`; `validate_execution_timing` rejects unassigned executable nodes. Timing is common node metadata rather than a separate payload field on selected `ASAPOp` variants. @@ -592,13 +542,6 @@ how they use the common operator model: | Plan selection | Evaluate complete physical candidates using workload requirements and deployment-provided models and capabilities. | | Deployment execution | Execute the selected graph, preserving its dependencies and assigned phases. | -Ordinary operators remain the same operations when their inputs are replaced by -summary computations. Replacing an aggregate below a projection, for example, does -not require a separate post-ASAP projection definition. - -This document changes the representation used by these stages, not their search, -accuracy, costing or selection policies. - ## 4. Export preserves the graph Export one node per operator and represent its input dependencies as edges. Export @@ -606,13 +549,10 @@ a shared producer once, with edges to all its consumers. Include dependencies referenced by scalar conversions and subqueries. Preserve scalar queries as expressions and their operator dependencies; do not invent a bridge node for export. -This keeps the graph visible to costing, physical compilation, execution and plan -inspection. Embedding a whole relational subtree in one exported node would hide -its internal sharing and recreate the original boundary problem. - Export preserves result kinds, resolved schemas, scalar value types and evaluation context, together with the applicable guarantees and assigned execution phases. -Physical compilation may lower one logical operation to several physical operations, but must preserve its dependencies and meaning. The execution layer +Physical compilation may lower one logical operation to several physical +operations, but must preserve its dependencies and meaning. The execution layer does not invent missing planning decisions. Changing the exported representation requires coordinated adoption by the planner @@ -631,23 +571,31 @@ The design is successful when: the unified representation. - Scalar expressions and conversions use the same representation before and after optimization, with no bridge nodes or hidden subplans. -- Validation includes result kind and query subgraphs referenced by scalar - expressions when checking schema, accuracy and execution constraints. +- Structural and timing validation include query subgraphs referenced by scalar + expressions; planner assessment includes their accuracy dependencies. - Export preserves visible dependencies and shared producers. ## 6. Scope and compatibility -This proposal defines the common operator structure, its operation-specific data -and how dependencies remain visible through export. Existing names and semantics -are retained except for the structural changes identified in §1.1. - -The pre-ASAP and post-ASAP versions of some operations carry different information. -The unified `BinaryOp` must retain existing checked-division requirements, and -`Limit` must retain the existing ability to limit within groups. These compatibility -requirements belong in this design because removing duplicate operator definitions -must not remove existing behavior. If a rewrite moves arithmetic into a scalar -expression, it must preserve applicable checked-division guards and exact fallback; -`ExprSemantics` selects language rules and does not replace those proof conditions. +The two documents define one resolved interface: this document owns +`OperatorNode`, `Operator`, operation payloads and schema/validation interfaces; +the companion owns `ScalarExpr`, its wrappers and language-semantic mappings. +Existing type names are retained where their meanings still apply. + +| Change from current code | Final representation and compatibility rule | +|---|---| +| Separate ordinary/summary models and relational wrappers | Both categories use `OperatorNode` dependencies; no wrapped relational subplan. | +| Mixed operator/scalar `QueryExpr` | Owned scalar expressions in operator fields; explicit scalar conversions/subqueries reference `OperatorNode`. | +| Different pre-/post-ASAP binary payloads | One `BinaryOp.operator: BinaryOperator`, retaining `kind`, `vector_match` and checked-division flags. `return_bool` on `BinaryOp` adds PromQL comparison mode. | +| Limit and time selection differences | Retain post-ASAP `Limit.partition_by`; optional `n` supports offset-only queries. `TimeRange.kind` distinguishes instant/range selection. | +| Reserved ASAP key references | `SummaryDelete.key` and `SummaryJoin.key` use resolved `ColumnId`s, like other column references in this graph. | +| Missing relation constructor | `Values` represents SQL literal rows and the one empty input row for SELECT without FROM. | +| Separate edge schemas | Common `Schema` uses `Column` while keeping the existing identity/time metadata (§2.1). | + +A rewrite that moves arithmetic into a scalar expression must preserve applicable +checked-division guards and exact fallback. `ExprSemantics` selects language rules; +it does not replace those proof conditions. Reserved ASAP operations still require +their own semantic/capability design. The [companion semantic tables](decoupling_op_and_expr.md#3-semantic-requirements) use DataFusion 55.1.0 and Prometheus 3.15.0 as design targets. Their gaps also apply From 79f0fae5b1cc46cba429bf28bfd8a325fc174905 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 22:20:01 +0000 Subject: [PATCH 44/46] renaming on some types related to schema --- .../design_docs/proposals/operator-sharing.md | 50 ++++++++++--------- 1 file changed, 26 insertions(+), 24 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index f63a8ce1e..bed6b2f40 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -150,13 +150,13 @@ Ordinary payload fields have these roles: that distinguishes branches together with their within-branch key. `ASAPOp` describes state construction, state operations and readout separately. -`SummaryFamilyType` retains its current name; state-producing operations use its +`FieldType` (§2.1) types every output field; state-producing operations use its summary or exact-accumulator cases, never its `Plain` case. ```rust enum ASAPOp { SummaryAgg { - child: Rc, family: SummaryFamilyType, input: SummaryUpdate, + child: Rc, family: FieldType, input: SummaryUpdate, reduction: Reduction, grouping: GroupingStrategy, }, SummaryEstimate { @@ -171,7 +171,7 @@ enum ASAPOp { SummarySubtract { left: Rc, right: Rc }, SummaryDelete { summary_input: Rc, key: ColumnId }, SummaryJoin { - outer: Rc, inner: Rc, key: ColumnId, family: SummaryFamilyType, + outer: Rc, inner: Rc, key: ColumnId, family: FieldType, }, Extension { child: Rc, name: String }, } @@ -286,7 +286,7 @@ Scan node: NonASAP(Scan) The second diagram abbreviates the same nesting: `ASAP(SummaryAgg)` means an `OperatorNode` whose `operator` is `Operator::ASAP(ASAPOp::SummaryAgg { ... })`. -Its family is `SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum)`; +Its family is `FieldType::ExactAggregate(ExactKind::Sum, ExactParams::Sum)`; its update reads `bytes`, and it uses the same ungrouped reduction. Finalization must preserve SQL SUM's NULL and empty-input behavior. This example assumes the existing capability and rewrite checks permit that exact implementation. @@ -366,28 +366,29 @@ caching or mutation mechanism. ### 2.1 One schema model for values and state -Use one `Schema` for operator outputs before and after optimization. Reuse the -existing `SummaryFamilyType` to distinguish ordinary values from state, and retain -the current `Schema` metadata. The following is the proposed resolved interface; -it is not the current Rust definition. +Use one `Schema` for operator outputs before and after optimization. Rename today's +`SummaryFamilyType` to `FieldType`: it types every field, and `Plain` is not a summary +family. Rename `Column` to `Field` and `Schema.columns` to `Schema.fields`: the struct +describes a column and holds none of its data. Retain the current `Schema` metadata. +The following is the proposed resolved interface; it is not the current Rust definition. ```rust -struct Column { +struct Field { name: String, - dtype: T, + dtype: FieldType, nullable: bool, table: Option, } struct Schema { - columns: Vec>, + fields: Vec, time_index: Option, unique_keys: Vec>, closed: bool, } -// Existing variants and payload names, reused without renaming. -enum SummaryFamilyType { +// Today's `SummaryFamilyType`, renamed; variants and payloads unchanged. +enum FieldType { Plain(DataType), ExactAggregate(ExactKind, ExactParams), Sketch(SketchKind, GroupingStrategy), @@ -420,22 +421,22 @@ impl ScalarExpr { } ``` -**Relationship to current types.** `Column` gains a type parameter so the common -operator schema can use `SummaryFamilyType`, while existing nested value types -such as `DataType::List` still use ordinary `Column`. The proposed common -`Schema` replaces the separate operator-edge roles of pre-ASAP `Schema` and -post-ASAP `SummarySchema` / `SummaryField`; it does not rename `DataType` or add a -second summary-family enum. A pre-ASAP value column becomes `Plain(dtype)`. +**Relationship to current types.** `Field` is today's pre-ASAP `Column` with `dtype` +widened from `DataType` to `FieldType`. `FieldType` is today's `SummaryFamilyType` +under a name that also fits its `Plain` case. The proposed common `Schema` replaces +the separate operator-edge roles of pre-ASAP `Schema` and post-ASAP `SummarySchema` / +`SummaryField`; it does not rename `DataType`. A pre-ASAP value column becomes +`Plain(dtype)`. Frontend validation permits only ordinary value columns, preserving the current pre-ASAP restriction even though the common schema can also express state. | Field | Meaning and requirement | |---|---| -| `columns` | Ordered named fields. `Plain(DataType)` is a readable value; other variants retain the identity and parameters of summary or exact-accumulator state. | -| `Column.nullable`, `Column.table` | Preserve SQL nullability and qualified column resolution. | +| `fields` | Ordered named fields. `Plain(DataType)` is a readable value; other variants retain the identity and parameters of summary or exact-accumulator state. | +| `Field.nullable`, `Field.table` | Preserve SQL nullability and qualified column resolution. | | `time_index` | Identifies the time column when present; it does not by itself distinguish an instant vector from a range vector. | | `unique_keys` | Proven column combinations identifying rows; an empty list asserts no known key. Recompute these proofs when a rewrite changes identity. | -| `closed` | Whether `columns` completely describes the output. An open PromQL schema must retain unlisted labels through the existing complete-series-identity contract. | +| `closed` | Whether `fields` completely describes the output. An open PromQL schema must retain unlisted labels through the existing complete-series-identity contract. | `OperatorResultKind` is derived from the operation and its inputs and retained as `OperatorNode.result_kind`. `State` describes an output carrying unfinalized state; its @@ -485,7 +486,7 @@ Its p99 readout outputs an ordinary `Plain(Float64)` column in the appropriate relation/vector schema. A numeric predicate can use that readout, but not the KLL state. Exact accumulator state similarly requires `FinalizeExactAccumulator`. An ordinary operator may pass state through only where its input/output contract -permits it. A bare-column projection can preserve the column's `SummaryFamilyType` +permits it. A bare-column projection can preserve the field's `FieldType` directly during `output_schema` derivation; `scalar_type` applies when that column is used as a scalar value and rejects state. Copying a state column does not turn it into a readable scalar. @@ -590,7 +591,8 @@ Existing type names are retained where their meanings still apply. | Limit and time selection differences | Retain post-ASAP `Limit.partition_by`; optional `n` supports offset-only queries. `TimeRange.kind` distinguishes instant/range selection. | | Reserved ASAP key references | `SummaryDelete.key` and `SummaryJoin.key` use resolved `ColumnId`s, like other column references in this graph. | | Missing relation constructor | `Values` represents SQL literal rows and the one empty input row for SELECT without FROM. | -| Separate edge schemas | Common `Schema` uses `Column` while keeping the existing identity/time metadata (§2.1). | +| Separate edge schemas | Common `Schema` uses `Field` / `FieldType` while keeping the existing identity/time metadata (§2.1). | +| `SummaryFamilyType`, `Column`, `Schema.columns` | Renamed `FieldType`, `Field`, `Schema.fields`; variants and payloads unchanged. | A rewrite that moves arithmetic into a scalar expression must preserve applicable checked-division guards and exact fallback. `ExprSemantics` selects language rules; From c9b0906c0bac9e49490912fe6d61b8d82314a60b Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 22:26:29 +0000 Subject: [PATCH 45/46] renaming again --- .../design_docs/proposals/operator-sharing.md | 22 +++++++++---------- 1 file changed, 11 insertions(+), 11 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index bed6b2f40..53e4d82f6 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -150,13 +150,13 @@ Ordinary payload fields have these roles: that distinguishes branches together with their within-branch key. `ASAPOp` describes state construction, state operations and readout separately. -`FieldType` (§2.1) types every output field; state-producing operations use its +`FieldDataType` (§2.1) types every output field; state-producing operations use its summary or exact-accumulator cases, never its `Plain` case. ```rust enum ASAPOp { SummaryAgg { - child: Rc, family: FieldType, input: SummaryUpdate, + child: Rc, family: FieldDataType, input: SummaryUpdate, reduction: Reduction, grouping: GroupingStrategy, }, SummaryEstimate { @@ -171,7 +171,7 @@ enum ASAPOp { SummarySubtract { left: Rc, right: Rc }, SummaryDelete { summary_input: Rc, key: ColumnId }, SummaryJoin { - outer: Rc, inner: Rc, key: ColumnId, family: FieldType, + outer: Rc, inner: Rc, key: ColumnId, family: FieldDataType, }, Extension { child: Rc, name: String }, } @@ -286,7 +286,7 @@ Scan node: NonASAP(Scan) The second diagram abbreviates the same nesting: `ASAP(SummaryAgg)` means an `OperatorNode` whose `operator` is `Operator::ASAP(ASAPOp::SummaryAgg { ... })`. -Its family is `FieldType::ExactAggregate(ExactKind::Sum, ExactParams::Sum)`; +Its family is `FieldDataType::ExactAggregate(ExactKind::Sum, ExactParams::Sum)`; its update reads `bytes`, and it uses the same ungrouped reduction. Finalization must preserve SQL SUM's NULL and empty-input behavior. This example assumes the existing capability and rewrite checks permit that exact implementation. @@ -367,7 +367,7 @@ caching or mutation mechanism. ### 2.1 One schema model for values and state Use one `Schema` for operator outputs before and after optimization. Rename today's -`SummaryFamilyType` to `FieldType`: it types every field, and `Plain` is not a summary +`SummaryFamilyType` to `FieldDataType`: it types every field, and `Plain` is not a summary family. Rename `Column` to `Field` and `Schema.columns` to `Schema.fields`: the struct describes a column and holds none of its data. Retain the current `Schema` metadata. The following is the proposed resolved interface; it is not the current Rust definition. @@ -375,7 +375,7 @@ The following is the proposed resolved interface; it is not the current Rust def ```rust struct Field { name: String, - dtype: FieldType, + dtype: FieldDataType, nullable: bool, table: Option, } @@ -388,7 +388,7 @@ struct Schema { } // Today's `SummaryFamilyType`, renamed; variants and payloads unchanged. -enum FieldType { +enum FieldDataType { Plain(DataType), ExactAggregate(ExactKind, ExactParams), Sketch(SketchKind, GroupingStrategy), @@ -422,7 +422,7 @@ impl ScalarExpr { ``` **Relationship to current types.** `Field` is today's pre-ASAP `Column` with `dtype` -widened from `DataType` to `FieldType`. `FieldType` is today's `SummaryFamilyType` +widened from `DataType` to `FieldDataType`. `FieldDataType` is today's `SummaryFamilyType` under a name that also fits its `Plain` case. The proposed common `Schema` replaces the separate operator-edge roles of pre-ASAP `Schema` and post-ASAP `SummarySchema` / `SummaryField`; it does not rename `DataType`. A pre-ASAP value column becomes @@ -486,7 +486,7 @@ Its p99 readout outputs an ordinary `Plain(Float64)` column in the appropriate relation/vector schema. A numeric predicate can use that readout, but not the KLL state. Exact accumulator state similarly requires `FinalizeExactAccumulator`. An ordinary operator may pass state through only where its input/output contract -permits it. A bare-column projection can preserve the field's `FieldType` +permits it. A bare-column projection can preserve the field's `FieldDataType` directly during `output_schema` derivation; `scalar_type` applies when that column is used as a scalar value and rejects state. Copying a state column does not turn it into a readable scalar. @@ -591,8 +591,8 @@ Existing type names are retained where their meanings still apply. | Limit and time selection differences | Retain post-ASAP `Limit.partition_by`; optional `n` supports offset-only queries. `TimeRange.kind` distinguishes instant/range selection. | | Reserved ASAP key references | `SummaryDelete.key` and `SummaryJoin.key` use resolved `ColumnId`s, like other column references in this graph. | | Missing relation constructor | `Values` represents SQL literal rows and the one empty input row for SELECT without FROM. | -| Separate edge schemas | Common `Schema` uses `Field` / `FieldType` while keeping the existing identity/time metadata (§2.1). | -| `SummaryFamilyType`, `Column`, `Schema.columns` | Renamed `FieldType`, `Field`, `Schema.fields`; variants and payloads unchanged. | +| Separate edge schemas | Common `Schema` uses `Field` / `FieldDataType` while keeping the existing identity/time metadata (§2.1). | +| `SummaryFamilyType`, `Column`, `Schema.columns` | Renamed `FieldDataType`, `Field`, `Schema.fields`; variants and payloads unchanged. | A rewrite that moves arithmetic into a scalar expression must preserve applicable checked-division guards and exact fallback. `ExprSemantics` selects language rules; From 62e90e41b260fada4968ebff392cfc0f001908e4 Mon Sep 17 00:00:00 2001 From: Selvomega Date: Thu, 1 Oct 2026 22:39:11 +0000 Subject: [PATCH 46/46] pruned some doc --- .../design_docs/proposals/operator-sharing.md | 72 +------------------ 1 file changed, 2 insertions(+), 70 deletions(-) diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 53e4d82f6..30bbe3f79 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -90,8 +90,7 @@ introduce state construction and readout. `CurrentTimestamp`, `EvalTimestamp` and `PromqlScalarFromVector` belong to `ScalarExpr`, defined in the [companion proposal](decoupling_op_and_expr.md#22-scalar-expressions). A constant needs no bridge operator. The sketches use resolved `ColumnId`s and -`Schema`; name resolution precedes construction of these nodes. Compatibility -changes from current types are collected in §6. +`Schema`; name resolution precedes construction of these nodes. `NonASAPOp` retains the query semantics needed before and after optimization: @@ -529,38 +528,7 @@ execution phases. `TimeShift`, subquery grids and `EvalTimestamp` retain their source-language evaluation context. A shared node identity alone does not permit reusing a result across different evaluation times. -## 3. Planning responsibilities - -These are the stages defined in -[#509](https://github.com/ProjectASAP/ASAPPlanner/pull/509), shown here only to explain -how they use the common operator model: - -| Stage from #509 | Use of the unified representation | -|---|---| -| Frontends | Represent queries using operators and scalar expressions; any referenced operator nodes are `NonASAP`. Preserve source-language semantics. | -| Logical ASAP-aware optimization | Form candidate graphs containing ordinary and summary operators, with no wrappers hiding their dependencies. | -| Physical ASAP-aware optimization | Determine executable alternatives, including materialization and execution timing, for those candidate graphs. | -| Plan selection | Evaluate complete physical candidates using workload requirements and deployment-provided models and capabilities. | -| Deployment execution | Execute the selected graph, preserving its dependencies and assigned phases. | - -## 4. Export preserves the graph - -Export one node per operator and represent its input dependencies as edges. Export -a shared producer once, with edges to all its consumers. Include dependencies -referenced by scalar conversions and subqueries. Preserve scalar queries as -expressions and their operator dependencies; do not invent a bridge node for export. - -Export preserves result kinds, resolved schemas, scalar value types and evaluation -context, together with the applicable guarantees and assigned execution phases. -Physical compilation may lower one logical operation to several physical -operations, but must preserve its dependencies and meaning. The execution layer -does not invent missing planning decisions. - -Changing the exported representation requires coordinated adoption by the planner -and downstream readers while preserving existing query semantics and the selected -plan's execution requirements. - -## 5. Acceptance criteria +## 3. Acceptance criteria The design is successful when: @@ -574,39 +542,3 @@ The design is successful when: optimization, with no bridge nodes or hidden subplans. - Structural and timing validation include query subgraphs referenced by scalar expressions; planner assessment includes their accuracy dependencies. -- Export preserves visible dependencies and shared producers. - -## 6. Scope and compatibility - -The two documents define one resolved interface: this document owns -`OperatorNode`, `Operator`, operation payloads and schema/validation interfaces; -the companion owns `ScalarExpr`, its wrappers and language-semantic mappings. -Existing type names are retained where their meanings still apply. - -| Change from current code | Final representation and compatibility rule | -|---|---| -| Separate ordinary/summary models and relational wrappers | Both categories use `OperatorNode` dependencies; no wrapped relational subplan. | -| Mixed operator/scalar `QueryExpr` | Owned scalar expressions in operator fields; explicit scalar conversions/subqueries reference `OperatorNode`. | -| Different pre-/post-ASAP binary payloads | One `BinaryOp.operator: BinaryOperator`, retaining `kind`, `vector_match` and checked-division flags. `return_bool` on `BinaryOp` adds PromQL comparison mode. | -| Limit and time selection differences | Retain post-ASAP `Limit.partition_by`; optional `n` supports offset-only queries. `TimeRange.kind` distinguishes instant/range selection. | -| Reserved ASAP key references | `SummaryDelete.key` and `SummaryJoin.key` use resolved `ColumnId`s, like other column references in this graph. | -| Missing relation constructor | `Values` represents SQL literal rows and the one empty input row for SELECT without FROM. | -| Separate edge schemas | Common `Schema` uses `Field` / `FieldDataType` while keeping the existing identity/time metadata (§2.1). | -| `SummaryFamilyType`, `Column`, `Schema.columns` | Renamed `FieldDataType`, `Field`, `Schema.fields`; variants and payloads unchanged. | - -A rewrite that moves arithmetic into a scalar expression must preserve applicable -checked-division guards and exact fallback. `ExprSemantics` selects language rules; -it does not replace those proof conditions. Reserved ASAP operations still require -their own semantic/capability design. - -The [companion semantic tables](decoupling_op_and_expr.md#3-semantic-requirements) -use DataFusion 55.1.0 and Prometheus 3.15.0 as design targets. Their gaps also apply -here: a common `Operator` type does not supply missing aggregate modifiers, value -types or function contracts. This proposal changes neither repository dependencies -nor the set of implemented language features. - -New accuracy metrics, accuracy-composition rules, computation-sharing algorithms and -lifecycle policies are outside this proposal. Planning responsibilities follow #509. -The scalar/operator separation is specified in the companion document. Storage, -traversal algorithms, serialization fields and a code migration sequence are also -outside this document.