From 6ac8334bffb3038d73e068f38f3f227f5ea18af3 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?T=C3=B5nis=20Ormisson?= Date: Wed, 9 Sep 2026 14:09:35 +0300 Subject: [PATCH 1/2] docs: document numeric target workflow --- docs/transformations.md | 132 ++++++++++++++++++++++++++++++++++++++++ 1 file changed, 132 insertions(+) diff --git a/docs/transformations.md b/docs/transformations.md index d51652e..a128d07 100644 --- a/docs/transformations.md +++ b/docs/transformations.md @@ -135,6 +135,138 @@ openstatspec apply-spss --database-url sqlite:///survey.sqlite \ On Dolt, also pass `--expected-branch` and `--expected-head`. +## Numeric targets: create versus replace + +`COMPUTE` chooses its canonical target mode from the live schema. An absent +numeric target lowers to `assign` with `target_mode=create`; it appends one +nullable numeric column and one catalog variable. An existing numeric target +lowers to `target_mode=replace`. Replace keeps the variable ID, ordinal, and +metadata that no later operation explicitly changes. `IF` requires an existing +target, so a new target must be initialized first. Variable-source `NULL` +propagates through assignment; a predicate changes a row only when its SQL +truth value is `TRUE`. + +For example, this is a create-or-replace program for a numeric target. On +SQLite/PostgreSQL it creates `target` when absent; with a pre-existing numeric +`target` it updates the same target instead: + +```text +COMPUTE target = source_a. +IF (source_a = 1 AND source_b = 1) target = 1. +IF (target = 1 AND source_b = 1) target = 2. +VARIABLE LABELS target 'Numeric target'. +VALUE LABELS target 0 'No' 1 'Yes'. +FORMATS target (F8.1). +VARIABLE LEVEL target (NOMINAL). +EXECUTE. +``` + +`RECODE ... INTO` always requests a new target and rejects an existing target +name. A canonical replace recode must use the same source and target name. +Without `ELSE`, create recodes use `system_missing` for unmatched rows, while +replace recodes use `copy`. To pre-provision `band` and preserve the create +semantics, initialize and recode it explicitly. The absent-target form is: + +```text +RECODE source_a (1 = 0.1) (2 = 1) INTO band. +``` + +The pre-existing-target form is: + +```text +COMPUTE band = source_a. +RECODE band (1 = 0.1) (2 = 1) (ELSE = SYSMIS). +``` + +The second form has an extra ordered assignment and is not a pure DDL-overhead +comparison. Omitting its `ELSE` would copy unmatched source values instead of +matching the create form. For canonical plans, create recodes use +`target_mode=create` and replace recodes use `source=target`, +`target_mode=replace`. + +On SQLite and PostgreSQL, numeric create is supported only where the adapter +has the atomic native transaction boundary. MySQL, MariaDB, and Dolt reject a +create target with `schema_change_not_atomic`. A caller must use a separate, +explicit, versioned provisioning action that creates both the nullable physical +column and its matching normative variable, including a unique variable ID, +the same physical binding, and the next ordinal. Provisioning is not an apply +step, and this package does not hide it behind a new helper or API. See the +specification's [In-Place Transformation Binding 0.2 target-creation +rules](https://github.com/OpenStatSpec/specification/blob/main/docs/transformation-plan-sql-binding-0.2.md#target-creation-by-sql-profile). +For Dolt, commit provisioning separately; the later apply starts from that +clean committed `HEAD` and supplies it as `expected_head`. The SQLite matrix +below is local evidence only and is not service or cross-engine evidence. + +## Local numeric-target measurement + +The following is a bounded observation of the existing workflow, not a +release-to-release benchmark. A disposable standard-library harness created a +fresh file-backed SQLite database for every variant and repetition, kept +fixture setup separate from the timed public call, and removed all temporary +files afterward. It covered two workloads (A: the `COMPUTE`/`IF` program above; +B: `RECODE`), create versus pre-provisioned replace, 100 versus 100,000 cases, +and both `apply_transformation_plan_in_place` and `apply_spss_in_place`. Each +cell had one counted fresh apply, one fresh warm-up, and five fresh timed +applies; create/replace order alternated by repetition. Both sizes repeated +this input cycle in order: +`(1,1)`, `(1,NULL)`, `(NULL,1)`, `(0,1)`, `(2,2)`, `(0.1,0)`, `(-9,1)`, +`(NULL,0)`. Each fixture had numeric `source_a` and `source_b`; `source_a` +carried label `Source A` and missing rule `99`, while `source_b` carried label +`Source B` and no missing rule. Replace fixtures added one nullable numeric +ordinal-3 target (`target` for A, `band` for B) with label `Old target`, `F4.0` +formats, ordinal measurement level, and missing rule `99`; create fixtures had +no target column. + +The counted value is an outermost SQLAlchemy `Connection.execute` or +`Connection.exec_driver_sql` invocation scoped to the owned database. It +includes profile probes, catalog checks, reflection, explicit `BEGIN`, DDL, +DML, and audit work; one executemany call counts once. It is not a driver wire +round-trip count. Timed values are the public apply call only: setup, +validation, and cleanup are excluded. Every counted and timed result checked +ordered rows, explicit `NULL`s, binary64 bits (including `0.1`), source and +target identity, catalog/audit provenance, and absence of forbidden copy or +snapshot artifacts before being retained. + +Environment and revisions: Python 3.13.2, SQLite 3.47.1, SQLAlchemy 2.0.52, +Linux 6.6.87.2-microsoft-standard-WSL2 on x86_64 with 12 logical CPUs and +local temporary-file storage; observed file settings were +`journal_mode=delete` and `synchronous=2` (`FULL`). The Python source was +`63a66d6589aeb8fdc4820d5e7217fc13c61ef3b9`; the reviewed specification +reference was `930345b922af4f43b0e622be02d4b12ebfeb08eb`. Fixture setup was +recorded separately and excluded from the table. + +| Workload | Cases | API | Mode | SQLAlchemy executes | Median ms | Min ms | Max ms | +| --- | ---: | --- | --- | ---: | ---: | ---: | ---: | +| A | 100 | canonical plan | create | 234 | 76.5 | 65.8 | 78.4 | +| A | 100 | SPSS | create | 237 | 69.3 | 66.8 | 78.2 | +| A | 100 | canonical plan | replace | 222 | 67.6 | 56.6 | 88.3 | +| A | 100 | SPSS | replace | 225 | 71.3 | 69.3 | 88.5 | +| A | 100,000 | canonical plan | create | 234 | 168.8 | 134.1 | 172.3 | +| A | 100,000 | SPSS | create | 237 | 160.2 | 149.4 | 170.3 | +| A | 100,000 | canonical plan | replace | 222 | 162.7 | 146.0 | 173.6 | +| A | 100,000 | SPSS | replace | 225 | 155.4 | 144.8 | 177.6 | +| B | 100 | canonical plan | create | 224 | 61.6 | 55.6 | 76.6 | +| B | 100 | SPSS | create | 227 | 77.1 | 65.5 | 312.8 | +| B | 100 | canonical plan | replace | 213 | 64.2 | 60.2 | 223.1 | +| B | 100 | SPSS | replace | 216 | 80.9 | 62.2 | 237.5 | +| B | 100,000 | canonical plan | create | 224 | 149.4 | 131.7 | 193.4 | +| B | 100,000 | SPSS | create | 227 | 144.1 | 128.3 | 158.4 | +| B | 100,000 | canonical plan | replace | 213 | 169.8 | 163.6 | 186.6 | +| B | 100,000 | SPSS | replace | 216 | 192.8 | 171.4 | 202.8 | + +The canonical plan operation/hash identities for the same runs were: + +| Workload | Mode | Operations | Contract | Canonical plan SHA-256 | +| --- | --- | ---: | --- | --- | +| A | create | 8 | v0.2 | `6b17a220f9be9d2d29b671dd5eb4598ea34ae459e4d3b8e84a7daafbe3ec11ef` | +| A | replace | 8 | v0.2 | `a239cac39d3fade1207228bee4296a444669eb31e154e6cea09f378ea75b12eb` | +| B | create | 1 | v0.1 | `2057e57117100ba3efd5becf5254f47a5696bc8631c35480a5a6d2f5b1327f83` | +| B | replace | 2 | v0.2 | `23b4bbf06383c5e73e04356af079db1e6fe0827f8afdb406f849169306fa2214` | + +These observations do not establish service latency, wire round trips, +row-independent cost, or a before/after speedup for any release. They are +owned SQLite evidence for the documented behavior at the revisions above. + ## Database and Dolt invariants Every successful apply preserves the logical `dataset_id` and physical From 361263ee833e00931999d9ceef702c5e58afc431 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?T=C3=B5nis=20Ormisson?= Date: Wed, 9 Sep 2026 14:15:45 +0300 Subject: [PATCH 2/2] docs: clarify measurement fixtures --- docs/transformations.md | 13 ++++++++----- 1 file changed, 8 insertions(+), 5 deletions(-) diff --git a/docs/transformations.md b/docs/transformations.md index a128d07..ea870bd 100644 --- a/docs/transformations.md +++ b/docs/transformations.md @@ -207,11 +207,14 @@ files afterward. It covered two workloads (A: the `COMPUTE`/`IF` program above; B: `RECODE`), create versus pre-provisioned replace, 100 versus 100,000 cases, and both `apply_transformation_plan_in_place` and `apply_spss_in_place`. Each cell had one counted fresh apply, one fresh warm-up, and five fresh timed -applies; create/replace order alternated by repetition. Both sizes repeated -this input cycle in order: -`(1,1)`, `(1,NULL)`, `(NULL,1)`, `(0,1)`, `(2,2)`, `(0.1,0)`, `(-9,1)`, -`(NULL,0)`. Each fixture had numeric `source_a` and `source_b`; `source_a` -carried label `Source A` and missing rule `99`, while `source_b` carried label +applies; create/replace order alternated by repetition. Each counted, +warm-up, and timed apply started from a newly created fixture database; no +apply reused a database from another trial. Rows were generated by cycling +this input sequence in order: `(1,1)`, `(1,NULL)`, `(NULL,1)`, `(0,1)`, +`(2,2)`, `(0.1,0)`, `(-9,1)`, `(NULL,0)`. Thus the 100-row runs used 12 full +cycles plus the first four rows of a cycle, while the 100,000-row runs used +12,500 full cycles. Each fixture had numeric `source_a` and `source_b`; `source_a` carried +label `Source A` and missing rule `99`, while `source_b` carried label `Source B` and no missing rule. Replace fixtures added one nullable numeric ordinal-3 target (`target` for A, `band` for B) with label `Old target`, `F4.0` formats, ordinal measurement level, and missing rule `99`; create fixtures had