Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
135 changes: 135 additions & 0 deletions docs/transformations.md
Original file line number Diff line number Diff line change
Expand Up @@ -135,6 +135,141 @@ openstatspec apply-spss --database-url sqlite:///survey.sqlite \

On Dolt, also pass `--expected-branch` and `--expected-head`.

## Numeric targets: create versus replace

`COMPUTE` chooses its canonical target mode from the live schema. An absent
numeric target lowers to `assign` with `target_mode=create`; it appends one
nullable numeric column and one catalog variable. An existing numeric target
lowers to `target_mode=replace`. Replace keeps the variable ID, ordinal, and
metadata that no later operation explicitly changes. `IF` requires an existing
target, so a new target must be initialized first. Variable-source `NULL`
propagates through assignment; a predicate changes a row only when its SQL
truth value is `TRUE`.

For example, this is a create-or-replace program for a numeric target. On
SQLite/PostgreSQL it creates `target` when absent; with a pre-existing numeric
`target` it updates the same target instead:

```text
COMPUTE target = source_a.
IF (source_a = 1 AND source_b = 1) target = 1.
IF (target = 1 AND source_b = 1) target = 2.
VARIABLE LABELS target 'Numeric target'.
VALUE LABELS target 0 'No' 1 'Yes'.
FORMATS target (F8.1).
VARIABLE LEVEL target (NOMINAL).
EXECUTE.
```

`RECODE ... INTO` always requests a new target and rejects an existing target
name. A canonical replace recode must use the same source and target name.
Without `ELSE`, create recodes use `system_missing` for unmatched rows, while
replace recodes use `copy`. To pre-provision `band` and preserve the create
semantics, initialize and recode it explicitly. The absent-target form is:

```text
RECODE source_a (1 = 0.1) (2 = 1) INTO band.
```

The pre-existing-target form is:

```text
COMPUTE band = source_a.
RECODE band (1 = 0.1) (2 = 1) (ELSE = SYSMIS).
```

The second form has an extra ordered assignment and is not a pure DDL-overhead
comparison. Omitting its `ELSE` would copy unmatched source values instead of
matching the create form. For canonical plans, create recodes use
`target_mode=create` and replace recodes use `source=target`,
`target_mode=replace`.

On SQLite and PostgreSQL, numeric create is supported only where the adapter
has the atomic native transaction boundary. MySQL, MariaDB, and Dolt reject a
create target with `schema_change_not_atomic`. A caller must use a separate,
explicit, versioned provisioning action that creates both the nullable physical
column and its matching normative variable, including a unique variable ID,
the same physical binding, and the next ordinal. Provisioning is not an apply
step, and this package does not hide it behind a new helper or API. See the
specification's [In-Place Transformation Binding 0.2 target-creation
rules](https://github.com/OpenStatSpec/specification/blob/main/docs/transformation-plan-sql-binding-0.2.md#target-creation-by-sql-profile).
For Dolt, commit provisioning separately; the later apply starts from that
clean committed `HEAD` and supplies it as `expected_head`. The SQLite matrix
below is local evidence only and is not service or cross-engine evidence.

## Local numeric-target measurement

The following is a bounded observation of the existing workflow, not a
release-to-release benchmark. A disposable standard-library harness created a
fresh file-backed SQLite database for every variant and repetition, kept
fixture setup separate from the timed public call, and removed all temporary
files afterward. It covered two workloads (A: the `COMPUTE`/`IF` program above;
B: `RECODE`), create versus pre-provisioned replace, 100 versus 100,000 cases,
and both `apply_transformation_plan_in_place` and `apply_spss_in_place`. Each
cell had one counted fresh apply, one fresh warm-up, and five fresh timed
applies; create/replace order alternated by repetition. Each counted,
warm-up, and timed apply started from a newly created fixture database; no
apply reused a database from another trial. Rows were generated by cycling
this input sequence in order: `(1,1)`, `(1,NULL)`, `(NULL,1)`, `(0,1)`,
`(2,2)`, `(0.1,0)`, `(-9,1)`, `(NULL,0)`. Thus the 100-row runs used 12 full
cycles plus the first four rows of a cycle, while the 100,000-row runs used
12,500 full cycles. Each fixture had numeric `source_a` and `source_b`; `source_a` carried
label `Source A` and missing rule `99`, while `source_b` carried label
`Source B` and no missing rule. Replace fixtures added one nullable numeric
ordinal-3 target (`target` for A, `band` for B) with label `Old target`, `F4.0`
formats, ordinal measurement level, and missing rule `99`; create fixtures had
no target column.

The counted value is an outermost SQLAlchemy `Connection.execute` or
`Connection.exec_driver_sql` invocation scoped to the owned database. It
includes profile probes, catalog checks, reflection, explicit `BEGIN`, DDL,
DML, and audit work; one executemany call counts once. It is not a driver wire
round-trip count. Timed values are the public apply call only: setup,
validation, and cleanup are excluded. Every counted and timed result checked
ordered rows, explicit `NULL`s, binary64 bits (including `0.1`), source and
target identity, catalog/audit provenance, and absence of forbidden copy or
snapshot artifacts before being retained.

Environment and revisions: Python 3.13.2, SQLite 3.47.1, SQLAlchemy 2.0.52,
Linux 6.6.87.2-microsoft-standard-WSL2 on x86_64 with 12 logical CPUs and
local temporary-file storage; observed file settings were
`journal_mode=delete` and `synchronous=2` (`FULL`). The Python source was
`63a66d6589aeb8fdc4820d5e7217fc13c61ef3b9`; the reviewed specification
reference was `930345b922af4f43b0e622be02d4b12ebfeb08eb`. Fixture setup was
recorded separately and excluded from the table.

| Workload | Cases | API | Mode | SQLAlchemy executes | Median ms | Min ms | Max ms |
| --- | ---: | --- | --- | ---: | ---: | ---: | ---: |
| A | 100 | canonical plan | create | 234 | 76.5 | 65.8 | 78.4 |
| A | 100 | SPSS | create | 237 | 69.3 | 66.8 | 78.2 |
| A | 100 | canonical plan | replace | 222 | 67.6 | 56.6 | 88.3 |
| A | 100 | SPSS | replace | 225 | 71.3 | 69.3 | 88.5 |
| A | 100,000 | canonical plan | create | 234 | 168.8 | 134.1 | 172.3 |
| A | 100,000 | SPSS | create | 237 | 160.2 | 149.4 | 170.3 |
| A | 100,000 | canonical plan | replace | 222 | 162.7 | 146.0 | 173.6 |
| A | 100,000 | SPSS | replace | 225 | 155.4 | 144.8 | 177.6 |
| B | 100 | canonical plan | create | 224 | 61.6 | 55.6 | 76.6 |
| B | 100 | SPSS | create | 227 | 77.1 | 65.5 | 312.8 |
| B | 100 | canonical plan | replace | 213 | 64.2 | 60.2 | 223.1 |
| B | 100 | SPSS | replace | 216 | 80.9 | 62.2 | 237.5 |
| B | 100,000 | canonical plan | create | 224 | 149.4 | 131.7 | 193.4 |
| B | 100,000 | SPSS | create | 227 | 144.1 | 128.3 | 158.4 |
| B | 100,000 | canonical plan | replace | 213 | 169.8 | 163.6 | 186.6 |
| B | 100,000 | SPSS | replace | 216 | 192.8 | 171.4 | 202.8 |

The canonical plan operation/hash identities for the same runs were:

| Workload | Mode | Operations | Contract | Canonical plan SHA-256 |
| --- | --- | ---: | --- | --- |
| A | create | 8 | v0.2 | `6b17a220f9be9d2d29b671dd5eb4598ea34ae459e4d3b8e84a7daafbe3ec11ef` |
| A | replace | 8 | v0.2 | `a239cac39d3fade1207228bee4296a444669eb31e154e6cea09f378ea75b12eb` |
| B | create | 1 | v0.1 | `2057e57117100ba3efd5becf5254f47a5696bc8631c35480a5a6d2f5b1327f83` |
| B | replace | 2 | v0.2 | `23b4bbf06383c5e73e04356af079db1e6fe0827f8afdb406f849169306fa2214` |

These observations do not establish service latency, wire round trips,
row-independent cost, or a before/after speedup for any release. They are
owned SQLite evidence for the documented behavior at the revisions above.

## Database and Dolt invariants

Every successful apply preserves the logical `dataset_id` and physical
Expand Down