perf: scope transformation label reads to datasets - #23
Conversation
|
@codex review |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (2)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe value-label query now filters through each linked variable’s dataset. Tests cover dataset-scoped reads across loader, canonical, and SPSS surfaces, including shared labels, read counts, and malformed label isolation. ChangesValue-label dataset scoping
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~20 minutes Merge Risk: ⚪ Minimal · up to Transformation schema loading now ignores value labels belonging only to sibling datasets, preventing unrelated malformed labels and excess reads from affecting target transformations. Target behavior is covered across loader, canonical, and SPSS paths, with no remaining merge-blocking risk identified. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
@codex review |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
@codex review |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Summary
Phase 6b: scope transformation input-schema value-label reads to the target dataset's variables.
_input_schema.value_label_set.dataset_idbelongs to another dataset; preserve label order, typed values, diagnostics, transactions and public APIs.wide.py, caches, schema/index changes, dependencies and other label reads remain out of scope.RED/GREEN and checks
1f4b89b: six intentional failures on base—unrelated label rows were fetched at 0/3/40 growth and unrelated malformed labels raised before binding; target-linked malformed behavior already passed.c195a25: production only; tests immutable after RED.@codex reviewwas requested at 5597433429, but the Codex service returned an account usage-limit message (5597434497) rather than running a review. This is recorded as an external review blocker; no Codex result is claimed.Evidence
Immutable base
c6cf9e593e9c30764dac585449db19a98de60b62versus headc195a25dd43b70c4b71fecfb69c2fbbf58578fed, same Python 3.13.2 / SQLAlchemy 2.0.52 / SQLite 3.47.1, disposable SQLite only.SQL statement counts remain unchanged: loader 3, canonical 213, SPSS 216. With fixed target labels, after-change label rows remain 5/5/10 for 0, 10,000 unrelated variables and 10,000 unrelated labels. Target data, metadata, table/dataset identity, audit contents and validation results matched base; only generated audit IDs/timestamps were normalized for comparison. 166 compatibility assertions passed across 24 scenarios per revision.
Malformed behavior is explicit: an unrelated-only numeric NULL no longer poisons the target load/apply; a target-linked NULL—including a foreign-owned/shared set—still raises the exact existing
TypeErrorbefore mutation/audit. Full raw evidence and commands:../python-label-reads-evidence.mdand/tmp/python-label-reads-phase6b-VSZrJV, outside the repository. No timing or server-performance claim is made.Next phase
After user merge, proceed to Phase 7: measure memory copies and bounded batches; full streaming only if justified separately. Overall phases 0–9 remain unfinished. Do not auto-merge.