feat(workflows): compose workflows with a unified execution tree - #4764
markuswondrak wants to merge 14 commits into
Conversation
Keep invocation results and workflow bindings on the same execution occurrence. Isolate fan-out contexts and resume persisted expansions through one executor. Assisted-by: OpenCode (model: gpt-6-astra, autonomous)
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Child step implementation exceptions are incorrectly converted into ordinary workflow failures and may be swallowed by continue_on_error.
Review effort: Balanced
Findings: 1
What changed in this PR
Introduces composable workflows backed by a unified, persisted execution tree.
Changes:
- Adds scoped workflow calls with typed inputs and declared outputs.
- Unifies execution, resume, fan-out, and nested-state persistence.
- Adds CLI reporting, documentation, and regression coverage.
| File | Description |
|---|---|
src/specify_cli/workflows/_commands.py |
Reports nested scopes and gates. |
src/specify_cli/workflows/_execution.py |
Implements tree-based execution. |
src/specify_cli/workflows/__init__.py |
Registers workflow steps. |
src/specify_cli/workflows/command_resume.py |
Documents crash recovery. |
src/specify_cli/workflows/command_status.py |
Displays composed scopes. |
src/specify_cli/workflows/composition.py |
Handles workflow boundaries. |
src/specify_cli/workflows/engine.py |
Integrates persistence and execution. |
src/specify_cli/workflows/step/gate/__init__.py |
Normalizes gate messages. |
src/specify_cli/workflows/step/workflow/__init__.py |
Defines the workflow step. |
tests/specify_cli/workflows/test_command_status.py |
Tests composed CLI lifecycle. |
tests/workflows/test_composition_execution.py |
Covers composition and resume. |
design/workflow-step.md |
Updates execution guidance. |
docs/reference/workflows.md |
Documents composition semantics. |
workflows/ARCHITECTURE.md |
Describes the execution tree. |
workflows/PUBLISHING.md |
Adds workflow-step validation guidance. |
workflows/README.md |
Lists the new step type. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: gpt-5.6-terra, autonomous)
Assisted-by: OpenCode (model: deepseek-v4.1-flash, autonomous)
Assisted-by: OpenCode (model: deepseek-v4.1-flash, autonomous)
Assisted-by: OpenCode (model: deepseek-v4.1-flash, autonomous)
Assisted-by: OpenCode (model: deepseek-v4.1-flash, autonomous)
Preserve the item traversal's projection decision for missing step types. Cover sequential and parallel execution, inherited aliases, unnamed templates, and resume after reinstalling the implementation. Assisted-by: OpenCode (model: gpt-6-astra, autonomous)
|
Updated through The cleanup narrows call-boundary recovery, preserves incomplete calls on rebind errors, removes Current workflow validation: 1,373 passed, 1 skipped; Ruff and The error-handling follow-up proposal is explained in the reply to the review thread. Posted on behalf of @markuswondrak by OpenCode (model: gpt-6-astra / github-copilot/gpt-6-astra, autonomous mode with user-directed scope); review summary and PR update fully AI-drafted, final fix and tests AI-authored and automatically verified. Earlier cleanup commits carry their own model disclosures. |
| with self.state._lock: | ||
| if self.state._checkpoint_failed: | ||
| raise CheckpointError("A previous checkpoint failed") | ||
| if node is not None: | ||
| node.update(changes or {}) | ||
| if context is not None and "result" in node: | ||
| self.project( | ||
| context, | ||
| name, | ||
| node["result"], | ||
| public=public, | ||
| qualified=qualified or name, | ||
| ) | ||
| self.state.save() |

Summary
Reimplements workflow composition from main (
c00dc055) with one persisted execution tree and one executor. Replaces #4724 and implements the scoped single-run model approved in #4680 (comment), with the explicit resume extension described below.Closes #4680.
Included workflows receive private, strictly bound inputs and return only declared outputs alongside workflow/status/error metadata. Targets must exactly match safe, installed, enabled IDs in the current project. Overlay-resolved definitions are bound at the call site, cycles are path-based, and included depth is limited to 16. The approved clarification supersedes the original issue's child-run and
run_id-output proposal.Current head:
6792bea1aaff941f1767a5c632da84d72f183337.Why replace #4724
The previous implementation accumulated separate scope identities, result keys, cursor state, snapshot files, and execution paths. Its review identified a real fan-out isolation bug: nested calls had distinct scope keys but shared the caller-result alias.
Each execution occurrence now owns its result, children, and workflow binding. Authored step IDs are local expression aliases; concurrent items have independent contexts. Execution and replay traverse the same tree, including fan-out. Unused legacy execution adapters were removed; their tests exercise the public engine path.
The current diff from the merge base (
c00dc055) is 17 files, 3,706 additions / 594 deletions, including 1,181 production additions / 486 deletions. These replace the smaller initial-PR figures: subsequent review added regression coverage and centralized traversal, projection, validation, and event rules. Local planning documents are not included.Deliberate deviations
fan:template:index) and the orderedfan.output.resultsremain available. This addresses the parallel-item collision in feat(workflows): compose installed workflows via a scoped workflow step #4724.Architecture and contracts
_execution.pyowns tree construction/validation, traversal/replay, control flow, workflow scope entry, projections, and event emission.composition.pyowns target resolution, strict binding, declared outputs, andCallError.engine.pykeeps public lifecycle, input coercion/merge, and RunState persistence.inside_fan_outare transitive. Child gates require explicit root-to-child verdict mapping.continue_on_errorcan handle reported child failures and initial binding/output contract violations. Step exceptions, expression errors (including call input/output expressions), rebind errors, and checkpoint errors propagate. Pauses, aborts, and unknown step types remain terminal. A resolverRuntimeErroris not converted into a call failure.PAUSEDandFAILEDruns can resume. The earlier crash-resume claim is withdrawn:RUNNINGcheckpoints are rejected. No ownership/lease mechanism is added. External effects before their completion checkpoint remain at-least-once.Evidence
test_concurrent_nested_calls_keep_downstream_aliases_localwas run against #4724 commit7ece7a16via an isolated import path. It failed with consumed outputs{1: 1, 2: 1}instead of{1: 1, 2: 2}and passes here.Regression coverage includes call exception propagation, rebind preservation, transitive fan-out restrictions, frozen branch/custom expansion replay, qualified aliases/events, shared snapshots, malformed-tree rejection before writes, checkpoint/log failures, strict targets/inputs/outputs, depth/cycles, and legacy adaptation. The composed-gate CLI test covers run → JSON/human status → resume with explicitly mapped input.
The final fan-out fix extends
test_unknown_fan_out_template_step_always_fails_despite_continue_on_error: four combinations (sequential/parallel, named/unnamed template) failed before the fix because missing implementations incorrectly published item aliases. All pass afterward, including a same-name inherited parent result and successful resume after re-registering the implementation. Direct comparison withc00dc055now produces only the fan-out container result, with matching per-item failure events.Current verification
uv sync --extra test, then this worktree's.venv/bin/python -m pytest tests/test_workflows.py tests/workflows tests/specify_cli/workflows -q -p no:cacheprovider: 1,373 passed, 1 skipped (26.25 seconds).uvx ruff@0.15.0 check src testsandgit diff --check: pass.c00dc055.Historical evidence and remaining limitation
c00dc055(preset-update missing-argument wording and three locale-sensitive checksum expectations). The full repository suite has not been rerun after this cleanup. The current result above is for all workflow suites.Intentionally changed tests
test_checkpoint_failure_never_overwrites_committed_progressbecametest_checkpoint_failure_leaves_running_run_not_resumable; crash-resume expectations were removed/inverted whenRUNNINGresume was withdrawn.test_rebind_failure_has_one_failed_caller_outcomenow asserts propagation, runFAILED, and an unchanged call node rather than a recoverable call failure.test_output_failure_retries_only_finalizationinjectsCallErrorfor a contract violation.test_output_expression_failure_retries_only_finalizationseparately covers propagating expression errors without repeating children.execute()while retaining their behavioral coverage.Out of scope
Crash recovery/run ownership/leases, a dedicated expression-error type and consistent recoverability policy, a direct occurrence-addressed gate-answer API, implicit input propagation or parent-default inheritance, invalidation of completed dependent work, rejecting unknown root resume inputs, an execution-position value object, and migration of private PR checkpoint formats.
AI disclosure
Implemented and updated on behalf of @markuswondrak using OpenCode in autonomous mode with user-directed scope. This update used gpt-6-astra (
github-copilot/gpt-6-astra) for review, the final fan-out fix and regression tests, automated verification, commit/push, and this fully AI-drafted PR description. Intermediate cleanup commits disclose gpt-5.6-terra and deepseek-v4.1-flash individually in theirAssisted-by:trailers. The original rewrite and its AI-assisted #4724 history are retained. Human line-by-line review or manual testing is not attested.