Skip to main content

PLAN: split checks out of the transform build — the minimum START_TIMEOUT unblock

IMPLEMENTATION RULES: Before implementing this plan, read and follow:

Status: Completed (2026-08-24) — verified by imac, round 6 PASS

Goal: Clear RunFailureReason.START_TIMEOUT with the smallest change that works, so integration criteria 10–12 can be verified — without building machinery that a later architectural decision would delete.

Last Updated: 2026-08-24

Investigation: INVESTIGATE-transform-job-decomposition.md

Priority: High — last blocker on a four-round integration test.


Problem Summary

transform_and_publish plans 711 events, 90.5% of them asset checks. Dagster builds that plan over gRPC before the run pod exists, and it doesn't finish inside start_timeout_seconds: 300.

Measured candidate shapes (from the real manifest; the model reproduces the tester's 711 exactly):

ShapeWorst plan
Today — monolith711
Build only, checks excluded65
Checks in one job644
Layer jobs carrying their checks344

Scope — deliberately narrow

Terje's decision (2026-08-24): land the minimum unblock now, park the elaborate machinery, then pilot the idiomatic Dagster shape (declarative automation + freshness) before committing to a direction.

In scope: separate the dbt build from its checks, and chain them in the correct order.

Explicitly parked — build only if hand-built jobs remain the primary mechanism after the pilot:

  • layer-based job splitting
  • the get_group_name translator override
  • the CI plan-size budget

Parked, not rejected. If declarative automation wins, all three become work we would build and then delete.


Phase 1: Make the dbt command follow the selection

The dbt asset runs dbt build unconditionally, so excluding checks from Dagster's plan would still leave dbt running all 784 tests inside the step and reporting them as untyped observations — the pre-round-2 behaviour we deliberately left.

Tasks

  • 1.1 In atlas_dbt_models, choose the dbt command from the execution context: assets only → dbt run; checks only → dbt test; both → dbt build (the local "materialise everything" case).
  • 1.2 Keep the api_v1_rowcount_matches_marts exclusion — still misordered inside a single dbt build.
  • 1.3 Verify all three selections locally.

Validation

Selecting assets runs dbt run; selecting checks runs dbt test; selecting both still runs dbt build.


Phase 2: Two jobs, correct order

Tasks

  • 2.1 transform_and_publish selects dbt assets and api_v1, both .without_checks().
  • 2.2 A transform_checks job selecting the checks.
  • 2.3 Chain them: checks run after the publish. dbt run rebuilds marts.mart_* by swapping tables and dropping the old ones CASCADE, destroying the api_v1.* views; rowcount_matches_marts compares the two, so it is meaningless until the publish has re-created them. Prefer a run-status sensor over offset schedules — an offset encodes a guess about duration.
  • 2.4 Both selections stay pattern-based, never enumerated lists (the zero-edit requirement).
  • 2.5 Measure the resulting plan sizes and confirm the build job is ~65.

Validation

Build job ≈65 events; both jobs run green end-to-end locally, checks after publish.

Outcome (2026-08-24) — phases 1 and 2 complete

Jobassetschecksplan
transform_and_publish65065 (was 711)
api_v1_checks011
transform_checks0644644

All three run green locally, each with the right dbt verb — logged so it is visible in the run, not inferred.

A judgment call beyond the brief: the checks are split in two, not one. The tactical fix would leave a single 645-event checks job, which is within noise of the 711 that already fails — so criterion 11 (the api_v1 rowcount check) would have ridden on exactly the risk this plan exists to remove. The split is semantic rather than convenient: the api_v1 check is a publish gate ("does the public surface match the marts it wraps?"), the 644 dbt tests are data quality ("is the data sound?"). Different questions, different audiences. Chaining them means a publish-gate failure is never queued behind, or hidden by, the bulk suite. Both selections stay pattern-based.

A regression caught before it shipped. The obvious implementation of "build without checks" is dbt run — and dbt run skips seeds. Atlas has 16 seed assets (ref_*, dim_postnummer, the sources manifest) that models join against, so on a fresh database — the tester's exact starting state — the models referencing them would fail. Verified by wiping marts and api_v1 and re-running: dbt run gave PASS=48, build --exclude-resource-type test gives PASS=64 (48 models + 16 seeds), still with zero tests. The build job now uses the latter.

Ordering is enforced by sensors, not offsets. transform_and_publishapi_v1_checkstransform_checks, chained on run success. An offset schedule would encode a guess about how long the build takes; a sensor encodes the actual dependency. The dependency is real: dbt swaps the marts tables and drops the old ones CASCADE, destroying the api_v1 views, so the publish gate is meaningless until the publish has re-created them.

⚠️ transform_checks is still 644 events and carries the same startability risk. That is the known limit of a tactical fix, it is why the platform's timeout bump matters as margin, and bounding it durably is the architectural question — not this plan's job.


Phase 3: Re-declare

Tasks

  • 3.1 Merge, record the image tag.
  • 3.2 Update ~/home/ai-developer/for-ops-atlas-testable.md: new tag, the two-job shape, and criteria 10–12 restated against it.
  • 3.3 State plainly that this is the tactical unblock and the architectural review is separate, so the tester isn't verifying a shape that may change.

Validation

A PASS on 10–12 from imac.

Outcome (2026-08-24) — criteria 10–12 passed before this shipped

The platform's start_timeout_seconds 300 → 900 landed first, and imac re-ran the old monolith under it: started, completed, dbt 820 PASS / 1 known WARN / 0 ERROR, 13 api_v1 views, rowcount_matches_marts success. Criteria 10–12 are closed, and the tester's own summary is the fair one: "your code was never the blocker; the plan just could not be built in time to prove it."

Which means this plan's value is no longer unblocking — it is headroom, and the tester measured exactly how much: 564s of the 900s budget consumed at 41 sources, leaving ~60% growth, or about 27 more sources before the limit returns. This split takes the write path to ~52s of that budget and leaves the checks path at ~511s.

Two honest consequences:

  1. Phase 3's re-declaration is not urgent — nothing is blocked on it. But the shape imac verified (one monolithic job) is no longer the shape Atlas ships, so leaving it unverified means the verified artefact and the real one have diverged. Declared as low priority.
  2. The remaining risk is now specifically the 644-event checks job, not the pipeline. That is the architectural question, and it is the one the pilot should answer.

Acceptance Criteria

  • transform_and_publish plans ≈65 events, not 711.
  • dbt tests still run — as dbt test in their own job, still surfacing as Dagster asset checks.
  • Checks execute after the api_v1 publish, never before.
  • Job membership stays pattern-based; adding a source touches no job definition.
  • dbt build, npm test, check-osmosis.sh still pass.

Out of Scope

  • Everything in the "parked" list above.
  • The declarative-automation / freshness pilot — its own plan after this lands.
  • The 41-source live ingest run — approval still open with Terje; do not run it.

Closed 2026-08-24 — verified

Round 6 PASS. The three-job shape holds in-cluster and the concurrency bound survives.

Worth recording that this plan's purpose changed mid-flight: it was written to unblock criteria 10–12, and the platform's timeout bump closed those first. What it delivered instead is headroom — the write path went from 564s of plan construction to about 52s. Against the tester's measurement (~15.8 plan events per source, ~1135-event budget at 900s) that is the difference between a limit that returns at ~68 sources and one that does not realistically return at all for the build path.

transform_checks remains ~644 events. Bounding that durably is still open, in INVESTIGATE-transform-job-decomposition.