Skip to main content

Installing Atlas on UIS

Atlas installs as one UIS application: a PostgreSQL database, a Dagster code location that runs the pipelines, and a PostgREST API over the curated api_v1 views.

This guide is written for someone doing it for the first time. It tells you what will look broken and is not — that section has cost more time than every real defect combined.

Two things to know before you start

Installing starts nothing. The schedules ship stopped. No data is fetched and no external service is contacted until you turn them on. That is a deliberate go-live decision.

Both verbs exist now — this box said they did not. It claimed "loading the data and going live currently need GraphQL; UIS has no uis dagster run verb yet". Both halves are false:

  • Going live: uis dagster automation --start shipped in UIS 1.6.90 and is verified on 1.6.106 (imac, urb-agents #1005, #1043, #1149). Step 4 is a command, not a mutation.
  • Loading data: uis dagster run <job> works — imac measured api_v1_checks at ~27 s. One exception: transform_checks takes 6–15 minutes to start — slow, not broken (urb-agents #1160). Why it is slow is platform-side and tracked in urb-agents #1147.

⚠️ This sentence was corrected further down the page on 2026-09-16 and survived here, at the top, where a first-time reader meets it before any of the corrected steps. Recorded rather than quietly deleted: a stale claim in a summary outlives the same claim in the body, because the body is what gets edited when someone checks a fact.

Step 0 — see the plan without installing anything​

uis template install atlas --dry-run

This pulls the definition and installs nothing. Run it first. (From UIS 1.6.52 template list and template info point you at it too.) The plan it prints is exactly what will happen:

1. uis deploy postgresql
2. uis configure postgresql --app atlas --database atlas --namespace dagster
--secret-name-prefix atlas-database --init-file -
3. uis configure postgrest --app atlas --database atlas --schemas api_v1 --url-prefix api-atlas
4. uis deploy postgrest --app atlas
5. uis deploy dagster
6. (write code location 'atlas-data' to .uis.extend/dagster-code-locations.yaml)
7. uis deploy dagster # again, so the overlay picks it up

uis template info atlas prints what the application does once installed — cadence, volume, and the external services it contacts.

Before you install​

  • Do not run uis pull, uis stop or uis restart while anything else is running. All three stop the container that uis template install runs inside. Mid-install that leaves a database created, some services deployed, and nothing recording which — no UIS command repairs it.

    From UIS 1.6.52 the tool refuses for you, naming the command it found running, and tells you to wait. If you are certain it is safe, UIS_FORCE=1 ./uis pull overrides. On anything older there is no guard and the warning above is the whole protection.

  • uis pull --check before uis pull. The version advertised as available is not always the version that installs; trust the installed version, not the advertised one.

Step 1 — install​

uis template install atlas

Expect roughly ok=31 changed=4 failed=0, three pods in the dagster namespace, two in postgrest, one database, one role, two secrets and one IngressRoute.

The database is named atlas. (An install made without --database gets atlas_db instead — if you are looking at an older install, that is why the name differs.)

Step 2 — expect an empty API​

Immediately after install the API answers and serves zero endpoints. meta_sources, meta_endpoints and meta_dimensions return 404.

That is correct, not a failed install. The schema and grants exist from install; the data arrives on the first pipeline run, which you start next.

Step 3 — load the data​

Six jobs, in this order:

annual_sources_refresh → klass_refresh → seed_sources_refresh
→ brreg_bootstrap → brreg_change_feed → transform_and_publish
This step said FOUR jobs and ~11 minutes until 2026-09-16

Both were wrong, and the omissions are the expensive half.

brreg_bootstrap and brreg_change_feed were missing from the list. brreg_bootstrap has no schedule and no automation condition, on purpose — re-running a 1.17M-record bulk load against a populated database is the one genuinely destructive operation in this pipeline, so nothing self-triggers it. It runs exactly once, here. Skipping it leaves the organisation register empty with nothing saying why, and because nothing will ever launch it for you, an install that missed it stays wrong until somebody notices.

~11 minutes / 2.9M rows was a measurement of the first FOUR jobs (urb-agents #507) taken before brreg_bootstrap existed. atlas-data/template-info.yaml has warned since that "an earlier figure of ~11 minutes is still in circulation and understates a cold install by nearly 3x" — this page was where it was circulating.

Takes ~30 minutes on a cold install: 1772 s wall, measured end to end by imac on a factory-reset cluster (urb-agents #1027), loading ~4.1M rows. The Brreg bulk load is 501 s of that. Individual jobs vary far more than the total suggests — annual_sources_refresh alone is 474 s and transform_and_publish 484 s — so a run that looks stalled at 8 minutes may be normal. To tell slow from stuck, run uis template check atlas rather than waiting: it reports what has been pulled against what has been applied.

Why this order​

seed_sources_refresh contains raw/_migrations and runs third. The migrations are idempotent, so applying them after two source jobs have already written is safe — but it is not what anyone would design, and reordering has not been tested (urb-agents #627).

brreg_change_feed runs immediately after brreg_bootstrap because the bootstrap seeds the feed's watermark from the snapshot's own date. The feed has nothing to start from until it has run, and it refuses to start without a watermark rather than silently walking history.

Running a job​

uis dagster run annual_sources_refresh

One at a time, waiting for each to succeed before starting the next.

One job is slow to start — do not mistake it for stuck

transform_checks takes 6–15 minutes to start. It returns exit 0 and then sits NOT_STARTED before running — imac measured 364 s and 885 s (urb-agents #1160). Slow, not broken: re-launching because it looks hung gives you duplicate runs. It is not one of the six above, and it also runs in the sensor chain once automation is on. Why the start is slow is platform-side (urb-agents #1147).

Fallback: the raw mutation, if your UIS predates uis dagster run
mutation {
launchPipelineExecution(executionParams: {
selector: {
repositoryLocationName: "atlas-data",
repositoryName: "__repository__",
jobName: "annual_sources_refresh"
},
mode: "default"
}) {
__typename
... on LaunchRunSuccess { run { runId } }
... on PythonError { message }
}
}

Change jobName for each of the six. Poll runOrError(runId:) until the status is SUCCESS before starting the next.

The Dagster UI can do the same through a browser, but it is internal-only with no authentication, so reaching it is a decision about your own cluster rather than something this guide can prescribe.

Step 4 — go live​

Enabling the schedules is what makes Atlas keep itself current. Until you do, it holds whatever you loaded in step 3.

Enabling the schedules does not backfill

Every automation condition is on_cron, which means next fire, not catch-up. Enabling them on a Thursday means the first automatic raw refresh is Sunday 02:00. That is why step 3 exists.

This page said the opposite until 2026-09-16

It said uis dagster automation cannot set state and sent you to raw startSchedule / startSensor GraphQL. That is wrong. imac used uis dagster automation --start and --stop on UIS 1.6.106 and they set it correctly (urb-agents #1149). The sentence was also copied into atlas-status.py, which printed it to operators, and into a unit test that asserted the flag must never be mentioned — so a claim about somebody else's CLI became unfalsifiable from inside this repo. All three are corrected together.

Use uis dagster automation --start (and --stop). Verified on UIS 1.6.106; if you are on an older UIS and it does not work, the raw GraphQL below is the fallback rather than the instruction.

Fallback: the raw mutations

Note the return types differ: startSchedule returns ScheduleStateResult, startSensor returns Sensor. Using the wrong one gives a bare HTTP 400.

Once enabled, Atlas polls on this cadence (Europe/Oslo):

whenwhat
Sunday 02:00~37 annual public-sector sources — SSB, FHI, Bufdir
1st of month, 01:00SSB Klass classifications
Daily 05:00dbt transform and publish — no external calls

Step 5 — verify​

-- structure
raw tables 47 · marts tables 64 · api_v1 views 13

-- the assertion to lead with
select count(*) from raw.brreg_enheter; -- 122

If brreg_enheter is 0, seed_sources_refresh did not run. 122 is a fixed reference list, so it is the one number worth remembering — everything else depends on upstream volumes.

For the rest, assert properties rather than remembered totals:

Every raw source that declares loaded_at has loaded, except redcross_branches and redcross_branch_activities, which are parked pending a credential.

API checks:

GET / -> 200, 14 OpenAPI paths (13 views + `/`)
GET /meta_sources -> 200
GET /meta_endpoints -> 200
GET /meta_dimensions -> 200

And:

uis dagster verify # A/B/C PASS, code location LOADED
uis dagster automation --expect running # verify alone passes whether or not schedules are on
A loaded install is not a validated one

Everything above checks that the data is there. Atlas also defines 679 asset checks — 675 from dbt, 4 on the api_v1 surface — and on a fresh install none of them has run.

They are reached only through a chain of three instigators that all ship stopped: transform_daily → the sensor to api_v1_checks → the sensor to transform_checks. So the suite does not first run at 05:00 tomorrow. It first runs after the next transform_daily following go-live, and if you never enable automation it never runs at all.

Four of them you can run now:

uis dagster run api_v1_checks # launches in ~0.5 s, runs in ~27 s

Those four cover the public API — row counts against the marts they wrap, column COMMENTs, and the anonymous role's grants — which is the part worth checking before anyone reads the data.

The other 681 can be launched too — budget 6–15 minutes. uis dagster run transform_checks returns exit 0; imac measured two launches at 364 s and 885 s to start, the first landing 681 checks SUCCEEDED (urb-agents #1160). The run sits NOT_STARTED for that whole time and then runs.

This paragraph told you not to try it

It said the call "does not return within 300 s, exceeds the client's 60 s budget, and leaves an unsubmitted run behind — do not retry it". True when written (#1064); false on this build. The 60 s budget went in UIS 1.6.104 and the deadline is now 900 s.

It is slow, not broken. Re-launching because it looks hung is how you get duplicate runs — which is the actual hazard, and the opposite of the one the old warning described.

:::

See your data​

The install is not finished until you have read a row out of it. Every endpoint below is live once step 3 has run:

/coverage_gap_barnefattigdom /indicator_latest_values /indicator_missing_kommuner
/indicator_summary /kommune_local_chapters /ngo_index
/ngo_overview /distrikt_summary /activity_catalog
/bufdir_indicator_alias /meta_sources /meta_endpoints
/meta_dimensions
GET /coverage_gap_barnefattigdom?limit=2
{"kommune_nr":"0301","kommune_name":"Oslo","fylke_name":"Oslo","year":2024,"value_pct":13,"personer":130343}
{"kommune_nr":"1101","kommune_name":"Eigersund","fylke_name":"Rogaland","year":2024,"value_pct":8.6,"personer":3270}

That is what you built — child-poverty coverage by kommune, joined to the NGOs that respond to it.

/indicator_missing_kommuner is worth opening early: it exposes the same discrepancy as the 17 warnings below, which is a friendlier introduction to them than a red test run.

Getting your machine to reach it​

The route matches on hostname, so api-atlas.<your-domain> must resolve to the ingress. On a real cluster with real DNS that is already true.

On a laptop this is the last undocumented step

How .localhost resolves depends on your Kubernetes distribution and your OS, and this guide will not guess for you. If the hostname does not resolve, port-forward instead — it works regardless of DNS:

kubectl port-forward -n postgrest svc/<postgrest-service> 8080:80
curl -H 'Host: api-atlas.localhost' 'http://localhost:8080/coverage_gap_barnefattigdom?limit=2'

The Host header is not optional: Traefik routes on it, so a request without it gets a 404 from Traefik rather than an answer from Atlas.

Things that look like failure and are not​

This section exists because every item in it has produced a wrong conclusion by someone experienced.

Every run you just launched alerts as failed, fifteen minutes later​

Fifteen minutes after your data load succeeds, monitoring may raise PodNotReady for every job you ran. On a production install today that was six alerts, one per job, all six having SUCCEEDED.

A completed Kubernetes Job pod is 0/1 and reports ready=false permanently — that is what a finished pod looks like. The alert rule matches any not-ready pod, and its own description says "Too long for a deployment": the annotation names the intended scope and the expression does not enforce it.

So: a PodNotReady alert naming a dagster-run-* pod is expected and is not a failure. Check the run's status in Dagster, not the pod's readiness.

This one has a compound cost

On that install the six false alerts were sitting beside one real alert that had been firing for four days — a dead metrics exporter — and the real one was invisible among them. A rule that cannot stop firing does not add noise; it spends the credibility of the channel. Fixing the rule is UIS's; knowing the alerts are false is yours.

transform_checks looks hung. It is slow.​

The launch call may not return for over 45 seconds, and the run sits at NOT_STARTED for more than a minute before succeeding in about 105 s.

All 647 dbt checks execute inside one op, so the two probes anyone reaches for — reading the definition and building the execution plan — both return "1 op, 1 step" in a fraction of a second and cannot see the weight. Two experienced testers concluded "hung" on different clusters, sixteen days apart. Both were wrong.

Judge it by polling the run, never by whether the launch call returns. Anything that alerts on it must poll the run, or it will report a false failure every night.

17 dbt checks do not pass, on a correct install​

transform_checks succeeds while reporting 17 non-passing checks out of 647. All 17 are relationships tests from indicator models to dim_kommune / dim_fylke.

They are configured severity: warn deliberately. Historical SSB series carry kommune and fylke codes retired in the 2020 reorganisation; the dimensions carry current Klass codes. The model schema says so in place: "Historical fylker (pre-2020 01-20 numbering) may appear — warn."

They surface as non-passing asset checks in Dagster because a dbt warning is not a pass. A run that succeeds with exactly these 17 is a correct install.

This is measured, not assumed. Two installs on different topologies, different hardware and different UIS versions both report 630 succeeded · 17 failed · 647 total, and the sets were compared rather than the counts — the same 12 tables, the same fylke-only and kommune-only split. Two lists that both totalled 17 and named different checks would have looked like agreement and meant the opposite.

⚠️ Documented is not the same as acceptable. The semantic question underneath — whether Atlas covers pseudo-regions or merely represents them — is open. The underlying question — whether Atlas covers pseudo-regions or merely represents them — is open and tracked in INVESTIGATE-ssb-pseudo-regions.

meta_endpoints has 91 rows against 13 views​

Not 78 phantom endpoints. The ratio has been 7 × views on every install measured. Watch the ratio, not the number — a changed ratio is the signal.

Row estimates read zero on a populated database​

pg_stat_user_tables.n_live_tup is a planner estimate and can read 0 for every table while the schema holds hundreds of megabytes. Pair every estimate with one count(*).

kubectl get ingress shows nothing​

The route is a Traefik IngressRoute CRD. kubectl get ingress returns an empty result rather than an error. Use kubectl get ingressroute -n postgrest.

A path-prefixed URL returns a Traefik 404​

The route matches on hostname, not a path prefix. api-atlas.<your-domain> must resolve to the ingress. Reaching it from your own machine is a DNS or hosts-file question about your setup, not something the install can do for you.

The API 404s even though the view exists, with rows and grants​

🔴 This is the one that will waste your evening, and it is not a false failure — it is a real outage with a one-line fix.

You check the database and everything is right: the view is there, it returns rows, atlas_web_anon has SELECT. And GET /brreg_enhet still returns 404 on every replica.

PostgREST answers from a cached schema. It does not read the catalogue per request. If that cache was last loaded at a moment when the view did not exist, it keeps serving that absence no matter what you do to the database.

recreate the view, no reload 404 — view exists, 1.17M rows, grants correct
issue the reload 200 — instant

The fix: materialise the api_v1 asset, or run transform_and_publish. That asset is what issues NOTIFY pgrst, 'reload schema'. It takes seconds.

⚠️ When this bites. A transform run that fails partway leaves some api_v1 views missing; nothing restores them until a run succeeds. If PostgREST reloads while they are missing — a pod restart is enough — the absence is cached. Repairing the database by hand then fixes everything except the endpoint, which is exactly the state that makes it confusing.

🔵 It is not a PostgREST fault. It is serving the last schema it was told about, correctly.

api_v1 views briefly lose their descriptions​

After a failed run, the views that were recreated carry no COMMENT, so PostgREST's OpenAPI descriptions come back empty — the data is right and the published documentation is gone. A successful transform_and_publish restores them, because the generated apply is what sets them.

Bounded, not permanent: it lasts as long as the failure does. Worth knowing so an empty API doc is not mistaken for a schema problem.

Removing and reinstalling​

uis template remove atlas works for installs made by uis template install — it reads the record that install writes. An application put in place before the template path existed has no such record and remove exits 1.

Do not hand-write a record to unlock removal

Per-app names derive from that record. A hand-written one that disagrees with reality can point a removal at the wrong live application.

No UIS command drops a database, at any version. --purge drops per-app roles and secrets; the atlas and dagster databases survive every removal, so a reinstall lands on the existing schema. Dropping is manual DDL and the order matters:

DROP DATABASE atlas; -- or atlas_db on an older install
DROP ROLE atlas; -- blocked until the database is gone
DROP DATABASE dagster;

Leave atlas_web_anon and atlas_authenticator — configure postgrest re-credentials them. Leave the dagster role.

Use uis undeploy dagster rather than helm uninstall; the playbook waits for pods to terminate and tells you what it preserves.

If you delete the dagster namespace, run uis secrets generate && uis secrets apply before reinstalling — the Dagster setup reads a generated password from a secret in that namespace, and it fails loudly with that instruction if it is missing. You are not expected to know that password.

Known gaps​

gapstatus
No uis dagster automation --startCLOSED — it shipped and this table did not notice. Verified working on UIS 1.6.106 (imac, urb-agents #1149). Listed as a gap long after it stopped being one, which is why atlas-status.py was still printing GraphQL instructions.
uis dagster run <job>Works. api_v1_checks ~27 s; transform_checks 6–15 min to start, exit 0 (imac, urb-agents #1160). Why the start is slow is platform-side, with tor-agent (urb-agents #1147).
Job order is documented, not enforcedsee step 3
transform_checks start latencytracked in INVESTIGATE-transform-job-decomposition