Installing Atlas on UIS
Atlas installs as one UIS application: a PostgreSQL database, a Dagster code location that runs the
pipelines, and a PostgREST API over the curated api_v1 views.
This guide is written for someone doing it for the first time. It tells you what will look broken and is not — that section has cost more time than every real defect combined.
Installing starts nothing. The schedules ship stopped. No data is fetched and no external service is contacted until you turn them on. That is a deliberate go-live decision.
Both verbs exist now — this box said they did not. It claimed "loading the data and going live
currently need GraphQL; UIS has no uis dagster run verb yet". Both halves are false:
- Going live:
uis dagster automation --startshipped in UIS 1.6.90 and is verified on 1.6.106 (imac, urb-agents #1005, #1043, #1149). Step 4 is a command, not a mutation. - Loading data:
uis dagster run <job>works — imac measuredapi_v1_checksat ~27 s. One exception:transform_checkstakes 6–15 minutes to start — slow, not broken (urb-agents #1160). Why it is slow is platform-side and tracked in urb-agents #1147.
⚠️ This sentence was corrected further down the page on 2026-09-16 and survived here, at the top, where a first-time reader meets it before any of the corrected steps. Recorded rather than quietly deleted: a stale claim in a summary outlives the same claim in the body, because the body is what gets edited when someone checks a fact.
Step 0 — see the plan without installing anything
uis template install atlas --dry-run
This pulls the definition and installs nothing. Run it first. (From UIS 1.6.52 template list
and template info point you at it too.) The plan it prints is exactly what
will happen:
1. uis deploy postgresql
2. uis configure postgresql --app atlas --database atlas --namespace dagster
--secret-name-prefix atlas-database --init-file -
3. uis configure postgrest --app atlas --database atlas --schemas api_v1 --url-prefix api-atlas
4. uis deploy postgrest --app atlas
5. uis deploy dagster
6. (write code location 'atlas-data' to .uis.extend/dagster-code-locations.yaml)
7. uis deploy dagster # again, so the overlay picks it up
uis template info atlas prints what the application does once installed — cadence, volume, and the
external services it contacts.
Before you install
-
Do not run
uis pull,uis stoporuis restartwhile anything else is running. All three stop the container thatuis template installruns inside. Mid-install that leaves a database created, some services deployed, and nothing recording which — no UIS command repairs it.From UIS 1.6.52 the tool refuses for you, naming the command it found running, and tells you to wait. If you are certain it is safe,
UIS_FORCE=1 ./uis pulloverrides. On anything older there is no guard and the warning above is the whole protection. -
uis pull --checkbeforeuis pull. The version advertised as available is not always the version that installs; trust the installed version, not the advertised one.
Step 1 — install
uis template install atlas
Expect roughly ok=31 changed=4 failed=0, three pods in the dagster namespace, two in postgrest,
one database, one role, two secrets and one IngressRoute.
The database is named atlas. (An install made without --database gets atlas_db instead —
if you are looking at an older install, that is why the name differs.)
Step 2 — expect an empty API
Immediately after install the API answers and serves zero endpoints. meta_sources,
meta_endpoints and meta_dimensions return 404.
That is correct, not a failed install. The schema and grants exist from install; the data arrives on the first pipeline run, which you start next.
Step 3 — load the data
Six jobs, in this order:
annual_sources_refresh → klass_refresh → seed_sources_refresh
→ brreg_bootstrap → brreg_change_feed → transform_and_publish
Both were wrong, and the omissions are the expensive half.
brreg_bootstrap and brreg_change_feed were missing from the list. brreg_bootstrap has
no schedule and no automation condition, on purpose — re-running a 1.17M-record bulk load
against a populated database is the one genuinely destructive operation in this pipeline, so nothing
self-triggers it. It runs exactly once, here. Skipping it leaves the organisation register empty
with nothing saying why, and because nothing will ever launch it for you, an install that missed it
stays wrong until somebody notices.
~11 minutes / 2.9M rows was a measurement of the first FOUR jobs (urb-agents #507) taken before
brreg_bootstrap existed. atlas-data/template-info.yaml has warned since that "an earlier figure
of ~11 minutes is still in circulation and understates a cold install by nearly 3x" — this page
was where it was circulating.
Takes ~30 minutes on a cold install: 1772 s wall, measured end to end by imac on a factory-reset
cluster (urb-agents #1027), loading ~4.1M rows. The Brreg bulk load is 501 s of that. Individual
jobs vary far more than the total suggests — annual_sources_refresh alone is 474 s and
transform_and_publish 484 s — so a run that looks stalled at 8 minutes may be normal. To tell
slow from stuck, run uis template check atlas rather than waiting: it reports what has been pulled
against what has been applied.
Why this order
seed_sources_refresh contains raw/_migrations and runs third. The migrations are idempotent,
so applying them after two source jobs have already written is safe — but it is not what anyone
would design, and reordering has not been tested (urb-agents #627).
brreg_change_feed runs immediately after brreg_bootstrap because the bootstrap seeds the
feed's watermark from the snapshot's own date. The feed has nothing to start from until it has run,
and it refuses to start without a watermark rather than silently walking history.
Running a job
uis dagster run annual_sources_refresh
One at a time, waiting for each to succeed before starting the next.
transform_checks takes 6–15 minutes to start. It returns exit 0 and then sits NOT_STARTED
before running — imac measured 364 s and 885 s (urb-agents #1160). Slow, not broken: re-launching
because it looks hung gives you duplicate runs. It is not one of the six above, and it also runs in
the sensor chain once automation is on. Why the start is slow is platform-side (urb-agents #1147).
Fallback: the raw mutation, if your UIS predates uis dagster run
mutation {
launchPipelineExecution(executionParams: {
selector: {
repositoryLocationName: "atlas-data",
repositoryName: "__repository__",
jobName: "annual_sources_refresh"
},
mode: "default"
}) {
__typename
... on LaunchRunSuccess { run { runId } }
... on PythonError { message }
}
}
Change jobName for each of the six. Poll runOrError(runId:) until the status is SUCCESS before
starting the next.
The Dagster UI can do the same through a browser, but it is internal-only with no authentication, so reaching it is a decision about your own cluster rather than something this guide can prescribe.
Step 4 — go live
Enabling the schedules is what makes Atlas keep itself current. Until you do, it holds whatever you loaded in step 3.
Every automation condition is on_cron, which means next fire, not catch-up. Enabling them on a
Thursday means the first automatic raw refresh is Sunday 02:00. That is why step 3 exists.
It said uis dagster automation cannot set state and sent you to raw startSchedule /
startSensor GraphQL. That is wrong. imac used uis dagster automation --start and --stop
on UIS 1.6.106 and they set it correctly (urb-agents #1149). The sentence was also copied into
atlas-status.py, which printed it to operators, and into a unit test that asserted the flag must
never be mentioned — so a claim about somebody else's CLI became unfalsifiable from inside this
repo. All three are corrected together.
Use uis dagster automation --start (and --stop). Verified on UIS 1.6.106; if you are on an
older UIS and it does not work, the raw GraphQL below is the fallback rather than the instruction.
Fallback: the raw mutations
Note the return types differ: startSchedule returns ScheduleStateResult, startSensor returns
Sensor. Using the wrong one gives a bare HTTP 400.
Once enabled, Atlas polls on this cadence (Europe/Oslo):
| when | what |
|---|---|
| Sunday 02:00 | ~37 annual public-sector sources — SSB, FHI, Bufdir |
| 1st of month, 01:00 | SSB Klass classifications |
| Daily 05:00 | dbt transform and publish — no external calls |
Step 5 — verify
-- structure
raw tables 47 · marts tables 64 · api_v1 views 13
-- the assertion to lead with
select count(*) from raw.brreg_enheter; -- 122
If brreg_enheter is 0, seed_sources_refresh did not run. 122 is a fixed reference list, so it
is the one number worth remembering — everything else depends on upstream volumes.
For the rest, assert properties rather than remembered totals:
Every raw source that declares
loaded_athas loaded, exceptredcross_branchesandredcross_branch_activities, which are parked pending a credential.
API checks:
GET / -> 200, 14 OpenAPI paths (13 views + `/`)
GET /meta_sources -> 200
GET /meta_endpoints -> 200
GET /meta_dimensions -> 200
And:
uis dagster verify # A/B/C PASS, code location LOADED
uis dagster automation --expect running # verify alone passes whether or not schedules are on
Everything above checks that the data is there. Atlas also defines 679 asset checks — 675
from dbt, 4 on the api_v1 surface — and on a fresh install none of them has run.
They are reached only through a chain of three instigators that all ship stopped:
transform_daily → the sensor to api_v1_checks → the sensor to transform_checks. So the suite
does not first run at 05:00 tomorrow. It first runs after the next transform_daily following
go-live, and if you never enable automation it never runs at all.
Four of them you can run now:
uis dagster run api_v1_checks # launches in ~0.5 s, runs in ~27 s
Those four cover the public API — row counts against the marts they wrap, column COMMENTs, and the anonymous role's grants — which is the part worth checking before anyone reads the data.
The other 681 can be launched too — budget 6–15 minutes. uis dagster run transform_checks
returns exit 0; imac measured two launches at 364 s and 885 s to start, the first landing 681
checks SUCCEEDED (urb-agents #1160). The run sits NOT_STARTED for that whole time and then runs.
It said the call "does not return within 300 s, exceeds the client's 60 s budget, and leaves an unsubmitted run behind — do not retry it". True when written (#1064); false on this build. The 60 s budget went in UIS 1.6.104 and the deadline is now 900 s.
It is slow, not broken. Re-launching because it looks hung is how you get duplicate runs — which is the actual hazard, and the opposite of the one the old warning described.
:::
See your data
The install is not finished until you have read a row out of it. Every endpoint below is live once step 3 has run:
/coverage_gap_barnefattigdom /indicator_latest_values /indicator_missing_kommuner
/indicator_summary /kommune_local_chapters /ngo_index
/ngo_overview /distrikt_summary /activity_catalog
/bufdir_indicator_alias /meta_sources /meta_endpoints
/meta_dimensions
GET /coverage_gap_barnefattigdom?limit=2
{"kommune_nr":"0301","kommune_name":"Oslo","fylke_name":"Oslo","year":2024,"value_pct":13,"personer":130343}
{"kommune_nr":"1101","kommune_name":"Eigersund","fylke_name":"Rogaland","year":2024,"value_pct":8.6,"personer":3270}
That is what you built — child-poverty coverage by kommune, joined to the NGOs that respond to it.
/indicator_missing_kommuner is worth opening early: it exposes the same discrepancy as the 17
warnings below, which is a friendlier introduction to them than a red test run.
Getting your machine to reach it
The route matches on hostname, so api-atlas.<your-domain> must resolve to the ingress. On a real
cluster with real DNS that is already true.
How .localhost resolves depends on your Kubernetes distribution and your OS, and this guide will not
guess for you. If the hostname does not resolve, port-forward instead — it works regardless of
DNS:
kubectl port-forward -n postgrest svc/<postgrest-service> 8080:80
curl -H 'Host: api-atlas.localhost' 'http://localhost:8080/coverage_gap_barnefattigdom?limit=2'
The Host header is not optional: Traefik routes on it, so a request without it gets a 404 from
Traefik rather than an answer from Atlas.
Things that look like failure and are not
This section exists because every item in it has produced a wrong conclusion by someone experienced.
Every run you just launched alerts as failed, fifteen minutes later
Fifteen minutes after your data load succeeds, monitoring may raise PodNotReady for every job you
ran. On a production install today that was six alerts, one per job, all six having SUCCEEDED.
A completed Kubernetes Job pod is 0/1 and reports ready=false permanently — that is what a
finished pod looks like. The alert rule matches any not-ready pod, and its own description says "Too
long for a deployment": the annotation names the intended scope and the expression does not enforce
it.
So: a PodNotReady alert naming a dagster-run-* pod is expected and is not a failure. Check the
run's status in Dagster, not the pod's readiness.
On that install the six false alerts were sitting beside one real alert that had been firing for four days — a dead metrics exporter — and the real one was invisible among them. A rule that cannot stop firing does not add noise; it spends the credibility of the channel. Fixing the rule is UIS's; knowing the alerts are false is yours.
transform_checks looks hung. It is slow.
The launch call may not return for over 45 seconds, and the run sits at NOT_STARTED for more
than a minute before succeeding in about 105 s.
All 647 dbt checks execute inside one op, so the two probes anyone reaches for — reading the definition and building the execution plan — both return "1 op, 1 step" in a fraction of a second and cannot see the weight. Two experienced testers concluded "hung" on different clusters, sixteen days apart. Both were wrong.
Judge it by polling the run, never by whether the launch call returns. Anything that alerts on it must poll the run, or it will report a false failure every night.
17 dbt checks do not pass, on a correct install
transform_checks succeeds while reporting 17 non-passing checks out of 647. All 17 are
relationships tests from indicator models to dim_kommune / dim_fylke.
They are configured severity: warn deliberately. Historical SSB series carry kommune and fylke
codes retired in the 2020 reorganisation; the dimensions carry current Klass codes. The model schema
says so in place: "Historical fylker (pre-2020 01-20 numbering) may appear — warn."
They surface as non-passing asset checks in Dagster because a dbt warning is not a pass. A run that succeeds with exactly these 17 is a correct install.
This is measured, not assumed. Two installs on different topologies, different hardware and different UIS versions both report 630 succeeded · 17 failed · 647 total, and the sets were compared rather than the counts — the same 12 tables, the same fylke-only and kommune-only split. Two lists that both totalled 17 and named different checks would have looked like agreement and meant the opposite.
⚠️ Documented is not the same as acceptable. The semantic question underneath — whether Atlas
covers pseudo-regions or merely represents them — is open. The underlying question — whether Atlas covers
pseudo-regions or merely represents them — is open and tracked in
INVESTIGATE-ssb-pseudo-regions.
meta_endpoints has 91 rows against 13 views
Not 78 phantom endpoints. The ratio has been 7 × views on every install measured. Watch the ratio, not the number — a changed ratio is the signal.
Row estimates read zero on a populated database
pg_stat_user_tables.n_live_tup is a planner estimate and can read 0 for every table while the
schema holds hundreds of megabytes. Pair every estimate with one count(*).
kubectl get ingress shows nothing
The route is a Traefik IngressRoute CRD. kubectl get ingress returns an empty result rather than
an error. Use kubectl get ingressroute -n postgrest.
A path-prefixed URL returns a Traefik 404
The route matches on hostname, not a path prefix. api-atlas.<your-domain> must resolve to the
ingress. Reaching it from your own machine is a DNS or hosts-file question about your setup, not
something the install can do for you.
The API 404s even though the view exists, with rows and grants
🔴 This is the one that will waste your evening, and it is not a false failure — it is a real outage with a one-line fix.
You check the database and everything is right: the view is there, it returns rows, atlas_web_anon
has SELECT. And GET /brreg_enhet still returns 404 on every replica.
PostgREST answers from a cached schema. It does not read the catalogue per request. If that cache was last loaded at a moment when the view did not exist, it keeps serving that absence no matter what you do to the database.
recreate the view, no reload 404 — view exists, 1.17M rows, grants correct
issue the reload 200 — instant
The fix: materialise the api_v1 asset, or run transform_and_publish. That asset is what issues
NOTIFY pgrst, 'reload schema'. It takes seconds.
⚠️ When this bites. A transform run that fails partway leaves some api_v1 views missing;
nothing restores them until a run succeeds. If PostgREST reloads while they are missing — a pod
restart is enough — the absence is cached. Repairing the database by hand then fixes everything
except the endpoint, which is exactly the state that makes it confusing.
🔵 It is not a PostgREST fault. It is serving the last schema it was told about, correctly.
api_v1 views briefly lose their descriptions
After a failed run, the views that were recreated carry no COMMENT, so PostgREST's OpenAPI
descriptions come back empty — the data is right and the published documentation is gone. A
successful transform_and_publish restores them, because the generated apply is what sets them.
Bounded, not permanent: it lasts as long as the failure does. Worth knowing so an empty API doc is not mistaken for a schema problem.
Removing and reinstalling
uis template remove atlas works for installs made by uis template install — it reads the
record that install writes. An application put in place before the template path existed has no such
record and remove exits 1.
Per-app names derive from that record. A hand-written one that disagrees with reality can point a removal at the wrong live application.
No UIS command drops a database, at any version. --purge drops per-app roles and secrets; the
atlas and dagster databases survive every removal, so a reinstall lands on the existing schema.
Dropping is manual DDL and the order matters:
DROP DATABASE atlas; -- or atlas_db on an older install
DROP ROLE atlas; -- blocked until the database is gone
DROP DATABASE dagster;
Leave atlas_web_anon and atlas_authenticator — configure postgrest re-credentials them. Leave
the dagster role.
Use uis undeploy dagster rather than helm uninstall; the playbook waits for pods to terminate and
tells you what it preserves.
If you delete the dagster namespace, run uis secrets generate && uis secrets apply before
reinstalling — the Dagster setup reads a generated password from a secret in that namespace, and it
fails loudly with that instruction if it is missing. You are not expected to know that password.
Known gaps
| gap | status |
|---|---|
uis dagster automation --start | CLOSED — it shipped and this table did not notice. Verified working on UIS 1.6.106 (imac, urb-agents #1149). Listed as a gap long after it stopped being one, which is why atlas-status.py was still printing GraphQL instructions. |
uis dagster run <job> | Works. api_v1_checks ~27 s; transform_checks 6–15 min to start, exit 0 (imac, urb-agents #1160). Why the start is slow is platform-side, with tor-agent (urb-agents #1147). |
| Job order is documented, not enforced | see step 3 |
transform_checks start latency | tracked in INVESTIGATE-transform-job-decomposition |