Skip to content

Governance & Tracking-Gap Register

Classification: CONFIDENTIAL — Internal Use Only Document: governance/gap-register.md · v1.32 · 2026-09-09 · GPUS-IT Owner: unassigned (see Owner field below) · Review cadence: not yet set


Purpose

The register of governance and tracking gaps — places where the estate's own records, controls or checks do not do what they are documented to do. It is one of three registers opened 2026-08-13:

Register ID prefix Lives in
Security findings VLN- security/vuln/tracker.md
Governance / tracking gaps GOV- this document
Program work items PRG- / HB- priorities/gpus-it-priorities.md

GOV- was chosen over GAP- because GAP-1 is already in use in the forms workstream with an unrelated meaning (GAP-1 renderer fix). There are no existing GOV- identifiers anywhere under docs/.


Schema

Every row carries: id · title · register · status · priority · owner · evidence · acceptance test · blocks · blocked_by · date_raised · date_verified.

Status vocabulary: open · in progress · blocked · done.

done means verified end-to-end via the real production path against live state. Authored, present, loaded or validated is not done.

Two rows are done: GOV-009 (verified live 2026-08-18, recorded 2026-08-21) and GOV-021 (verified live 2026-08-24, recorded the same day). Every other row here remains open or in progress.

evidence records what was read, where, and on what date — a command output or a file path, never "per notes". acceptance test states, in advance, the observable condition that closes the row.

A row missing evidence or an acceptance test is marked ⚠ INCOMPLETE on that field and is not filled with a plausible guess. Four rows below are incomplete. They are listed as such deliberately; see Incomplete rows at the foot of this document.

Owner field. No owner was supplied for any row. Every row reads unassigned. This is a gap in its own right and is not resolved by defaulting to the document owner.

Priority field. No priority was supplied for any row. Every row reads not assessed. None was inferred.


Index

ID Title Status Evidence Acceptance test Blocked by
GOV-001 Rewritten 2026-08-17 — declared source of truth covers ~a fifth of the estate open ✅ confirmed first-hand
GOV-002 validate_portal_presence() checks two portals against a written standard of three open ✅ confirmed first-hand INCOMPLETE
GOV-003 cloud-services-render sentinel satisfies presence unconditionally, with no expiry open ✅ confirmed first-hand INCOMPLETE
GOV-004 No declared-vs-actual reconciliation control and no GL5 liveness monitoring open INCOMPLETE
GOV-005 ADC and gcloud CLI identities diverge; Terraform reads ADC open ⚠ partial
GOV-006 VLN-011 had a published finding page and nav entry but no tracker row in progress INCOMPLETE GOV-009
GOV-007 priorities/session-log.md stale since 2026-04-24 in progress ✅ confirmed first-hand INCOMPLETE GOV-009
GOV-008 Pre-existing Done rows predate the status rule and are un-re-adjudicated open
GOV-009 Register changes cannot be verified as rendered — render confirmed live done 2026-08-18 ✅ confirmed first-hand ✅ met
GOV-010 inventory.yaml covers roughly a fifth of the estate — 82 declared vs 481 live open
GOV-011 PuppetDB enforcement gap — DERIVED is not enforced open INCOMPLETE
GOV-012 Five nodes applying against a dead Puppet master open INCOMPLETE
GOV-013 Control repo does not reproduce the running configuration open INCOMPLETE
GOV-014 Three repos on deprecated CSR, one unassessed at 39 GB open INCOMPLETE
GOV-015 Twelve power devices never enumerated by any Stage 1 surface open INCOMPLETE
GOV-016 Hostname is not a reliable join key in this estate open INCOMPLETE
GOV-017 Property disagreements between inventory.yaml and live state open INCOMPLETE
GOV-018 Orphan count superseded twice — record 50 with the method stated open
GOV-019 Snapshot configuration unknown on all three Synology units (403) open INCOMPLETE
GOV-020 Two tools share one working tree — each push publishes the other's unreviewed commits open ✅ confirmed first-hand
GOV-021 The travel form told every submitter a workflow it does not perform — corrected and verified live done 2026-08-24 ✅ confirmed first-hand ✅ met
GOV-027 required is enforced as key-presence, not as a value — every form open ⚠ read-confirmed, not exercised against the live endpoint
GOV-028 gpus-reports deploy drift guard — deployed at 59a253c, both acceptance commands pass done 2026-08-27 ✅ confirmed first-hand on the host ✅ met, verified live
GOV-029 /opt/gpus-reports is cloudadmin-owned, so its root-owned modules can be replaced without sudo open ✅ confirmed first-hand, probed on the host ✅ runnable, currently failing
GOV-030 Fact D of the ASVS scope rests on a revocation path the estate does not have open ✅ confirmed first-hand, read at 2939812 ✅ runnable, currently failing
GOV-031 A stall at the final approval step has no chaser, and the one open case is the coordinator's own trip open ✅ confirmed first-hand, live DB read ✅ runnable, currently failing
GOV-032 Conditional display — the third instance arrived 2026-09-02; asked for three times, compromised around three times, DECLINED 2026-09-01. Also now carries Shereyll's second capability gap: a live client-side cost total open — declined, recorded ✅ confirmed first-hand; ⚠ the two requests are not preserved verbatim anywhere ✅ runnable, currently passing
GOV-033 The September executive monthly did not send, and no signal exists that would have said so open ✅ confirmed first-hand on the host ✅ three, runnable, all failing
GOV-034 MAILTO=root reaches nobody on all seven hosts — six have no MTA at all, MAPLE has 439 KB unread open ✅ confirmed first-hand, probed on all seven ✅ two, runnable, both failing
GOV-035 /metrics is public and unauthenticated, and a container boot can hold it to the 60 s platform timeout open — rescoped 2026-09-01, the incident was a memory outage ✅ confirmed first-hand, Cloud Logging + live config ✅ three, runnable, all failing
GOV-036 Cloud Run cost defaults were set for a portal with no traffic; the estate now runs live travel approvals on them open ✅ confirmed first-hand, all 10 services surveyed live ✅ runnable review, currently failing on all three for all ten
GOV-037 The travel reports take a revised trip's values from the submission that was rejected — R1 and R3; R5 immune open ✅ confirmed first-hand, full key diff run live ✅ runnable, fails on today's code in both modules
GOV-038 10.8.0.0/28 became a default allow-list by accretion — one origin now fronts 8 services across both cloud hosts, and nobody has asked which are load-bearing open ✅ confirmed first-hand
GOV-039 Every WDC syslog event is indexed 4 hours in the past — the collector parses EDT timestamps as UTC, and the grok/date/mutate filters are dead code open ✅ confirmed first-hand
GOV-040 BIND on SKY and RAIN is not chrootednamed-chroot.service is inactive and disabled on both, while Puppet manages that unit open ✅ confirmed first-hand

GOV-001 — The declared source of truth covers roughly a fifth of the estate

REWRITTEN 2026-08-17 — the finding is scope, not competing claims

This row was rewritten per B-15 (Stage 2 A1 census). The original 2026-08-13 wording is retained in full below as a dated superseded note, not deleted: it is accurate, and anything citing GOV-001 before 2026-08-17 means that framing.

Field Value
id GOV-001
title The declared source of truth covers roughly a fifth of the estate — 82 declared entries against 481 deduped live entities
register Governance / tracking gaps
status open
priority not assessed — none supplied
owner R. Chhetry
evidence Stage 2 A1 census, 2026-08-17. inventory.yaml holds 82 entries against 481 deduped live entities. gpusa has zero entries against 174 live entities. gpus-it has 9 against 57. last_updated: 2026-06-26. Prior first-hand evidence from 2026-08-13 is preserved in the superseded note below. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage2-A1-inventory-census.txt. Confirmed: TOTAL ENTRIES IN inventory.yaml: 82, version=1 last_updated=2026-06-26, and the bucket breakdown showing gpusa 0 against wdc 24, gpus-it 9, gpus-infra 7, meraki/network 33, gcp-unattributed 9. ⚠ Discrepancy — figure did not reproduce. The artifact's own header reads Captured 2026-08-13, not 2026-08-17. This row and GOV-010 both record the census as 2026-08-17 because that is what the backlog stated. The evidence date is wrong by four days; the figures are not. Corrected here rather than in the backlog, which is source text. The 481 figure is not in this file — it comes from the A2 dedupe pass, raw/stage2-A2-dedupe.txt.
acceptance test inventory.yaml contains an entry for every entity in enumerated.yaml, or a dated exception with an expiry for each omission. (Widened from "one claim retired in writing", which addressed only the secondary framing.)
blocks
blocked_by
date_raised 2026-08-13 (rewritten 2026-08-17)
date_verified

Why the reframing matters. The competing-claims finding was a documentation inconsistency. This one explains a control failure: it is why the coverage gate has been passing. The gate validates only what is declared, and nothing declares the largest location. A gate over 82 of 481 entities cannot fail on the 399 it cannot see. Read with GOV-003 — the cloud-services-render sentinel — the gate is narrow and its one enforced category is unconditionally satisfied.

GOV-010 carries the same census as the row raised from B-15 in the 2026-08-17 load. The two are deliberately not merged: GOV-001 is the identifier already cited elsewhere and keeps its history, GOV-010 is where the new load filed it. Treat GOV-001 as canonical and GOV-010 as its entry in the 2026-08-17 batch.

SUPERSEDED 2026-08-17 — original GOV-001 wording, raised 2026-08-13

Retained verbatim. Accurate but secondary, per B-15.

title — Two artifacts each claim to be the single source of truth: inventory.yaml vs the host registry CSVs

evidence — Supplied as "inventory.yaml vs wdchostregistry.csv, per information-asset-registry.md". Confirmed first-hand 2026-08-13, and the citation corrected: there is no information-asset-registry.md in the repo. The competing claim is at mkdocs-portal/docs/hostregistry/index.md:13"Each registry is the single source of truth". Against it, three documents assert the opposite: docs/governance/inventory-schema.md:10, docs/governance/component-coverage-standard.md:58, and docs/governance/adding-new-infrastructure-quickstart.md:9, each stating "inventory.yaml at the repository root is the single source of truth".

acceptance test — One claim retired in writing.

The conflict is 3 documents to 1, which suggests which way it resolves — but the acceptance test is a written retirement, not a majority count, so the row stays open until that is done.

This sub-finding is not closed by the rewrite. It remains true and its original acceptance test is unmet.


GOV-002 — validate_portal_presence() checks two portals against a written standard of three

Field Value
id GOV-002
title validate_portal_presence() checks PORTAL_DIRS = ("status-site", "soc-site") — two portals against a written standard of three
register Governance / tracking gaps
status open
priority not assessed — none supplied
owner unassigned
evidence Confirmed first-hand 2026-08-13. scripts/check-component-coverage.py:564PORTAL_DIRS = ("status-site", "soc-site"); consumed by validate_portal_presence() at :584, and iterated at :596, :603, :644; invoked from :744. The written standard is docs/governance/component-coverage-standard.md:33-40, which requires presence across three surfaces and states that anything missing an "entry, status-site card, or soc-site tile is incomplete and must not" be treated as covered. The check enforces two of the three.
acceptance test INCOMPLETE — none supplied. No closing condition was stated for this row, and none is invented here. Settling it requires a decision the record does not contain: whether the standard drops to two surfaces or the check rises to three.
blocks
blocked_by
date_raised 2026-08-13
date_verified

GOV-003 — cloud-services-render sentinel satisfies presence unconditionally, with no expiry

Field Value
id GOV-003
title The cloud-services-render sentinel satisfies portal presence for every cloud_services entity unconditionally and without expiry — broader than .coverage-exceptions.yaml permits
register Governance / tracking gaps
status open
priority not assessed — none supplied
owner unassigned
evidence Confirmed first-hand 2026-08-13. scripts/check-component-coverage.py:557PORTAL_RENDER_SENTINEL = "cloud-services-render", with the comment at :551 describing a portal as rendering an entity when its source merely names the sentinel. Against that, .coverage-exceptions.yaml (repo root) states in its header that "Every entry MUST include an expires: date" and that "Indefinite exceptions are not allowed by design." Its exceptions: list is currently empty (exceptions: []), so the sentinel is the only bypass in effect — and it is exactly the indefinite, unexpiring kind the exceptions file forbids.
acceptance test INCOMPLETE — none supplied. No closing condition was stated, and none is invented here.
blocks
blocked_by
date_raised 2026-08-13
date_verified

The design intent of .coverage-exceptions.yaml is that no opt-out outlives its expiry. The sentinel is an opt-out that never expires and was never entered in that file, so it is not visible as an exception at all.


GOV-004 — No declared-vs-actual reconciliation control and no GL5 liveness monitoring

Field Value
id GOV-004
title No reconciliation control (declared vs actual) and no GL5 liveness monitoring — the phoebe mechanism, still open
register Governance / tracking gaps
status open
priority not assessed — none supplied
owner unassigned
evidence INCOMPLETE. What was supplied — "This is the phoebe mechanism, still open" — is a characterisation of the gap, not a record of what was read, where, or on what date. No artifact, command output or file path was given, and the absence of a control cannot be evidenced by asserting it. Not filled with a guess. To complete: the check or scheduler listing that demonstrates no such job exists, dated.
acceptance test A scheduled read-only reconciliation job whose output diffs against inventory.yaml.
blocks
blocked_by
date_raised 2026-08-13
date_verified

Related but not a dependency: HB-004 (phoebe's real consumers) and PRG-008 (phoenix into inventory.yaml and under a liveness check) are downstream of the same missing control. No blocks relation is asserted between them because none was stated.

Design constraint attached 2026-08-17 — GOV-016 (B-21)

Any reconciliation control built to close this row must join on serial/MAC for devices and IP for DNS. Hostname is not a reliable join key in this estate. This was learned in both directions at Stage 2:

  • False absence — 21 Meraki devices were reported DECLARED_NOT_FOUND because inventory uses descriptive names (wdc_stack_0, mdec_wap_garage) while Meraki reports operational ones (WDC-STACK-0-NOPOE, mdec garage). Re-matching on serial and MAC — both already present in inventory.yaml — resolved all 21, which would otherwise have inflated DECLARED_NOT_FOUND from 23 to 44.
  • False presenceforms.us.gl3 matched gpus_forms_frontend, whose FQDN is forms.greenpeace.us, conflating the legacy portal with its replacement and hiding a retirement candidate behind a confident match.

A control that joins on hostname would have reported forms.us.gl3 as reconciled. See GOV-016 for the full row, and PRG-025 for the enumeration surfaces such a control must additionally cover.


GOV-005 — ADC and gcloud CLI identities diverge; Terraform reads ADC

Field Value
id GOV-005
title Application Default Credentials (.org, dated 2026-07-22) and the gcloud CLI identity (.us) diverge; Terraform reads ADC
register Governance / tracking gaps
status open
priority not assessed — none supplied
owner unassigned
evidence Partial. Supplied: ADC resolves to the .org identity and is dated 2026-07-22; the gcloud CLI resolves to .us; Terraform reads ADC. The identity strings and the ADC date are specific, but no read date or command output was given for the observation itself. To complete: the ADC file path and gcloud auth list output, dated. Related first-hand context: VLN-016 records the same person holding three identity strings with inverted access. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): access-map.md — this upgrades the row from partial to sourced. It states directly: "gcloud auth application-default login was not run. CLI and ADC now genuinely diverge; expected for this stage," and records that accounts were selected per-call with --account while core/account and ADC were left unmodified. Raw probe output at raw/access-probe-both-accounts-20260813-v2.txt, dated 2026-08-13.
acceptance test Reconciled and verified before any Terraform run.
blocks
blocked_by
date_raised 2026-08-13
date_verified

The acceptance test carries an implicit ordering constraint rather than a date: any Terraform run before reconciliation executes as an identity nobody chose.


GOV-006 — VLN-011 had a published finding page and nav entry but no tracker row

Field Value
id GOV-006
title VLN-011 had a published finding page and a nav entry but no row in the vulnerability tracker — retained as a closed record of the defect class
register Governance / tracking gaps
status in progress — the missing row is committed but not verified rendered. Not done: a commit is not the production path.
priority not assessed — none supplied
owner unassigned
evidence The defect was recorded in security/vuln/tracker.md v1.6 (2026-08-10) as the second of two register-hygiene defects. Remedied in commit 7ff38ae, 2026-08-13, which added the VLN-011 row linking the pre-existing finding page. Verified structurally the same day: every VLN- id in the tracker now appears exactly once (12 ids, no duplicates).
acceptance test INCOMPLETE — none supplied. What was given — "retain as a closed record of the defect class" — is a disposition instruction, not an observable close condition. The nearest defensible test would be the VLN-011 row confirmed rendered live, but that is GOV-009's test, not this row's, and it is not imported here as a substitute.
blocks
blocked_by GOV-009 — while the render cannot be verified, this row cannot advance past in progress regardless of what its acceptance test turns out to be.
date_raised 2026-08-13
date_verified

The defect class this row preserves: an artifact can be published, navigable and cross-referenced while the register that is supposed to account for it has no entry at all. Nothing in the toolchain caught it; it was found by reading.


GOV-007 — priorities/session-log.md stale since 2026-04-24

Field Value
id GOV-007
title priorities/session-log.md stale since 2026-04-24
register Governance / tracking gaps
status in progress — a 2026-08-13 entry was added this session. Not done: the entry is committed, not verified rendered, and one entry does not close a ~16-week gap.
priority not assessed — none supplied
owner unassigned
evidence Confirmed first-hand 2026-08-13. mkdocs-portal/docs/priorities/session-log.md read on 2026-08-13: 48 lines, single entry headed "## 2026-04-24 — Phase 2 forms backend LIVE", file mtime 2026-04-24. No entry existed for any session between 2026-04-24 and 2026-08-13 — a span over which gpus-it-priorities.md alone advanced from v1.1 to v1.18, i.e. 17 documented version bumps with no corresponding session record.
acceptance test INCOMPLETE — none supplied. No closing condition was stated. It is genuinely ambiguous whether this closes on "the log is current from now on" (a cadence, needing a defined interval) or on "the 16-week gap is backfilled" (a body of work). Those are different rows with different costs, and choosing between them is not a gap-filling exercise. Not guessed.
blocks
blocked_by GOV-009
date_raised 2026-08-13
date_verified

GOV-008 — Pre-existing Done rows predate the status rule and are un-re-adjudicated

Field Value
id GOV-008
title Pre-existing Done rows in the priorities tracker predate the status rule and have not been re-adjudicated against it
register Governance / tracking gaps
status open
priority not assessed — none supplied
owner unassigned
evidence Flag list produced 2026-08-13 against priorities/gpus-it-priorities.md v1.18: 13 rows flagged — 2 Tier A (Done contradicted by later evidence in this repo) and 4 Tier B (Done asserted while the row's own stated gate is unmet), plus 7 Tier C (no verification evidence recorded). Committed as priorities/flag-list-unverified-done-2026-08-13.md. Flag list only — no row in gpus-it-priorities.md was changed.
acceptance test A dedicated re-adjudication pass, with evidence gathered per row.
blocks
blocked_by
date_raised 2026-08-13
date_verified

The two Tier A rows are the load-bearing ones: the SKY portal-backup row is self-contradicted by its own warning text, and the forms go-live row asserts a delivery claim that VLN-011 disproves. That second one is also carried separately as VLN-022, because the false claim is currently published.


GOV-009 — Register changes cannot be verified as rendered — CLOSED AND VERIFIED 2026-08-18

Field Value
id GOV-009
title Register changes could not be verified as rendered; structural validation is not a render — resolved, render confirmed live
register Governance / tracking gaps
status done — verified live 2026-08-18
priority not assessed — none supplied
owner R. Chhetry
evidence Confirmed first-hand 2026-08-13. mkdocs build --strict run in the Cowork workspace against mkdocs-portal/ aborted before building: ERROR - Config value: 'theme'. Error: Unrecognised theme name: 'material'. The available installed themes are: mkdocs, readthedocsAborted with 1 Configuration Errors!. mkdocs itself is present at /usr/bin/mkdocs; mkdocs-material is not installed and the workspace has no network access to install it. What was run instead, same date: table pipe-count consistency (no mismatches), internal link resolution (5 of 5 targets exist on disk), admonition block indentation, and VLN- id uniqueness (12 ids, no duplicates). None of that is a render.
acceptance test Changes confirmed live on infra.greenpeace.us after push. MET 2026-08-18 — see the verification below. One clause amended; see the amendment note.
blocks GOV-006, GOV-007, VLN-022 — and, in practice, the verification of every row written on 2026-08-13. All three are now unblocked and await re-adjudication; none is re-statused here.
date_raised 2026-08-13
date_verified 2026-08-18

This row was the reason nothing dated 2026-08-13 was done anywhere in the three registers. It is now closed, and the constraint it imposed is lifted for rows verified against the live render. Rows still carrying unmet acceptance tests remain open on their own merits, which is the great majority of them.

The two anchor links added to tracker.md on 2026-08-13 — VLN-004VLN-013, and the VLN-009VLN-012 redirect — have both now been checked in the built site. The first was broken and fixed (5d177fc); the second resolves to its section heading and is addressed in the amendment note above.

CLOSED 2026-08-18 — 73 register rows confirmed rendering live

Verified on infra.greenpeace.us after push, at tracker v1.11, gap-register v1.2, priorities v1.21:

  • All six pages return HTTP 200.
  • Nav entries present for each.
  • All 51 internal anchors resolve, after commit 5d177fc and build f02200d4.

This is the production path against live state, so it satisfies the status rule as written. It is the first row in any of the three registers to reach done.

5d177fc is why this row existed. The VLN-004VLN-013 anchor written on 2026-08-13 was broken — an em-dash slugifies to a single hyphen, not two — and no amount of structural validation could have caught it, because the target file existed and only the slug was wrong. The render check found it; the pre-render checks could not.

Acceptance-test amendment — the VLN-009VLN-012 redirect

One clause of the original test asked for a resolving anchor on the VLN-009VLN-012 redirect in tracker.md. That clause is amended, because it asserted a capability the structure never had — not because it failed.

VLN-012 is a table row, and table rows receive no heading id. There is therefore no row-level anchor target, and none can exist without restructuring the Remediated-vulnerabilities table into per-finding headings. The link as written points at the section heading (#remediated-vulnerabilities), which does resolve and is among the 51 confirmed — and the surrounding prose names VLN-012 explicitly, so a reader arriving from a stale "VLN-009 = chronyd" reference lands in the right place and can see it.

The prose redirect serves the intent. Amended rather than failed, and recorded here so the amendment is visible rather than silent — a test that demanded the impossible would otherwise have held this row open forever.

Three rows are unblocked by this closure and are NOT re-statused here

GOV-006, GOV-007 and VLN-022 each carried blocked_by: GOV-009. That blocker is gone. None of them is advanced by this commit — each has its own acceptance test to meet, and two of the three (GOV-006, GOV-007) are marked INCOMPLETE on that field, so what would close them is still an open question. They need a deliberate re-adjudication pass, not an automatic promotion.



GOV-010 — inventory.yaml covers roughly a fifth of the estate

Field Value
id GOV-010
title The declared source of truth covers roughly a fifth of the estate — 82 declared entries against 481 deduped live entities
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 2 A1 census, 2026-08-17. inventory.yaml holds 82 entries against 481 deduped live entities. gpusa has zero entries against 174 live entities. gpus-it has 9 against 57. last_updated: 2026-06-26. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage2-A1-inventory-census.txt. Confirmed: 82 total entries, gpusa bucket 0.Discrepancy — figure did not reproduce. The artifact header reads Captured 2026-08-13, not the 2026-08-17 recorded here. See GOV-001 for the full note. The census also carries a THE FIVE NAMED HOSTS block independently confirming that gannet, emu, ostrich, catbird and phoebe are all absent from inventory.yaml — corroborating the first-hand check of 2026-08-13 that left HB-001HB-004 without host rows — and adds thrush as a sixth absent host, with kingfisher present.
acceptance test inventory.yaml contains an entry for every entity in enumerated.yaml, or a dated exception with an expiry for each omission.
blocks
blocked_by
date_raised 2026-08-17
date_verified

This row rewrites GOV-001. The original framing — competing single-source-of-truth claims at hostregistry/index.md:13 against three documents asserting inventory.yaml — is accurate but secondary, and is retained there as a dated superseded note rather than deleted.

The primary finding is coverage, and it explains something the 2026-08-13 build could not: why the coverage gate has been passing. It validates only what is declared, and nothing declares the largest location. A gate over 82 of 481 entities cannot fail on the 399 it cannot see. Read with GOV-003 — the cloud-services-render sentinel — the two together mean the gate is narrow and its one enforced category is unconditionally satisfied.


GOV-011 — PuppetDB enforcement gap: DERIVED is not enforced

Field Value
id GOV-011
title PuppetDB enforcement gap — 11 certnames against 39 FQDN node definitions; thirteen RUNNING VMs have never produced a retained catalog
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3, 2026-08-17. 11 certnames against 39 FQDN node definitions. Thirteen RUNNING VMs with node definitions have never produced a retained catalog — including phoenix itself. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage3-puppetdb-nodes.json. Confirmed: 11 certnames.
acceptance test INCOMPLETE — none supplied. The row states a design requirement — "this must gate disposition, not merely annotate it" — which is a constraint on how the classification is used, not an observable condition that closes the row. Not converted into one.
blocks PRG-020derived. GOV-011 gates the Stage 3 DERIVED result.
blocked_by
date_raised 2026-08-17
date_verified

Confidence and enforcement are different axes

Recorded as supplied: a DERIVED function statement on a host with no retained catalog is a hypothesis, not evidence. PRG-020 records that 24 of 25 gpusa VMs moved UNKNOWN → DERIVED at Stage 3; this row is why that movement is not a movement toward verified. See also M-07 in Method learnings.


GOV-012 — Five nodes applying against a dead Puppet master

Field Value
id GOV-012
title Five nodes declare puppetMaster => raven.cloud.us.gl3; raven has no backing host
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3, 2026-08-17. Five nodes declare puppetMaster => raven.cloud.us.gl3. raven is a DNS A record at 10.1.96.25 with no backing host. Classified ASSERTED-STALE.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Recorded as supplied: this is managed-in-name-against-nothing, a third state that is neither managed nor unmanaged. Any control that classifies hosts as managed/unmanaged will mis-sort these five, because the declaration is present and the enforcement target does not exist.


GOV-013 — The Puppet control repo does not reproduce the running configuration

Field Value
id GOV-013
title The Puppet control repo does not reproduce the running configuration — untracked but load-bearing files
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3 Task 1a, 2026-08-17. Untracked but load-bearing: production/environment.conf (defines modulepath), hiera.yaml, an emacs autosave #environment.conf#, and five vendored modules (apt, concat, inifile, stdlib, systemd).
acceptance test INCOMPLETE — none supplied. The row states a remediation method"mirror /etc/puppetlabs/code as a filesystem capture, not a git clone" — but no observable condition that closes the row. The method is recorded; a close condition is not invented from it.
blocks
blocked_by
date_raised 2026-08-17
date_verified

The emacs autosave file being load-bearing is worth keeping in the record: a clone of the repo reproduces neither the modulepath nor the vendored modules, so "we have the repo" is not "we can rebuild the master".


GOV-014 — Three repos on deprecated Cloud Source Repositories, one unassessed at 39 GB

Field Value
id GOV-014
title Three repos on phoenix across two CSR projects, all on deprecated Cloud Source Repositories
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3b, 2026-08-17. gpus-puppet and gpus-dist in gpus-it-infrastructure (scheduled for deletion), and /opt/puppet/static served as puppet:///static/ from gpusa-it-infrastructure/gpusa-puppet, branch testing, last commit 2019-08-30, 39 GB — unassessed.
acceptance test INCOMPLETE — none supplied. The row states an ordering constraint — "mirror before any teardown work touches those projects" — which is a sequencing rule, not a close condition. Not converted into one.
blocks
blocked_by
date_raised 2026-08-17
date_verified

The ordering constraint is the urgent part and is easy to lose in a backlog: one of the two projects is already scheduled for deletion, and 39 GB of the content has never been assessed. VLN-029 (key material in the control repo) depends on this mirror happening first.


GOV-015 — Twelve power devices never enumerated by any Stage 1 surface

Field Value
id GOV-015
title Twelve power devices (PDUs/UPS) were never enumerated by any Stage 1 surface
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 2 A3, 2026-08-17. Twelve power devices absent from every Stage 1 surface; six have 192.168.122.x management IPs that nothing probed. Firmware and credentials unknown. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage2-A1-inventory-census.txt. Confirmed: power_devices 12 in the by-category census, all twelve mapping to the wdc bucket. The census establishes the twelve are declared; the row's point stands unchanged — no Stage 1 surface enumerated them.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

This is an enumeration gap, not an estate fact

Recorded verbatim as supplied: do not infer state from absence. Nothing here says the devices are unmanaged or insecure — it says nothing looked. See M-06 and M-10 in Method learnings.


GOV-016 — Hostname is not a reliable join key in this estate

Field Value
id GOV-016
title Join-key constraint — hostname is not a reliable join key; any reconciliation control must join on serial/MAC for devices and IP for DNS
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 2, 2026-08-17. Learned in both directions. False absence: 21 Meraki devices were reported DECLARED_NOT_FOUND because inventory uses descriptive names (wdc_stack_0, mdec_wap_garage) while Meraki reports operational ones (WDC-STACK-0-NOPOE, mdec garage); re-matching on serial and MAC — both already present in inventory.yaml — resolved all 21, which would otherwise have inflated DECLARED_NOT_FOUND from 23 to 44. False presence: forms.us.gl3 matched gpus_forms_frontend, whose FQDN is forms.greenpeace.us — conflating the legacy portal with its replacement and hiding a retirement candidate behind it.
acceptance test INCOMPLETE — none supplied. Not invented. This row is a constraint on a control that does not yet exist (GOV-004), so its close condition is properly a property of that control's design.
blocks
blocked_by
attached to GOV-004 as a design constraint — stated in the source row.
date_raised 2026-08-17
date_verified

The false-presence direction is the more dangerous of the two and the easier to miss: a bad join did not merely lose an entity, it produced a confident match that concealed one. A reconciliation control that joins on hostname would have reported forms.us.gl3 as reconciled.


GOV-017 — Property disagreements between inventory.yaml and live state

Field Value
id GOV-017
title Property disagreements between inventory.yaml and live state — presence was never the issue, the declared attributes are wrong
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 2 Part B, 2026-08-17. See the disagreement table below. Compute instances showed no disagreement on ip, machine_type or state. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage2-B-property-disagreements.txt, the Stage 2 Part B capture this row's table is drawn from.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
justifies the widening of VLN-013's acceptance test — stated in the source row.
date_raised 2026-08-17
date_verified
Entity Field Declared Live
fire vms [] sky, rain, sun, wind
water hypervisor ESXi 6.7 ESXi 8.0.3 (build 24022510)
water vms ["ocean"] none
flower hypervisor ESXi 6.8 (no such release) ESXi 6.7.0 (build 8169922)
flower vms [] ocean, desert, river, star
fire / water / flower hardware_model "" R610 / R660xs / R630
vmstorage model "Synology" RS1619xs+
solstorage model "Synology" DS1823xs+
synstorage not declared at all SA3600, 3 volumes
SKY / RAIN / SUN / WIND project gpus-infra on-prem WDC VMs on fire

Recorded as supplied: fire declared as hosting nothing while carrying all four core WDC servers is the sharpest. Any blast-radius reasoning from that field is wrong, and that field is what a coverage gate reads. Note flower is declared as ESXi 6.8 — a release that does not exist — which no schema check caught.


GOV-018 — Orphan count superseded twice; record 50 with the method stated

Field Value
id GOV-018
title Orphan count superseded twice — 35 → 52 → 50
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3, 2026-08-17. 35 (Stage 1b raw A-record cross-reference) → 52 (Stage 2 matcher, which did not know the estate spans two GCP projects for Puppet purposes) → 50 (Stage 3, short-name method accounting for the sibling project). Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task1-arecord-crossref-20260813.txt for the first figure. Its header reads "Stage 1b Task 1 — wdc.us.gl3 A records vs enumerated.yaml (Stage 1 state). Captured 2026-08-13", recording 138 A records in zone (sky and rain identical) against 159 IPv4-bearing rows across 117 distinct IPs in enumerated.yaml. The 52 and 50 figures come from the later Stage 2 and Stage 3 passes.
acceptance test Record 50 with the method stated; both prior figures marked superseded. (Stated in the source row; recorded as given.)
blocks
blocked_by
date_raised 2026-08-17
date_verified

The count moved twice because the method changed, not because the estate did. PRG-021 records the two-project fact that corrected the middle figure. See M-08 — every count in a register needs a source.


GOV-019 — Snapshot configuration unknown on all three Synology units

Field Value
id GOV-019
title Snapshot configuration unknown on all three Synology units — error 403, the querying account lacks rights
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Synology DSM API, 2026-08-17. Error 403 on all three units — the querying account lacks rights to that API. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed verbatim: Snapshot list: FAILED err=403 (permission denied for this account). The 403 is recorded in the capture itself, so the blindness is documented rather than inferred — exactly the discipline M-02 and M-06 require.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks VLN-026derived. Snapshot configuration unknown blocks assessing the severity of the no-backup finding.
blocked_by
date_raised 2026-08-17
date_verified

Unknown, not absent

Recorded verbatim as supplied. A 403 means the enumerator is blind, not that snapshots are missing. VLN-026 is a positive finding of configured absence for backup tasks; this row is the reason its severity cannot yet be assessed. See M-02 — permission-denied is not unreachability.


GOV-020 — Two tools share one working tree; each push publishes the other's unreviewed commits

Field Value
id GOV-020
title Two tools operate on the single checkout at ~/gpus-infra-portals; a push from either publishes whatever the other has committed, unreviewed
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Confirmed first-hand 2026-08-24. Four independent signals, all from local state in this repository. (1) Committer timezone splits cleanly. August 2026 carries 76 commits: 68 at +0545, all forms/reports engineering, and 8 at +0000, all governance-register documents — 7ff38ae, 6a1821c, 02c9f87, 44275b8, ebdb99c, ebfde58, e0b2bf5, 49c4cdd. (2) .git/claude-parked-junk/ holds seven HEAD.lock.<pid>.5 files whose mtimes match the +0000 commits to the second once converted to local time — 01:06:46ebfde58, 13:29:06e0b2bf5, 00:28:1249c4cdd, and the same for 08-13 and 08-17. These are local files in this .git, so the UTC commits were made against this tree, not fetched from a clone. (3) An older .git/_stale_locks/ generation dated 2026-07-29 uses a different naming scheme (nanosecond-epoch suffixes, plus a maintenance.lock) — two distinct lock-parking implementations, so more than one tool version has done this. (4) .git/index.lock, 0 bytes, dated 2026-08-21 00:28 — the same second as 49c4cdd — blocked a git stash on 2026-08-24 with error: could not write index, with no git process running.
acceptance test One tool per working tree, or separate clones. Closed when, over 30 consecutive days, (a) every commit reaching origin/main carries a single committer timezone offset, and (b) no new lock-parking artifacts appear under .git/ — no claude-parked-junk/, no _stale_locks/, no orphaned index.lock or HEAD.lock. Both clauses are required: (a) alone would pass if two tools were merely configured to the same TZ.
blocks
blocked_by
date_raised 2026-08-24
date_verified

The mechanism is the finding — not that two tools committed, but that either one publishes the other

Git pushes a branch, not a session's commits. Whatever the other tool has left on main is an ancestor of the tip, so it goes out with the next push from either side, silently and unreviewed.

This already happened. The push at 2026-08-20 23:33:59 +0545 published three commits: 6744e6b and fcbc9b4 — the R2 work that session had just written — and e0b2bf5, a governance commit made sixteen hours earlier by the other tool. That session did not write e0b2bf5, in all likelihood never saw it, and certainly did not review it. It rode along.

Note what this defeats. Reviewing your own diff before pushing is not enough, because the commit that ships is not in your diff. git log origin/main..HEAD is the only check that would have shown it, and nothing prompts anyone to run it.

This is not a hypothetical risk — it has already produced a live finding

gpus-reports/ and soc-backend/ hold copies of the same three report scripts and have drifted 99 lines apart in the generator alone (979 vs 880). MAPLE runs the gpus-reports/ copies. That drift is tracked separately; this row is the mechanism that produced it — two writers editing one tree with neither seeing the other's working state.

The same shape recurred on 2026-08-24: 49c4cdd, a commit from the other tool marking VLN-014 and GOV-009 done, sat unpushed on main and could not be excluded from a push of unrelated work in forms-backend/ and forms-frontend/ — because it is an ancestor of the tip. Reviewing it was a deliberate act, not something the workflow forced.

Remediation — two options, one recommended

Option A — one tool per working tree (RECOMMENDED). Only one of Claude Code, Cowork, or a terminal git session holds ~/gpus-infra-portals at a time; the others are closed before the next is opened. Already the standing rule in CLAUDE.md ("One tool in the repo at a time"). It is the recommendation because the failure here is not concurrent writes — it is concurrent unreviewed state, and only serialising the tools removes that. It costs a habit and no infrastructure, and it is the option that also fixes the lock contention.

Option B — separate clones, one per tool, each pushing to origin independently. Removes the lock contention and makes each tool's history its own. It does not remove the finding: both clones still push to the same main, so unreviewed commits still reach origin — they arrive by fetch and merge instead of by riding along, which is more visible but not prevented. It also costs disk, divergence, and a merge discipline nobody has written down.

Option A is the recommendation. Option B without Option A is worse than either — it multiplies the trees while leaving the publish path shared.


GOV-021 — The travel form told every submitter a workflow it does not perform — CLOSED 2026-08-24

Field Value
id GOV-021
title travel-request-001's own instructions described a three-step approval order that is not the order the resolver performs — served live, to every submitter, for the entire life of the chain
register Governance / tracking gaps
status done — corrected and verified against live 2026-08-24
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Confirmed first-hand 2026-08-24, all three sources read directly. (1) The YAMLforms/travel-request.yaml at version: 1 states: "Your request goes to Admin Ops, who will route it to the SMT member you name and then to Finance." (2) The live databaseSELECT md5(instructions), length(instructions), current_version FROM forms WHERE id='travel-request-001' on gpus_forms returned c425ff6ee1beea6e8e6c1a30fd64f61c, 476, 1, byte-identical to the YAML. This column had never been verified against live before; the known DB↔repo drift did not apply here, which makes the finding worse rather than better — the wrong text is exactly what was served. (3) The implemented chainapproval_dispatch.STEP_DESCRIPTIONS is {1: "SMT approval", 2: "Admin Ops", 3: "Final approval"}, seeded in 014_approval_workflow.sql as traveller-selected smt_*travel_admin_ops (Shereyll Woodley) → travel_final (Kevin Toruno).
acceptance test The live forms.instructions column for travel-request-001 describes the implemented chain, verified by query against gpus_forms — not by reading the YAML and not by a green build. MET 2026-08-24 15:05:54 UTC — see the verification below.
blocks The all-staff travel-portal announcement — the email would describe one order while the form described another.
blocked_by
date_raised 2026-08-24
date_verified 2026-08-24

CLOSED 2026-08-24 — corrected text confirmed in the live database

Queried at the moment of closing, not carried over from an earlier check — VLN-014 is the worked example of why that distinction matters.

current_version md5(instructions) length
Before 1 c425ff6ee1beea6e8e6c1a30fd64f61c 476
After 2 e9b6e1acc6bd0b0bf3ce8f2aacc1bee9 570

The in-database md5 matches the md5 of the YAML's UTF-8 encoding exactly. Pushed 7d35371 at 15:01:16 UTC; the row flipped at 15:05:54 UTC, 4m38s later. The poll recorded nine consecutive reads of the old value and then the new one, so this is an observed boot-time load_all() reload rather than a single ambiguous read.

Verified by DB query, not by build status — deliberately. Cloud Build was not observable to either non-interactive identity, and a green build is not a booted container in any case: load_all() runs at boot, so the database is the only place the answer actually exists.

What closing this row does NOT undo

Correcting the text fixes every future submission and no past one. The attestations already recorded against the false description stay exactly as they are and cannot be remediated retroactively. This row closes because its acceptance test — live text matches the implemented chain — is met, not because that residual has been addressed.

DECIDED 2026-08-25 — the corrected name stands: step 3 is \"Final approval\"

The stale text this row corrected said Finance. It was wrong, it is gone, and the name is not reopening.

Recorded here as well as in the workflow doc because this row is where a reader lands when they go looking for why the name changed — and the answer is that it did not change. travel_final, STEP_DESCRIPTIONS[3], the corrected YAML and all four staff guides have always said "final approval"; only the served instructions text ever said otherwise. Nothing in the repository ties Kevin Toruno to Finance.

This is the residual the row's own warning is about, in its mildest form: the false text was live long enough that people read it, and some of them remember it. Correcting the column did not correct the memory. That is what this note is for.

Not a typo — a confident, specific, wrong claim, to exactly the people who needed it right

The text did not omit the order or describe it vaguely. It named a first approver (Admin Ops), named what that approver would do (route it to the SMT member), and named a third party (Finance). All three are wrong: Admin Ops is step 2, it routes nothing, and Finance is not in the chain at all.

The audience is the whole point. Form instructions are read by submitters — the one group whose correct expectation of "who has it now" determines whether they chase the right person, and whether they book travel believing it is nearly approved when it has not yet left step 1.

Submitters attested to it — this is a SECOND finding, not a detail of the first

The Acknowledgement field is required, and its only permitted value is "I understand and confirm the above." The above is the instructions block.

So every submitted travel request carries a signed attestation that the submitter understood a description of the workflow that was false. That is a different finding from "the help text was wrong", and it is the more serious of the two: the first is a defect in guidance, the second put a person's confirmation against a statement the system knew to be untrue.

It cannot be remediated retroactively. Correcting the text fixes every future submission and no past one; the attestations already recorded stay as they are. Nothing here suggests they were made in bad faith — the submitters were the only party who could not have known.

The acknowledgement makes no independent routing claim — it inherits this one by reference, which is exactly why a scan for routing language would never have found it, and why the check in test_travel_instructions.py targets the instructions rather than the fields. Nothing else in the form asserts an order: the field descriptions, the pulldowns and the travelrequest001 email template are all clean, and the form has no dividers.

Why the four staff guides did not catch it

docs/guides/travel-request-submitting.md documents the implemented order correctly and always has. The guides and the form were written from different sources and never reconciled against each other — so the estate held a correct description and an incorrect one simultaneously, in different places, with nothing comparing them.

The gap was that nothing compared the two, though both are machine-readable. The acceptance test above is deliberately a query rather than a review, so that closing this row leaves behind something repeatable.

Class fix — forms-backend/test_travel_instructions.py

Added 2026-08-24. Asserts the instructions describe the chain the resolver actually performs, by reading STEP_DESCRIPTIONS rather than restating it:

  • the stated approver count equals TOTAL_STEPS — adding a fourth step fails until the text says "four"
  • step 1 reads as traveller-selected ("the SMT member you name"), not a fixed role, which is the distinction the original text got wrong
  • the chain's landmarks appear in chain order — the load-bearing check
  • no uninvolved destination is named, which is what "then to Finance" was
  • a returned request is described as needing a new submission

Calibration. Counting three approvers would pass a rewrite that reverses them; pinning exact wording would fail on every harmless edit and be deleted the first time someone rephrased a sentence. It therefore asserts structure and vocabulary and never prose. Proven by mutation: restoring the original wording fails 4 of 5, and a rewrite naming all three approvers in the wrong order — which a count-only test passes — fails the order assertion alone.

It runs in CI. forms-backend/cloudbuild.yaml gained a travel-instructions-check step, hard-gated with no || true, alongside the existing field-exposure gate. This was necessary, not decorative: the forms-backend build ran no unit tests at all — the security scan is || true and nothing invoked unittest — so a test added without a gate would have run only when someone happened to run it locally, which is indistinguishable from not existing. The step runs this one file rather than the suite, because it needs only stdlib and pyyaml and so cannot be broken by a dependency the slim image lacks.

What it does not do. It checks the YAML that populates the column, not the live column itself, and its uninvolved-destination list is curated, so it cannot catch an invented destination nobody listed. Verifying the deploy remains a query against gpus_forms — this row's acceptance test, kept separate on purpose.

GOV-027 — "Required" is enforced as key-presence, not as a value

Field Value
id GOV-027
title create_submission tests f.field_key not in fields_in — the KEY must appear in the posted fields object, but its VALUE is never checked. A submission carrying {"FundingSource": ""} passes the required-field gate. This affects every required: true field on every form in the portal, not only travel.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence forms-backend/routes_phase2.py:286-294, read 2026-08-27 while verifying that FundingSource is enforced required for the v3 field exercise. fields_in = body.get("fields") or {} (line 219) is a dict of {field_key: value}; the gate is [f.field_key for f in form.fields if f.required and f.field_key not in fields_in and f.field_type not in _SKIP_FIELD_TYPES]. A key present with "" or null is not "missing". The live fields row confirms FundingSource.required = t, so the schema says required and the server does not enforce it as one. Confirmed by reading, NOT by a live POST — no Okta token was available in this session, so the runtime behaviour is inferred from unambiguous code rather than observed.
acceptance test A POST to /api/submissions with a required field present and empty ({"FundingSource": ""}) is rejected with missing_required_fields, demonstrated against the live endpoint. Until that POST is made, both the finding and any fix are unverified.
blocks Nothing today. The SPA supplies values, so this is reachable only by an API client — which is every future integration, and the migration tooling.
blocked_by
date_raised 2026-08-27
date_verified

Why this is separate from the known cross-field hole

The travel form already documents that FundingSource=Grant with a blank GrantName is accepted, because the platform has no conditional display and no cross-field validation. That is a stated design limit. This is different: it is the single-field required check not doing what its own column name claims, on every form. One is a capability the platform never had; the other is a control that exists and is weaker than it reads.

GOV-028 — gpus-reports has no deploy drift guard, and MAPLE ran three commits behind for two days

Field Value
id GOV-028
title The routing worker carries a VERSION file stamped into its boot log precisely because "MAPLE is not a git client". gpus-reports, deployed to the same host by the same copy-a-file method, carries nothing equivalent — no version stamp, no hash, no line in the report or the cron log naming the code that produced it. Deployed-vs-repo drift there is undetectable by design.
register Governance / tracking gaps
status done 2026-08-27 — deployed to MAPLE at 59a253c; both acceptance commands pass on the live host
priority / severity not assessed — none supplied
owner unassigned
evidence Found 2026-08-27 while staging R1/R3. /opt/gpus-reports/report_generator.py sha256 c659c86a… matches commit 10a48be (2026-08-24), three commits behind 776dc69. Two are behaviour-affecting: 65350f8 added volume_report.count_trips (the trips-vs-rows distinction) and 3cb2c5a corrected the mailer footer's sender claim. Confirmed on the host: grep -c "def count_trips" /opt/gpus-reports/volume_report.py0, grep -c trips_month_to_date report_generator.py0, and the live log line reads Volume data OK (day=2026-08-26 submitted_yday=3 submitted_mtd=12 …) with no trips_mtd= or revisions_mtd= field. R5 shipped daily to its recipient for two days without the trip-deduplication fix, over an August dataset that contains exactly the case it exists to handle (12 rows, 11 trips, 1 revision).
acceptance test Two commands, both re-runnable, both able to fail. test -f VERSION is NOT the test — the drift below would have passed it.

(a) Internal consistency — the version a LIVE RUN EMITS equals the deployed marker:
EMITTED=$(/opt/gpus-reports/venv/bin/python3 /opt/gpus-reports/report_generator.py --type volume 2>/dev/null \| grep -oE 'version=[0-9a-f]+' \| head -1 \| cut -d= -f2)
test "$EMITTED" = "$(cat /opt/gpus-reports/VERSION)"
Fails on a partial or split deploy — modules landing without VERSION, or the mailer landing without the generator.

(b) Repo currency — is there any commit newer than the deployed marker that touches a file this deploy actually ships?
git log --oneline "$(ssh maple cat /opt/gpus-reports/VERSION)"..origin/main -- 'gpus-reports/*.py' 'gpus-reports/report_cron.sh' 'gpus-reports/requirements.txt' ':!gpus-reports/test_*.py'
Empty output is the pass. Any line is an undeployed change, named.

CORRECTED TWICE ON 2026-08-27, BEFORE FIRST USE, AND THE SECOND CORRECTION IS THE INTERESTING ONE. It first read test "$deployed" = "$(git rev-parse --short origin/main)" — wrong in a monorepo, because origin/main advances for reasons unrelated to these files, and it advanced past this very row while the row was being written. Rewritten as a range over gpus-reports/, it still cried wolf on the next commit, which touched only the README. The scope has to be the staging list, not the directory: the README, the tests and this register are all in the repo and none of them are copied to MAPLE, so a change to any of them is not drift. A check that reports drift on a host that is perfectly current gets ignored within a week — which is the exact failure GOV-028 describes, reintroduced by its own fix, twice, in the space of an hour.

This is the one that fails on the drift actually found: run from 10a48be it names all four undeployed commits, including 65350f8 (which added count_trips) and 3cb2c5a. (a) alone would pass a host that is stale but self-consistent, which is exactly what two-days-behind looks like.

Both must pass. Demonstrated able to fail: as of 2026-08-27 both FAIL/opt/gpus-reports/VERSION does not exist and the modules are three commits behind — which is the finding, not a defect in the test.
blocks Nothing mechanically. It is the reason nobody could have noticed the above.
blocked_by
date_raised 2026-08-27
date_verified 2026-08-27

Mechanism built 2026-08-27 — why the row is NOT closed

The guard is implemented and tested (306 tests pass), mirroring the routing worker's rather than inventing a second pattern: a gitignored VERSION written at deploy from git rev-parse --short HEAD, read by a report_version() with the same lazy / "unknown" / OSError shape as worker_version(), emitted at four surfaces —

  • the first stdout line of every run, all eight types, before any fetch, render or upload can fail (a cron job has no boot; this is the nearest equivalent, and an aborted run still says which code aborted);
  • each degraded= summary line, for adjacency to the health signal — not sufficient alone, because those lines only exist for the four FORMS_DB_TYPES and the four SOC reports never emit one;
  • the mailer's start line, read independently of the generator, so a deploy landing one and not the other shows up as two log lines that disagree;
  • the PDF footer on every page — the only surface that outlives /var/log/gpus-reports.log. These PDFs are archived to gs://…/reports/YYYY-MM-DD/; a year on the archived PDF is the only artifact of the run, and the disputed-figure case has already happened.

CLOSED 2026-08-27 ON LIVE EVIDENCE, NOT ON "THE CODE IS WRITTEN". The row was deliberately held open while the guard existed only in the repo — the finding was about a host, and closing it on a merged commit would have been the same error it describes. It closed when the deploy ran.

Deployed to MAPLE at 59a253c by R. Chhetry (interactive sudo; cloudadmin is (ALL) ALL but password-required). Verified on the host:

  • cat /opt/gpus-reports/VERSION59a253c
  • (a) live run emits version=59a253c, deployed marker 59a253cPASS
  • (b) no commit touching a shipped file since 59a253cPASS
  • SELinux contexts system_u:object_r:usr_t:s0 on all installed files after restorecon (the three genuinely new files — summary_report.py, prebooking_report.py, VERSION — arrived as unconfined_u and were relabelled; overwritten files kept their existing correct context)

The deploy also cleared the drift the finding was about. grep -c "def count_trips" volume_report.py went 0 → 1, and the first live R5 run afterwards reported trips_mtd=11 revisions_mtd=1 submitted_mtd=12 — the trips/rows distinction that had been missing for two days, over exactly the dataset it exists to handle. The stamp is visible at all four surfaces on that run, including build 59a253c in the PDF footer.

What this does NOT solve

It makes drift visible. It does not prevent it. Nothing here stops someone deploying a stale build, forgetting to deploy at all, or writing a VERSION that does not match the modules beside it. The guard turns an undetectable condition into a detectable one, and detection is only realised by a human reading the line.

That is a thin margin here, and it should be stated rather than assumed: R5 goes to one person. The version now appears in his daily mail's footer and in a log he does not routinely read. If he does not look, the drift is exactly as invisible as it was — the difference is that the evidence exists to be found afterwards, which is what made the routing worker's near-miss attributable.

The prevention-grade fix is different work and is not attempted here: acceptance command (b) run on a timer with an alert on mismatch, rather than by hand. Raised as PRG-028 rather than left as a residual line — a residual field has no owner and no acceptance test, so it reads as closed to anyone scanning the register, and a reader who sees "drift guard: done" will reasonably conclude drift is handled. It is legible, which is a different property, and the gap between the two is exactly where the original two-day drift lived.

The stale staging objects are LEFT IN PLACE — the 021 argument, and its limit

Six objects dated 2026-08-24T18:12 sit loose in gs://gpus-infra-backups-wdc/tmp/, byte-identical to what was deployed. Hashed 2026-08-27 so the evidence survives the bucket:

c659c86a1dde5a548a999e0775851cecc3800a75d15494db7ec93930f8d25755  report_generator.py
2b46a37d3985e51d5ed7f9440d5bc2b8fe594140733c88f23a5728257e8c7889  report_mailer.py
8642fe1f10dad716357ed40b27ab1c2f37aaa2774ac135c2a1aaf69f35d67705  volume_report.py
ef681fbcbf01acc6e75a2145c35da6ef216eae6fa76fa66d4a70118977a11b55  approval_report.py
27d20f2b74f30b09189f993047da58a798115884990d4faeaf9ecff5488dd389  report_cron.sh
489dbd850b896e25842984b19370a41b2771632361c8a7aa33efef06bf100d4b  requirements.txt

Migration 021 left notification_sent_at untouched because those stamps "are the evidence of the outage … clearing them here would erase the only remaining trace." Same reasoning: these objects are the physical record of what was pushed and never refreshed.

But the precedent needed a condition to transfer, and it is worth naming. 021's stamps sat inside the system of record, where keeping them cost nothing and they could not be mistaken for live state. These sit in a path a deploy reads from — so as evidence they are also a loaded gun, and "it is evidence" would not have been sufficient on its own. What makes leaving them correct is that the new procedure fetches from tmp/gpus-reports/, a different prefix: nothing on any deploy path reads them any more. Namespacing defused them first; only then did the 021 argument apply. Had the prefix stayed tmp/, they would have had to go regardless of their evidentiary value.

Every signal said fine — this is M-13 in a second system

The daily reports arrived on schedule, rendered correctly, and logged degraded=False. Every one of those is true of the old code as well, so none of them could distinguish "current" from "two commits behind". M-13's constraint was written about the routing worker four days after the near-miss it came from; the same absence was sitting in the report pipeline the whole time, in a directory whose README documents the copy-to-MAPLE deploy in detail and never mentions verifying what landed. The lesson did not transfer, because nothing made it transfer.

GOV-029 — /opt/gpus-reports is owned by cloudadmin, so root-owned modules are not protected

Field Value
id GOV-029
title The report modules in /opt/gpus-reports/ are root:root 0644, which reads as "only root can change these". The directory containing them is cloudadmin:cloudadmin 0755, and write permission on a directory is what governs creating, deleting and renaming the files inside it. cloudadmin cannot overwrite a module in place, but can rm it and cp a replacement — no sudo involved. The root ownership of the files confers no protection it appears to confer.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Found 2026-08-27 while deploying the GOV-028 drift guard. Live on MAPLE: ls -ldZ /opt/gpus-reportsdrwxr-xr-x. 5 cloudadmin cloudadmin; ls -l report_generator.py-rw-r--r--. 1 root root. Probed non-destructively — touch /opt/gpus-reports/.perm-probe-$$ succeeded as cloudadmin and the probe was removed. test -w report_generator.py fails, confirming the in-place write is blocked and the directory route is not. Contrast /opt/gpus-forms-routing, which is root:root on both the directory and its contents.
acceptance test stat -c '%U:%G' /opt/gpus-reports returns root:root, and touch /opt/gpus-reports/.probe fails as cloudadmin. Re-runnable; currently fails. Changing it must keep /opt/gpus-reports/output/ writable by whoever the cron runs as, or every report stops — verify a full report_cron.sh run afterwards, not just the ownership.
blocks Nothing today. It weakens GOV-028: a drift guard assumes the version marker and the modules beside it were written by the same authorised deploy, and here both can be replaced without privilege.
blocked_by
date_raised 2026-08-27
date_verified

Why this is raised alongside the drift guard rather than separately

GOV-028's guard answers "is the deployed code what was pushed?" — it does not, and cannot, answer "was it put there by the deploy?" Those are different questions, and the second one only becomes interesting once the first is being asked. The guard is still worth having: this is a note on what it is evidence of, not an argument against it.

The fix is small (chown root:root /opt/gpus-reports) and is deliberately not bundled into the drift-guard deploy. That deploy already moves the host three commits and touches all eight reports; adding an ownership change to it would mean a single rollback could not separate "the new code broke something" from "the permissions change broke something", and output/ writability is exactly the kind of thing that breaks quietly.


GOV-030 — Fact D of the ASVS scope rests on a revocation path the estate does not have

Field Value
id GOV-030
title security/asvs-scope-forms-approval.md fact D argues the frozen dispatch snapshot is acceptable rather than a defect because revocation exists as "an administrative re-dispatch, which voids the outstanding step and issues a new one, and is a recorded event", and instructs the assessor to "evaluate whether that path exists and is auditable, rather than evaluating the snapshot itself as a defect". That path does not exist. VOIDED is written in exactly one place in application code, inside the decision transaction; there is no admin re-dispatch, no reassignment and no withdraw endpoint. This is not a code defect — the snapshot behaviour is deliberate and tested. It is an assurance document resting its argument on a capability the estate lacks, and it steers an assessor toward auditing a control instead of recording its absence.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Read 2026-09-01 at commit 2939812. The claim: security/asvs-scope-forms-approval.md:450-453. Restated a second time, independently: forms-backend/approval_access.py:28-30 — "Revocation is an administrative re-dispatch, which voids the outstanding step, issues a new one, and is a recorded event." Two documents, one nonexistent mechanism. The single write: forms-backend/approval_decide.py:482 (.values(status=STATUS_VOIDED)), reachable only from apply_decision's transaction and only for plan.void_step_indexes — steps ahead of a step someone just returned or declined. Nothing voids the dispatched step itself. The whole approval API surface is three routes — routes_phase2.py:881 (GET approval-view), :1062 (GET /api/approvals), :1097 (POST approval-decision). No withdraw, no reassign, no re-dispatch; grep -rn -i "withdraw\|reassign\|redispatch" forms-backend/*.py returns only prose. The estate's own admission: forms-backend/schema/021_withdraw_test_chains.sql:9-13 — "A submitter-facing withdraw is real future functionality … and it will be built. It is not what is needed today." WITHDRAWN has been in the 014 CHECK constraint since 2026-08-18 and its only writer is that hand-run SQL file (:182), which writes no audit_log row — so even the one exercised instance is not "a recorded event". Live instance: a574e4cb step 1 has been DISPATCHED to sraman@greenpeace.org since 2026-08-26 13:31 UTC — 6 days at the time of writing — with resolved_okta_email frozen. The only exits are (i) Sushma decides, or (ii) direct SQL. If she were removed from smt_sraman today the step would still be hers and nothing in the product could take it back.
acceptance test Two ways to close, and the row must not be closed by doing neither. (1) Build the path: grep -rn "approval_withdrawn" forms-backend/schema/*.sql returns an audit_action enum value, AND an endpoint exists that transitions a DISPATCHED step for a reason other than a decision, AND it is exercised end-to-end against a live row. (2) Or amend the document: sed -n '448,455p' mkdocs-portal/docs/security/asvs-scope-forms-approval.md no longer asserts that revocation exists, and instead states that snapshot authorization is unrevocable in the current build — with the same correction applied to approval_access.py:28-30, because M-14 says a correction is not complete until every co-located claim agrees with it. Both are runnable today and both currently fail.
blocks Any ASVS §2.1/§2.2 assessment. An assessor following fact D's instruction will look for an audit trail on a re-dispatch path, find nothing, and has been pre-framed to read that as "not yet evidenced" rather than "the mitigation does not exist".
blocked_by
date_raised 2026-09-01
date_verified

Why the snapshot itself is not the finding

Fact D's engineering reasoning is sound and is not disputed here: authorizing against live membership would silently strand every in-flight approval the moment a role changed, and test_membership_removal_does_not_revoke exists to stop someone "fixing" it. The freeze is the right trade.

The finding is narrower and is about assurance, not code: a scope statement is allowed to say "this is deliberate"; it is not allowed to say "and here is the compensating control" when there is no such control. The sentence converts a known, accepted limitation into a claimed capability, and it does so in the one document an external assessor is told to work inside. 021 — written 2026-08-19, two weeks before this was noticed — describes the absent withdraw path honestly and at length. The two documents have disagreed since the day the second was written.


GOV-031 — A stall at the final approval step has no chaser, and the one open case is the coordinator's own trip

Field Value
id GOV-031
title R2, the daily stalled-approvals report, is the estate's only mechanism that surfaces an approval nobody has acted on. It goes to Shereyll Woodley alone, and Kevin Toruno is excluded from it by an explicit, documented decision. Both halves are correct in isolation and together they leave one uncovered case: a request stalled at step 3, the final step, is held by the one person the chase-list is designed not to chase, and is reported only to a person with no lever over him. The live instance is worse than the general case — e5ff2b59 is Shereyll's own travel request, so the daily mail asks her to chase a step she cannot act on, on a trip that is hers. Nothing escalates it to anyone else, at any threshold, ever.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Read live 2026-09-01 from gpus_forms on MAPLE. The stall: e5ff2b59, submitted by swoodley@greenpeace.org on 2026-08-20 16:41 UTC. Its chain is 1: SKIPPED_DUPLICATE (smt_ktoruno), 2: SKIPPED_SELF (travel_admin_ops → swoodley), 3: DISPATCHED (travel_final → ktoruno@greenpeace.org) at 2026-08-20 16:41:27, decided_at NULL. The round is IN_APPROVAL current_step=3, updated_at unchanged since dispatch — 12 days. Both earlier steps were skipped because she is the submitter and Kevin was named as her SMT approver, so the chain collapsed to a single decider on the first day. The threshold: gpus-reports/approval_report.py:102, STALLED_AFTER_DAYS = 5 — it has qualified for section (a) of every R2 since 2026-08-25. The recipients: gpus-reports/report_mailer.pyapprovalsswoodley@greenpeace.org, plus rchhetry@ as a pilot CC that expires 2026-09-18. The exclusion, verbatim: gpus-reports/README.md:369-375 — "Kevin Toruno is deliberately not a recipient. This is a decision, not an oversight — please do not add him because the list looks short. R2 is a chase-list and chasing is the coordinator's job; Kevin is the final approver, so a request stalled at step 2 never reaches him and most of what this report lists is not his to act on." The gap in that reasoning: it reasons about step 2 and does not consider step 3, where the roles invert. No other mechanism exists: grep -rn -i "escalat" forms-backend/ gpus-forms-routing-worker/ gpus-reports/*.py returns only R2's own in-PDF colour band (report_generator.py:1015,1171) and two comments warning against building a timer on notification_sent_at (schema/019_outcome_notified_at.sql:92; ASVS fact E). There is no reminder mailer, no timer, no second recipient.
acceptance test Runnable now, currently failing. Every approval_steps row that is DISPATCHED, older than STALLED_AFTER_DAYS, and at the terminal step of its chain must be reported to at least one identity that is neither the step's resolved_okta_email nor the submission's submitter_email. Query: SELECT st.submission_id, st.step_index, st.resolved_okta_email, s.submitter_email, now()::date - st.dispatched_at::date AS waited FROM approval_steps st JOIN submissions s ON s.id = st.submission_id JOIN submission_approvals a ON a.submission_id = st.submission_id AND a.version = st.version WHERE st.status = 'DISPATCHED' AND st.dispatched_at < now() - interval '5 days' AND st.step_index = (SELECT max(step_index) FROM approval_steps x WHERE x.submission_id = st.submission_id AND x.version = st.version); — for each row returned, name the recipient who is told. Today it returns e5ff2b59 and the answer is nobody: the only reader is swoodley@, who is that row's submitter_email. The row closes when the query returns zero rows or when every row it returns has a named, delivered chaser. It does not close by deciding the case is rare.
blocks Nothing mechanically. It is the operational half of ASVS fact E (notification_sent_at proves handoff, not notification): a final-step approver who never saw the mail produces exactly this state, and the estate cannot tell that state apart from one where he saw it and has not decided.
blocked_by
date_raised 2026-09-01
date_verified

Both design decisions stay — this is not an argument to add Kevin to R2

README.md:369 is right that a daily mail of things you cannot act on stops being read, and that is exactly the failure that would cost him the one that matters. Adding him to R2 would fix this case by breaking the reason R2 works.

The uncovered case is narrow enough to name precisely: a DISPATCHED step at the terminal index of its chain, where the only party who could chase it is either the holder or the submitter. SKIPPED_SELF and SKIPPED_DUPLICATE make it reachable in one hop rather than three — any request where the submitter names Kevin as their SMT approver collapses both earlier steps and lands here on day one. That is a documented, supported path (travel-approval-workflow.md:882 warns submitters off it for a different reason), not an edge case.

Whatever closes this should be a different instrument from R2 — a low-frequency exception notice with a different recipient rule — not a widened chase-list.


GOV-032 — Conditional display: asked for three times, compromised around three times, DECLINED 2026-09-01

Field Value
id GOV-032
title The forms platform cannot show a field conditionally on another field's value, and the server cannot validate a cross-field condition. A dependent field can therefore be neither shown when it applies nor enforced when it does. Shereyll Woodley has now asked twice for the same shape — ask a question, then show a field if the answer is Yes — and both times the missing capability forced a different compromise that she accepted. A third instance arrived 2026-09-02 (Kevin's StayingInGPApartments, relayed by Shereyll) — see the dated note below; it was NOT requested as conditional display, which is what makes it the more useful data point. This row is a decision record, not a work item. Conditional display was raised, considered and declined on 2026-09-01 by Rajesh Chhetry and Shereyll Woodley together. It is written down so the third request is recognised as the third, and so the reasoning is available to whoever eventually reverses it.
register Governance / tracking gaps
status open — declined, recorded (no work item; not blocked, because nothing is waiting on it)
priority / severity not assessed — none supplied
owner unassigned
evidence The two requests, as recorded by Rajesh on 2026-09-01: (1) for funding, "ask whether it's a grant or a cost centre, then show the grant name field only if they pick grant"; (2) for the registration fee, "ask if there's a registration fee, and if yes show a box for the amount."The original wording is not preserved anywhere in the estate — not in the repo, not in a ticket, not in a reachable mailbox (searched 2026-09-01). These are Rajesh's account of them and are marked as such rather than presented as quotations. That the estate keeps no record of what a stakeholder actually asked for is a smaller gap inside this one. What shipped instead — request 1 (v3, 2026-08-27), forms/travel-request.yaml:52-84: three always-visible fields, FundingSource required, GrantName and CostCenter both optional, with the rule carried only by the description text — "complete this only if Funding Source is Grant. Leave blank if you selected Cost Center." The yaml comment states the consequence in advance: "a submission can be accepted with FundingSource=Grant and GrantName blank, or with both GrantName and CostCenter filled in. Nothing rejects either." What shipped instead — request 2 (v4, 2026-08-31), forms/travel-request.yaml:100-128: one required numeric, "Event registration fee (USD) — enter 0 if none." The yaml comment records that a Yes/No + amount pair was considered and rejected precisely because it "would reproduce FundingSource/GrantName exactly". Both surfaces confirmed live 2026-09-01: forms.current_version = 4 for travel-request-001. The observed cost, one data point: c84bfb45, Madison Carter, 2026-08-31 — FundingSource='Cost Center', CostCenter='23051 — COMMUNICATIONS CORE', GrantName absent. She complied with the rule correctly and unprompted. Across all four August submissions carrying FundingSource, the rule-violation query in the acceptance test returns 0. That is evidence the descriptions worked once, not evidence they scale.
acceptance test Two, and the first is the one that matters. (1) The unenforceable rule stays unbroken. Run against gpus_forms: SELECT count(*) FROM submissions WHERE form_id='travel-request-001' AND searchable_values ? 'FundingSource' AND ((searchable_values->>'FundingSource'='Grant' AND coalesce(searchable_values->>'GrantName','')='') OR (searchable_values->>'FundingSource'='Cost Center' AND coalesce(searchable_values->>'CostCenter','')='') OR (coalesce(searchable_values->>'GrantName','')<>'' AND coalesce(searchable_values->>'CostCenter','')<>''));returns 0 as of 2026-09-01 over 4 rows. A non-zero result is the first hard evidence that a description cannot carry a rule, and is the condition under which this decision should be re-opened rather than re-affirmed. Nobody is watching this number today; it is not in R1, R2, R3 or R5. (2) The written constraint stays true. grep -rn "div_lock" forms-frontend/src/ returns nothing. If it ever returns a hit, conditional display has been built and this row is superseded rather than closed.
blocks Nothing. Both requests shipped.
blocked_by
date_raised 2026-09-01
date_verified

Do not build or scope this. The decision reasoning, which is the useful part

A show_if on the field spec is not a change to one form. Fields replace-on-bootforms-backend/app.py runs load_all() on every container start — so a schema change is exercised against all 29 forms at once, on the next deploy, whether or not anyone intended to touch them. The travel form has been live two days with real requests in flight (e5ff2b59, 1b4dfd50, a574e4cb mid-chain; c84bfb45 at step 2). That is the wrong week to change the platform every form loads.

A correction to the cost estimate as it was reasoned, from source. The decision assumed the change would touch four layers: yaml_schema.py, the fields table, the form-schema API, and the SPA render. Three of the four already carry it. A div_lock key exists end to end — forms/_schema.yaml:85-92, forms-backend/yaml_schema.py:22, forms-backend/models.py:68, schema/001_init.sql:119 (div_lock TEXT), yaml_loader.py:121 and routes/forms.py:71 (it is served in the form-schema response) — and forms/README.md:74 documents it as "Conditional show/hide" — corrected 2026-09-01; that row now reads INERT — DOES NOTHING. It is inert. grep -rn "div_lock" forms-frontend/src/ returns nothing: the SPA has never read it, so every field carrying one renders unconditionally today (forms/it-support-request.yaml:41-43; finding-2026-08-06-forms-happyfox-never-dispatched.md:157-160).

⚠ COUNT CORRECTED 2026-09-01. This paragraph first read "eight legacy forms carry values for it; new-employee-notification alone has thirteen". Both figures were wrong, eyeballed from a truncated grep listing rather than counted. Counted properly — and confirmed against the live fields table, not just the yaml — it is five forms and nineteen fields: new-employee-notification 14, new-employee-notification-contractor-intern 2, and one each in rate-position-change-notification, employee-termination-notification and contract-extension-notification. forms/_schema.yaml:92 carries a twentieth as an example and is not a form — yaml_loader.py:174 excludes it from the load glob by name, which is exactly the kind of detail an eyeballed count misses. M-08 applies: a count supplied from reading rather than from a live read is provisional, and this one was published as if it were not.

This does not reopen the decision — the missing layer is the renderer, which is the expensive and risky one, and a div_lock that silently started working would change the rendering of five forms nobody asked to change, on the next deploy, with fields disappearing for users of the HR forms. That is worse than the four-layer estimate, not better: building conditional display fresh touches only the forms that opt in, while activating div_lock touches every form that already carries a value — and one of them carries fourteen. A larger blast radius than building it from scratch. The smaller corrected count does not weaken this; the fourteen-field form is the whole risk and it is unchanged.

It is recorded because the decision should rest on the true shape of the gap, and because forms/README.md:74 documented a capability the product does not have — now corrected in place, with the inventory of dead values and an explicit do not add new ones.

The second half of the constraint is independent and is not fixed by a renderer. create_submission performs no cross-field validation, so even with show_if working, "GrantName is required when FundingSource is Grant" would be enforced in the browser and nowhere else — and GOV-027 already records that required itself is enforced as key-presence, not as a value. A conditional-display build that did not also add server-side conditional validation would move the rule from unenforced and visible to unenforced and hidden, which is the worse of the two.

OBSERVATION 2026-09-01 — the first real evidence about whether EventRegistrationFee captures what was asked for, and it points the wrong way

Not a defect. Nobody did anything wrong, and no change is proposed here — the wording is Shereyll's to decide and Rajesh will raise it. Recorded because it is one real data point about a field that had existed for two days, and one data point on a two-day-old field should not be lost.

af7ea350, Jack Sundius, 2026-09-01. He entered 0 for "Event registration fee (USD) — enter 0 if none." His Purpose text says:

"Each NRO will be recharged the updated standard meeting attendance fees (1,600 Euros per participant attending physically)."

So the approval mail and the queue-45 ticket both show Event registration fee (USD): 0 above a cost total that excludes a €1,600 recharge. He answered the question exactly as asked. The question may not be the one that was meant.

What this is evidence of. EventRegistrationFee was the v4 compromise for Shereyll's second conditional-display request — one required numeric instead of a Yes/No plus an amount, on the reasoning recorded above that 0 says "none" unambiguously. That reasoning holds for the mechanics and is unaffected. What it did not settle is the semantics: a submitter can read the label as "a fee I am paying" rather than "the cost of attending this event", and a recharge billed to the NRO is invisible under the first reading. Jack read it the first way.

Why it belongs in this row rather than as a new one. GOV-032 records the cost of shipping a compromise in place of the capability that was asked for, and the entry above already says the descriptions "worked once" on Madison's submission and that this is "not evidence they scale". This is the second data point and it goes the other way: the same field family, a different submitter, and the prose did not carry the meaning. Two submissions, one each way — which is the honest state, and a better argument for revisiting the field than either one alone.

It is NOT evidence for building conditional display. The reading problem here is a label-wording problem, and a show_if would not have fixed it — a conditionally-revealed amount box asked the same way would have collected the same 0. Recorded so the two are not conflated when the third request arrives.

Watch for: the acceptance query in this row counts rule violations and returns 0; it cannot see this, because 0 is a valid answer. Nothing in R1, R2, R3 or R5 would surface it either. It was found by reading one submission's free text beside its numbers, which does not scale and is not a control.

What a third request of this class looks like, so it is recognised

Any request of the form "ask X, then show/require Y depending on the answer". Plausible next ones, from the same form: per-diem rate only if international; visa details only if a passport is needed; approver's cost-code override only if the trip exceeds a threshold; "other" free text beside any pulldown that has an Other option. All four are the same request.

When the third arrives, the choice is not "build it or refuse it" — it is between the two compromises already made, and they are not equally good. Request 2's shape is the better one and should be reached for first: a single required field whose empty case has an explicit, meaningful value (enter 0 if none) removes the dependency instead of documenting it. Request 1's shape — optional dependent fields governed by prose — is only necessary when the dependent value cannot be collapsed into the parent, and it is the shape that carries the silent-bad-data risk. Reach for a collapse before reaching for a rule in a description.

Three requests in five weeks from one stakeholder on one form is the signal that changes the arithmetic — not one more workaround, but the evidence that the platform's field model is a poor fit for the forms being asked of it. Record the third here before deciding anything.

THIRD INSTANCE, 2026-09-02 — StayingInGPApartments. It arrived without being asked for, which is the finding

What shipped. forms/travel-request.yaml v5 → v6, one field: StayingInGPApartments, a required pulldown reusing the shared Select YES/NO list, "Are you staying in Greenpeace Apartments?", placed directly after CostLodging. Requested by Kevin via Shereyll on 2026-09-02. It is not a cost field — not in COST_FIELDS, not in READ_KEYS — so it does not reach the approval mail and the routing worker was not redeployed for it.

Why it is an instance of this row at all. Nobody asked for conditional display this time. The pairing appeared anyway: if the answer is Yes, lodging should be 0 or reduced, and nothing enforces that in either layer — no show_if in the SPA, no cross-field check in create_submission. A submission with StayingInGPApartments='Yes' and CostLodging='2400' is accepted by both and nothing warns. That is the same shape as FundingSource/GrantName (request 1) and the EventRegistrationFee collapse (request 2), reached from a different direction.

This is the stronger evidence, and it is stronger precisely because it was not a request. The first two instances can be read as one stakeholder wanting a feature. This one is a plain Yes/No question from a different person that acquired an unenforceable dependency on contact with the form. The dependency is a property of the questions, not of who is asking: add a fact that qualifies a number and the platform has no way to hold the two together. Three instances, two of them unsolicited in this sense, is what the note above meant by "the evidence that the platform's field model is a poor fit".

The description deliberately does NOT say "if Yes, enter lodging as 0", and this was the live decision. The case for adding it is request 1's precedent — prose in a description is the only carrier this platform has, and it demonstrably worked on Madison's c84bfb45. The case against, which won:

  • It inverts the field's direction. Every other description on this form tells the submitter what to put in that field. This one would tell them what to put in a different, earlier field they have already passed. The SPA renders in field_index order and cannot scroll them back; by the time they read it, CostLodging is filled in above.
  • It overloads a bounded question with a cost instruction. Yes/No is the entire answer domain. Appending a monetary rule makes the label longer than the thing it labels and asks a lodging question inside an accommodation question.
  • Request 2's own reasoning applies against it. The v4 comment rejected a Yes/No + amount pair because it "would reproduce FundingSource/GrantName exactly". Putting the amount rule into the Yes/No description reproduces it in prose instead of in fields — the same unenforceable rule, one layer less visible.
  • The prose-carries-the-rule approach is 1-for-2 on live evidence — Madison complied, Jack's af7ea350 shows the reading failure — and this row already says the descriptions worked once, not that they scale. Adding a third rule to a description is spending the same untested mechanism a third time.

What was done instead. The submitter-facing guide (travel-request-submitting.md) states the pairing in its existing "nothing stops you answering inconsistently" warning, which is where the form's other unenforceable rule already lives and where a submitter reads prose that has room to explain itself. The approver sees the lodging figure and the apartments answer adjacently on /approvals and in the queue-299 ticket — that adjacency is the whole reason for the field's placement, and it is the only enforcement mechanism the platform actually has: a human reading two lines next to each other.

Acceptance test — a third query, same shape as the one above. Run against gpus_forms: SELECT count(*) FROM submissions WHERE form_id='travel-request-001' AND searchable_values->>'StayingInGPApartments'='Yes' AND coalesce(searchable_values->>'CostLodging','0')::numeric > 0; A non-zero result is the third data point on whether a rule nobody enforces survives contact with submitters. Nobody is watching this number — it is not in R1, R2, R3 or R5, exactly as the two rules before it are not. ⚠ Expect it to be 0 with no rows for some time: this field is on v6 and the five in-flight requests predate it.

SECOND CAPABILITY GAP FROM THE SAME STAKEHOLDER, 2026-09-02 — a live running cost total in the form. SCOPED, NOT BUILT

The request. Shereyll asked for a field in the form showing the total cost, updating live as the traveller types. Rajesh confirmed with her that she means the form and not the approval mail, which already itemises the four cost lines and totals them (_cost_block, approval_render.py:173). So this is not a duplicate of something that exists — it is the same information at a different moment, which is the point: the traveller cannot see the total while deciding what to enter.

This is a field request in wording only. What it needs is a computed display field — a field spec the SPA sums client-side from CostAirfare, CostGroundTransport, CostLodging and EventRegistrationFee and renders as a running total that is NOT stored: no submission_fields row, no searchable_values key, nothing to encrypt, nothing for an approver to read as an answer. That last part is what makes it a platform change rather than a form change. It touches the same four layers the show_if decision named — yaml_schema.py, the fields table and loader, the form-schema API, and the SPA render — and fields replace-on-boot, so it is exercised against all 29 forms on the next deploy whether or not anyone meant to touch them.

How it DIFFERS from the show_if decision, and the difference is real. This one is ADDITIVE. No form in forms/ declares a computed field today, so on the deploy that ships the capability, zero existing forms change how they render. Compare div_lock: five forms and nineteen fields already carry values, one of them fourteen, and activating the renderer would have made fields disappear for users of the HR forms nobody asked to change. That was a larger blast radius than building the thing from scratch. This is the opposite — the blast radius is the forms that opt in, which is initially one.

Same care required, for two reasons that survive the difference. (a) It still changes the schema every form loads at boot, so a defect in yaml_loader or yaml_schema reaches all 29 regardless of how few declare the new key — the additive property covers rendering, not loading. (b) A computed field is a new kind of field: the first one whose value is not an answer. Every consumer that assumes "a field spec means a stored value" has to be checked — the loader, create_submission's fields_in.get(f.field_key) loop (routes_phase2.py:331), the ticket template guard (which asserts placeholders equal the field list, and would now be asserting over a field that has no value to render), R1/R3/R5's key handling, and READ_KEYS. That enumeration is the actual cost, and it is not visible from the request.

Not built. Not scoped further. Recorded here because this row exists so that the case for changing the field model is assembled from real requests rather than argued in the abstract when somebody eventually proposes it — and this is the second capability gap the same stakeholder has hit in two weeks, from a different direction than conditional display. Two distinct gaps from one person on one form, plus the unsolicited third pairing above, is a materially different argument than three requests for the same feature.


GOV-033 — The executive monthly report did not send, and no signal exists that would have said so

Field Value
id GOV-033
title The monthly (executive) report aborted at 08:00:01 UTC on 2026-09-01 on a 30-second read timeout to /api/soc. No PDF, no mail, no error line — the log simply stops mid-run. Two separate defects hold it up. (1) _fetch_soc_data() returns {} for every failure, so "the SOC API did not answer" and "the SOC API answered with nothing" are the same value by the time anything can act on them. (2) The wrapper's explicit failure path is unreachable: set -euo pipefail kills report_cron.sh at the PDF_PATH=$(…) assignment when the generator exits non-zero, so log "ERROR: PDF generation failed" never runs — it has never been printed, across five successful monthlies and this failure. The report is not down; it is undetectably absent, and it will recur on the 1st of every month.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Read live on MAPLE 2026-09-01. The failure, /var/log/gpus-reports.log: 08:00:01 [report_cron] === Starting monthly report ===[report_generator] ERROR: No SOC data returned. Aborting.Read timed out. (read timeout=30) — and then nothing. No === monthly report complete ===, no gpus-monthly-report-2026-09-01.pdf in /opt/gpus-reports/output/, no [report_mailer] lines. The backend is not down: curl -o /dev/null -w '%{http_code} %{time_total}' from MAPLE at 11:05 returned 200 in 14.326 s — the timeout is sized at roughly 2× the real latency of a service that cold-starts and then SSH-fans-out to seven hosts. Defect (1): report_generator.py:313-320except Exception: return {} — and :2125-2127, if not soc_data: … sys.exit(1). A transport failure and an empty payload are indistinguishable at the decision point. Defect (2): report_cron.sh:169 set -euo pipefail; :254 PDF_PATH=$("$PYTHON" "$GENERATOR" … \| tee -a "$LOG" \| tail -1); :256-259 the if [ ! -f "$PDF_PATH" ] guard with its log "ERROR: …". grep -c "ERROR: PDF generation failed" /var/log/gpus-reports.log0, against grep -c "monthly report complete" → 5. The guard was written for an empty capture, which pipefail now prevents from ever being reached.
acceptance test Three, all runnable, all failing today. (1) The two cases are distinguishable. A run with the SOC API unreachable produces a delivered report whose first band says the SOC API could not be read — the degraded= contract R2 and R5 already use — rather than sys.exit(1). Test by pointing SOC_API at a black-hole address and asserting a PDF is produced and mailed. (2) The wrapper reports its own failures. Force a non-zero generator exit and assert ERROR: PDF generation failed appears in $LOG. Today the wrapper dies before it. (3) Absence is detectable without reading a log. gsutil stat gs://gpus-infra-backups-wdc/reports/executive/latest.pdf returns an object whose age is within the report's cadence — a freshness check that covers every report type at once, including the ones with one reader.
blocks Nothing mechanically. It is the concrete instance of the risk gpus-reports/README.md states in its own deploy warning — "a report failing at import is quieter and therefore worse: the cron job exits non-zero into a log nobody is watching, and the only visible symptom is a report that did not arrive." The same sentence describes a report failing at fetch, and the estate had no more defence against the second than the first.
blocked_by
date_raised 2026-09-01
date_verified

Which signal would have shown it, and to whom — the answer is none, to nobody

Enumerated on the host 2026-09-01, not reasoned from the design:

  • The wrapper's own error line. Unreachable — see defect (2). It has never printed.
  • /var/log/gpus-reports.log. The only tell is a === Starting monthly report === with no matching === complete ===. Nobody reads this file on a schedule, and nothing parses it.
  • cron's MAILTO. /etc/crontab sets MAILTO=root and cron did generate mail. /etc/aliases has no root: entry, and Postfix on MAPLE is loopback-only, so it lands in /var/spool/mail/root439 KB and still growing (mtime 12:15 today, the R2 run). Every report run has been mailing an unread local mailbox for months. This is the one signal that technically fired, and it is indistinguishable from the successes piled on top of it.
  • Prometheus. No metric. /var/lib/node_exporter/textfile_collector/ does not exist on MAPLE, and /etc/prometheus/rules/ holds one file (soc-log-shipper.yml) that says nothing about reports. No alert in forms-backend/grafana/prometheus-alerts.yml concerns report output.
  • Wazuh. /var/ossec/etc/ossec.conf is unreadable without an interactive sudo, so whether the agent tails this log is UNVERIFIED — recorded as unknown rather than assumed either way (M-10).
  • The recipients. CTO / CISO / VP IT would have to notice a monthly PDF that did not arrive, on a date they have no reason to hold in mind. It has arrived every month since before this workstream, which is exactly what makes its absence unremarkable.

So the only real signal was a report that did not arrive, to three people with no reason to expect it at a particular minute — and it was found by someone looking for something else entirely. That is M-11: findings come from checking what looks correct.

Recommended fix, in priority order — NOT implemented, stated for a decision

(1) Stop collapsing the two cases. This is the fix; the rest is hardening. _fetch_soc_data() should return the reason alongside the payload — or raise and let the caller decide — so the generator can tell "the SOC API did not answer" from "the SOC API said there is nothing". The house pattern already exists and should be reused rather than reinvented: R2 and R5 carry a degraded flag and still render and send, with a first band saying the data could not be read. gpus-reports/README.md states the reasoning for the travel reports — a report that arrives saying it could not read its source is information; a report that does not arrive is silence, and silence is the signal reserved for "the job stopped running". The executive monthly is the only report family that aborts instead. Aligning it is a behaviour change to a CONFIDENTIAL board-adjacent report and should be Rajesh's call, not a quiet edit.

(2) Size the timeout from measurement, and retry. 30 s against a measured 14.3 s is not a margin. A larger number alone is still a guess — but this job runs twelve times a year and nothing waits on it, so the generous option costs nothing: a connect timeout around 10 s, a read timeout of 120 s, and two retries with backoff. Do not raise the number without (1), or the next slow morning fails the same silent way, just later.

(3) Make absence detectable, once, for every report. Every type already uploads gs://gpus-infra-backups-wdc/reports/<type>/latest.pdf. One freshness check over those objects — is each newer than its own cadence allows — covers all eight reports including the two with a single named reader, and is the only item here that would have caught this failure rather than making it louder next time. It belongs with GOV-028, which made drift visible; this makes silence visible. Raised as PRG-029 on 2026-09-01, following the GOV-028PRG-028 precedent, and scoped deliberately so that it does not depend on GOV-034 — report silence should become detectable without first fixing mail on seven hosts.


GOV-034 — MAILTO=root is the estate's only unattended failure channel, and on seven hosts of seven it reaches nobody

Field Value
id GOV-034
title Every host runs cron, and cron's contract is that a job's output becomes mail to MAILTO. That is the one mechanism in this estate by which an unattended job can say it failed without anyone having built something. It is inert on all seven hosts, in two different ways. No host has a root: line in /etc/aliases, so the mail has no onward destination. Six of the seven do not run an MTA at all — Postfix is inactive on CEDAR, OAK, SKY, RAIN, SUN and WIND — so their cron output is discarded outright. MAPLE, the only host with a live MTA, keeps it: /var/spool/mail/root is 439,418 bytes and growing, roughly six months of accumulated job output that nobody has opened. Raised out of the travel workstream because it is not a travel problem, not a reports problem, and it will be the same gap under whatever runs next.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Probed across all seven hosts 2026-09-01 13:23 UTC over ssh -o BatchMode=yes. grep -c ^root: /etc/aliases0 on every host. systemctl is-active postfixactive on MAPLE only; inactive on CEDAR, OAK, SKY, RAIN, SUN, WIND. systemctl is-active crondactive on all seven. stat -c%s /var/spool/mail/root439418 on MAPLE, no such file on the other six. MAILTO=root is set in /etc/crontab:3 and /etc/cron.d/0hourly:4. The jobs this covers are not trivial: on MAPLE alone /etc/cron.d carries the AIDE integrity check (0 3 * * *), the Lynis scan, gpus-backup.sh, gpus-forms-db-backup.sh and the GCP ID-token refresh, plus the eight gpus-reports entries in root's crontab. SKY and RAIN run the AIDE checks whose promote step is already known broken. The instance that exposed it: GOV-033 — the September executive monthly died at 08:00:01 on 2026-09-01 and cron's mail was the only channel that fired. It fired into /var/spool/mail/root, indistinguishable from the months of successful runs stacked above it.
acceptance test Two, both runnable, both failing today. (1) The channel terminates at a person. getent aliases root resolves to a monitored address on every host, and an MTA capable of relaying off-host is running there. Demonstrate by forcing a non-zero exit from a throwaway cron.d entry on one host and confirming the mail arrives in a human inbox — a configuration that looks right is not the test; the delivered message is. (2) Nothing is silently discarded in the meantime. Until (1) holds, each of the six MTA-less hosts must have its cron output redirected to a file that something reads, and ls /var/spool/mail/root on MAPLE must not be the estate's only record of six months of job output.
blocks GOV-033 in part — the executive monthly's failure had one channel and this is why that channel did not work. It does not block PRG-029, which routes around it deliberately.
blocked_by
date_raised 2026-09-01
date_verified

The two failure modes are different and only one of them is visible

It is tempting to record this as "MAPLE's root mailbox is unread", because that is the instance that was found. MAPLE is the best-case host. It has an MTA, so its cron output is at least retained and can be read retrospectively — 439 KB of it is sitting there now, and the September monthly's failure is somewhere inside it.

On the other six there is no MTA, so there is nothing to retain and nothing to read later. Whether their cron output is captured anywhere at all is UNVERIFIED/etc/cron.d entries on those hosts redirect to their own log files in some cases (aide.log, gpus-backup.log) and not in others, and that has not been enumerated. Recorded as unknown (M-10) rather than assumed either way, and enumerating it is the first piece of work this row needs.

A silent discard is worse than an unread file and looks identical from the outside. Both present as "the job ran and nobody said anything", which is also what success looks like.

Why this is a GOV- row and not program work

The register's purpose is places where the estate's own records, controls or checks do not do what they are documented to do. MAILTO=root is documented in two files on every host, is the default contract of cron itself, and does nothing. Nobody built it wrongly — it was inherited, reasonable, and never tested end to end.

The building that follows from it — an alias, a relay decision for six MTA-less hosts, a monitored destination — is program work and should get a PRG- row when someone owns it. Raising it here first is deliberate: the gap is a fact about the estate today, and it should be on the record whether or not anyone schedules the fix. PRG-029 is not that fix; it is a narrower thing that deliberately does not depend on this one.


GOV-035 — /metrics is public, unauthenticated, and can be held open to the platform timeout by a container boot it has nothing to do with — RESCOPED 2026-09-01

Field Value
id GOV-035
title GET /metrics on gpus-forms-backend returned 504 at 59.99 s — Cloud Run's request timeout — on eight scrapes between 14:47 and 14:51 UTC on 2026-09-01. The handler did not block. It cannot. It is generate_latest(registry) over four in-process collectors and touches no database, no KMS, no network and no lock. What blocked was create_app(), which runs load_all() at module scope during gunicorn worker boot and took 24–34 s against Cloud SQL that afternoon. Gunicorn's master binds :8080 before its workers finish importing, so the container looks ready while it cannot serve, Cloud Run routes to it, and requests queue in the accept backlog until they age out at 60 s. The endpoint that pays for this is unauthenticated on a service granted run.invoker to allUsers with ingress: all — the same service that carries live travel approvals.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Cloud Logging, gpus-forms-backend revision 00117-d5x, 2026-09-01, read 2026-09-01. Three instances were in play and only two of them failed. Instance …9223fce6 served every scrape it received at ~4 ms, 200, right through the window. Instances …3ba075cd and …7764d0ed returned the 504s while booting, then recovered on a decaying curve — …3ba075cd went 59.99 s (504) → 51.10 s → 21.10 s → 6.09 s → 0.004 s. The causal link is timestamped to the millisecond: db.engine.created 14:51:59.556yaml_loader.done 14:52:23.856 (24.3 s in load_all) → boot.approval_orphan_check 14:52:24.306 → and then thirteen queued /metrics requests all completed 200 between 14:52:24.319 and 14:52:24.368 — a 45 ms drain of a backlog that had been accumulating for over three minutes. Thirteen queued scrapes at one per 15 s is ~3¼ minutes of Prometheus backlog held on a single container. Config: containerConcurrency: 80 against a --workers 2 --threads 4 gunicorn — 8 real slots, and Cloud Run believes there are 80. maxScale: 3, no minScale annotation (= 0), startup-cpu-boost: true, --timeout 60 in the Dockerfile matching the platform's 60 s exactly. Exposure: gcloud run services get-iam-policyroles/run.invoker includes allUsers; ingress: all; routes/health.py:45 is a bare @bp.get("/metrics") with no require_role — only /api/admin/reload is guarded. Database: gpus-forms-db is db-f1-micro, shared-core and burstable.
acceptance test Three, all runnable, all failing today. (1) A booting container receives no traffic. Configure a Cloud Run startup probe against /health and demonstrate it by forcing a slow boot (e.g. FORMS_SKIP_YAML_BOOT unset against a throttled database) while scraping: no request may be served by, or queued on, an instance whose workers have not finished importing. (2) Concurrency matches capacity. containerConcurrency equals the gunicorn slot count (workers × threads = 8) — or the worker count is raised to meet 80, deliberately, with the Cloud SQL connection maths done first on a db-f1-micro. Assert them equal in a test that reads both the Dockerfile and the service config, so they cannot drift apart again. (3) /metrics is not anonymously reachable. curl -s -o /dev/null -w '%{http_code}' https://<url>/metrics from an unauthenticated client returns 403, not 200.
blocks Nothing today. It is a latent availability and abuse surface on the service carrying travel approvals, not an active outage.
blocked_by
date_raised 2026-09-01
date_verified

The exposure, stated separately from the incident — it is the part that outlives it

The 504s were self-inflicted and harmless: Prometheus scraping its own service, and one healthy instance served throughout. The property they revealed is not harmless.

/metrics is reachable by anyone on the internet — allUsers holds run.invoker, ingress is all, and the route carries no require_role. It is on the same service and the same eight worker slots as /api/submissions, /api/approvals and approval-decision. A request that can be made to hang for 60 s by an unauthenticated caller, on a container with 8 slots that Cloud Run will send 80 concurrent requests to, and a maxScale of 3, is 24 slots total for the whole estate's travel approvals.

There IS a rate limit, and the correction matters more than the absence would have. Flask-Limiter applies RATE_LIMIT_API = "300 per hour" as a default limit to every route including this one — but its storage_uri is memory://. That makes it per instance, so three instances are 900/hour, and — the part that matters — a freshly booted instance starts with an empty limiter store. The limit is weakest at exactly the moment the queue is deepest. Prometheus alone, at 4 scrapes a minute, is already using 80 % of one instance's allowance.

Nothing here requires a sophisticated attacker. It requires someone finding the URL.

Do the metrics handler and the reports contend on the database? — NO, and the honest limit of what can be said

They cannot contend, because the handler never touches the database. generate_latest() serialises four in-process collectors. That part is settled by reading twelve lines of routes/health.py, not by inference.

What does touch the database on that service is load_all() at container boot — 29 forms, 52 pulldowns, 46 templates, in one session. So the real question is whether the reports slowed the boots, and here is every measurement that exists for today:

boot load_all context
12:31:35 8.0 s quiet — R2 finished 12:15, R5 not until 13:40
14:49:02 34.0 s inside the incident window
14:51:59 24.0 s inside the incident window

Boots during the window were 3–4× slower than the quiet one, on identical work. That is a correlation on three data points and it is not attribution. R1 and R3 each occupied the database for about three seconds (14:25:02–05 and 14:50:01–04); three seconds of reporting does not obviously produce a 26-second regression.

A better-supported explanation is available and does not involve the reports at all: the boots contended with each other. Two completed within seven seconds at 14:49:36 and 14:49:43, meaning two load_all() runs were executing concurrently against a db-f1-micro — a shared-core instance whose CPU is burst-credited. That is self-amplifying: a slow boot keeps requests queued, Cloud Run scales out, the new instance boots slower still.

State plainly: the cause of the 24–34 s boots is not established. It needs Cloud SQL CPU and connection metrics across the window, which were not captured and cannot be reconstructed after the fact. Anyone repeating this should pull them first.

min-instances=1 is MOSTLY A RED HERRING here, and this is the second time the estate's 0 default has produced a confusing signal

The instinct is reasonable: minScale is unset (= 0), so cold starts are expected, and a cold start is the usual explanation for a slow first request. The logs refute it as the explanation for these 504s.

Instance …9223fce6 was warm and healthy for the entire window, answering in ~4 ms. There was no scale-from-zero to survive. min-instances=1 guarantees one warm instance — which already existed — and does nothing about the actual failure, which is that Cloud Run routed scrapes to additional instances that were still booting. minScale: 1 does not prevent scale-out, and it is scale-out that hurt.

Where it is not a red herring: it reduces how often a boot happens at all, so it lowers the frequency of the window in which this can occur. It is a mitigation of exposure, not of mechanism, and it costs money continuously. Do not buy it as a fix for this.

The mechanism fix is a startup probe. Cloud Run's readiness signal today is the open socket, and gunicorn's master binds :8080 before its workers finish importing app:app — so a container advertises readiness it does not have. A startup probe against /health closes exactly that gap, costs nothing, and is the only recommendation here that addresses the cause rather than the odds.

Recorded as a pattern because it has now happened twice: the estate's min-instances=0 cost default makes "cold start" the first hypothesis for every latency anomaly, and it has been the wrong one both times. Check whether a warm instance was serving before reaching for it.

RESCOPED 2026-09-01 — the incident this row was raised from was a MEMORY-EXHAUSTION OUTAGE, not a boot window

The boot-window analysis below stands and is unchanged. The mechanism is real, it is evidenced to the millisecond, and it is why /metrics can be held to the platform timeout. What was wrong was the incident it was attached to.

Later the same afternoon the service went genuinely down for roughly twenty-five minutes, and the container-lifecycle log says why:

16:40:33  Starting new instance          (AUTOSCALING)
16:40:43  terminated: Application failed to start:
          Waited too long for connection to be ready
16:46:00  terminated: Application failed to start   (again)
16:46:38  terminated: Application failed to start   (again)
16:47:08  terminated: Application failed to start   (again)
16:47:38  terminated: Application failed to start   (again)
16:47:54  Uncaught signal: 7, pid=1, tid=1, fault_addr=139856018710528
16:47:54  Container terminated on signal 7
16:48:43  Uncaught signal: 7, pid=1  →  Container terminated on signal 7
16:51:41  Container called exit(1)

Signal 7 is SIGBUS at pid 1, and on Cloud Run that is the signature of a container hitting its memory ceiling. The revision was 512Mi. Between 14:55:58 and 17:10:49 not one container completed create_app() — no yaml_loader.done in that entire window — so every instance Cloud Run started died mid-boot and every request queued behind one and aged out at 60 s.

Three things this changes, and the third is the important one.

  • It began before the deploy. The SIGBUS crashes at 16:47–16:51 are on revision 00117-d5x. The push that built 00118-9cs did not start until 16:58:43. The outage was not caused by shipping the destination switch, and 00118 then failed identically — rollback would not have helped, which is worth knowing because rollback is the reflex.
  • The database was fine. Measured from MAPLE during the outage: 0.3 s for SELECT 1, 11 of 25 connections in use, nothing long-running. Every earlier hypothesis that reached for Cloud SQL contention — including this row's own, recorded as "cause not established" — was looking in the wrong place.
  • THE STARTUP PROBE IS NOW THE FOLLOW-UP, NOT THE FIX. A probe would have turned a fleet of silently-504ing instances into a revision that visibly refused to go live, which is a large improvement in legibility — but the service would still have been down. The fix was 512Mi → 1Gi. Recorded because this row previously named the probe as "the mechanism fix", and against a memory ceiling it is not a fix at all.

Resolution. Revision 00119-wls, memory 512Mi → 1Gi, nothing else changed. yaml_loader.done at 17:10:49 — 29 forms, 52 pulldowns, 46 templates, 0 errors, the first completed load since 14:55 — orphan check CLEAN (6 of 6), /health 200 in 0.76 s where it had been 504 at 60 s.

The probe still belongs, and its priority is now argued on different grounds: see the alerting note below.

AMENDED 2026-09-01 — \"nothing else in the project failed\" is not evidence, and checking it changed the picture

The scoping argument offered for this incident was: only gpus-forms-backend, only /metrics, nothing else in the project in that window — therefore not a platform fault. The conclusion is right, but that argument does not support it, and the check that was asked for is the reason.

Cloud Logging for 14:43–14:52 UTC, per service:

Service /metrics scrapes any requests at all
gpus-status-backend 0 0
gpus-soc-backend 0 0
gpus-forms-frontend 0 0
gpus-mkdocs-portal 0 0
gpus-forms-clamav-worker 0 1

They were not being scraped, so they had no opportunity to fail. Their silence carries no information whatsoever. Confirmed at the source rather than inferred: /etc/prometheus/prometheus.yml on MAPLE has exactly one Cloud Run job — gpus-forms-backend. Every other job is a node_exporter on a host (:9100) or a local exporter. One Cloud Run service out of roughly ten is under metric monitoring at all.

This does not weaken the finding. The mechanism was established positively — from per-instance latencies, the load_all timing and the thirteen-request drain — not by elimination. But the elimination leg was load-bearing in how the incident was first described, and it was hollow.

A quiet neighbour is only evidence if someone was listening to it. Before reading "nothing else failed" as scoping, establish that anything else would have been able to report a failure.

It also sharpens the detection finding below: the estate has one Cloud Run metrics target, so this class of fault is invisible on nine other services, including gpus-soc-backend — whose slow /api/soc is the subject of GOV-033.

FormsBackendDown was a TRUE POSITIVE, and it is the only reason anyone looked

Recorded prominently because it was dismissed on first reading, and the dismissal was wrong. FormsBackendDown fired at 16:44:12 via Alertmanager and opened #USSEC00368982 in US - IT Security. It was initially called a boot-window false alarm and closed as noise. It was not. Scrapes began failing around 16:41; the rule's for: 3m elapsed at 16:44; the service was in fact down and stayed down for another twenty-five minutes.

An alert dismissed as noise turned out to be the one signal in the estate that worked. No report, no log watcher, no cron mail and no other alert surfaced this — and by the register's own accounting (GOV-034, PRG-029, and the scrape-coverage amendment below) none of them could have.

The temptation to close an alert as a known false positive is strongest exactly when a known false positive shares its signature. The boot-window analysis in this row supplied a ready explanation that fit, and it was wrong. Before closing an alert against a known benign cause, confirm the benign cause is present, not merely plausible — here that was one curl to /health.

The rule is correctly built and needs no change. up{job="gpus-forms-backend"} == 0, for: 3m, severity: critical. up goes 0 on any scrape failure including a timeout, which is what makes it the only rule in the file that can see this class of fault at all.

No other rule has the same exposure, checked rather than assumed. The other four in forms-backend/grafana/prometheus-alerts.ymlFormsBackendHighErrorRate, FormsDecryptFailures, HappyFoxIntegrationDegraded, FormsRoutingWorkerDown — all key on rate() or increase() over metric values. When scrapes fail there is no data, so they go stale and silent rather than firing. FormsBackendDown is the only rule that treats scrape failure itself as the condition. That is correct design, and it is precisely why it is the one that fires on a boot window.

The cost of that correctness: every boot window bills a manual close to the SECURITY queue

The Alertmanager path does not auto-resolve — one incident = one ticket — so each firing costs a human closing a ticket in US - IT Security, the queue where real security alerts land. And every gpus-forms-backend deploy produces a boot window, so every deploy produces one.

That raises the startup probe's priority on a different argument from the one this row opened with. It is no longer "cleaner signals". It is stop generating security-queue tickets that a human must triage and close, because alert fatigue aimed at that queue is paid for by the next real security alert being read more slowly.

Note the tension and do not resolve it by weakening the rule: the same property that makes FormsBackendDown noisy on deploys is what made it the only detector of a real outage today. Fix the boots, not the alert.

Nobody would have seen this WITHOUT that alert, and that is the third finding

It was seen because Grafana's "Server Down" alert happened to fire at 21:12 local and resolve. That alert exists for the service being down — the service was not down; one of three instances was answering normally throughout.

Nothing else would have surfaced it. There is no alert on 5xx rate, none on request latency, and none on scrape failure. GOV-034 records that cron's MAILTO=root reaches nobody, and PRG-029 covers report silence — this is a third instance of the same shape: the estate cannot tell anyone that something failed quietly, and each time it has been found by a person looking at something else.

A 5xx-rate alert on the Cloud Run services is the smallest thing that would have caught this one. It is not raised as work here because PRG-029 and GOV-034 are already open on the same theme and should be scoped together rather than three times.

Not fixed in this pass, deliberately — and what NOT to do first

Recommended order, cheapest and most causal first:

  1. Cloud Run startup probe on /health. Addresses the mechanism. No cost, no code change.
  2. Reconcile containerConcurrency (80) with the 8 real gunicorn slots, and pin the two together in a test. Today Cloud Run may send ten times what a container can serve.
  3. Authenticate or remove /metrics from public reach. Prometheus scrapes from MAPLE, which already holds run.invoker as maple-agent@ — the anonymous grant is not needed for this path.
  4. Move load_all() off the request-serving boot path, or make it fail fast. 24–34 s of database work before a worker can answer is the underlying cost; the probe hides it from callers but does not remove it.

Do not start by raising --timeout or the Cloud Run request timeout. That converts a 504 into a longer hang and consumes a worker slot for longer, on a service with eight of them.


GOV-036 — The estate's Cloud Run cost defaults were set for a portal with no traffic, and it now carries live travel approvals

Field Value
id GOV-036
title memory=512Mi, minScale=0 and containerConcurrency=80 are carried by almost every Cloud Run service in gpus-infra. All three were correct when they were set — on portals that served a handful of internal readers and cost nothing to keep small. They are no longer being asked the same question. gpus-forms-backend now carries the estate's live travel-approval workflow and is scraped every 15 s, and on 2026-09-01 two of those three defaults produced failures on the same afternoon: minScale=0 sent the first diagnosis of a latency anomaly toward "cold start" twice, wrongly both times, and 512Mi took the service down for twenty-five minutes with SIGBUS at pid 1. This row is not "raise the numbers". It is that the numbers are inherited defaults being read as decisions, and nobody has re-asked them against what the estate actually runs now.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Surveyed live 2026-09-01 across all ten Cloud Run services in gpus-infra. containerConcurrency=80 on every one of the ten, without exception — including gpus-forms-backend, whose gunicorn is --workers 2 --threads 4 = 8 real slots, so Cloud Run may route ten times what a container can serve. minScale=0 on all ten. memory=512Mi on seven of tengpus-mkdocs-portal, gpus-security-backend, gpus-security-site, gpus-soc-backend, gpus-soc-site, gpus-status-backend, gpus-status-site — with gpus-forms-frontend at 256Mi, gpus-forms-clamav-worker at 2Gi (sized deliberately, for ClamAV signatures) and gpus-forms-backend now at 1Gi only because of today's outage. The failures: 512MiUncaught signal: 7, pid=1 on 00117-d5x at 16:47:54 and 16:48:43, Container called exit(1) 16:51:41, no create_app() completing between 14:55:58 and 17:10:49; fixed by 512Mi → 1Gi alone (00119-wls), nothing else changed. minScale=0 produced the wrong first hypothesis on both the 14:47 boot window (GOV-035 — a warm instance was serving at ~4 ms throughout) and again during the outage triage.
acceptance test Not "the numbers are bigger" — that is unfalsifiable and would close on the wrong thing. Each service carries a recorded, dated answer to what workload is this sized for, and the setting is derivable from it. Runnable as a review: for every Cloud Run service, containerConcurrency is either equal to the container's real request-slot count or documented as a deliberate over-subscription; memory names the workload that sets its ceiling (for gpus-forms-backend that is load_all() over 29 forms at boot); and minScale states whether cold starts are acceptable for that service's consumers, naming them. Currently failing on all three for all ten — no service has any of it written down.
blocks Nothing mechanically. It is the shared cause behind GOV-035's concurrency mismatch and today's outage, and it is why GOV-033's /api/soc timeout is worth re-examining: gpus-soc-backend is 512Mi, concurrency 80, minScale 0 and takes 14.3 s to answer a scrape — the same shape, on a service nothing monitors.
blocked_by
date_raised 2026-09-01
date_verified

The defaults are not wrong. They are unexamined, and that is a different problem with a different fix

It would be easy to read this row as "512Mi was too small" and close it by raising numbers. That misses it.

Every one of these settings was a correct decision when made. A docs portal and three read-only dashboards, serving an internal audience, genuinely should scale to zero and genuinely do not need a gigabyte. The defaults saved real money for real reasons and there is no error to apportion.

What changed is the question, not the answer. The estate now runs:

  • a live travel-approval workflow on gpus-forms-backend, where a failed boot means approvers cannot decide and submitters cannot submit;
  • four scheduled report jobs against the same database, on a cadence nobody watching a dashboard would predict;
  • a Prometheus scrape every 15 s, which is a synthetic caller that never stops and turns any slow boot into a queue.

None of those existed when 512Mi and minScale=0 were chosen. A default becomes a decision the moment the workload changes and nobody re-asks it — and the tell is that it took a SIGBUS to find out, not a review.

What else in the estate carries defaults set under different assumptions

Enumerated, with what each is now being asked to do. None of these is a recommendation to change anything today — the point is that each is an unexamined inheritance and should be adjudicated, not adjusted by reflex.

Setting Where Set for Now asked to
containerConcurrency=80 all ten Cloud Run services the platform default; nobody chose it route up to 80 concurrent requests to gpus-forms-backend's 8 gunicorn slots (GOV-035)
memory=512Mi 7 of 10 services static sites and thin read APIs gpus-soc-backend, which fans out SSH polls to seven hosts and answers in 14.3 s (GOV-033)
minScale=0 all ten portals nobody was watching a service scraped every 15 s and carrying live approvals
db-f1-micro gpus-forms-db Phase 1, a forms portal with no submissions the approval workflow, submission_fields encryption, and four report jobs. Shared-core and burst-credited — max_connections is 25
--timeout 60 gpus-reports Dockerfile gunicorn matching Cloud Run's request timeout the same, but it now equals the platform timeout exactly, so a slow worker and a platform timeout are indistinguishable in the logs
MAILTO=root all seven hosts inherited from cron itself the estate's only unattended failure channel (GOV-034)
RATE_LIMIT_API memory:// forms backend a single instance up to 3 instances, each with an empty limiter store on boot (GOV-035)

The Cloud SQL tier is the one to look at next. db-f1-micro was sized for a portal with no submissions and now backs the approval chain and four scheduled readers. It has not failed, and this row does not claim it will — but it is the same class of inheritance as the one that failed today, and it is the only one on this list whose failure mode is data.


GOV-037 — The travel reports take a revised trip's values from the submission that was rejected

Field Value
id GOV-037
title A revised trip is two submissions rows: a root that was returned, and a child that replaced it. The reports count the chain once, by its root — summary_report.py:271, trips = [s for s in rows if s.get("parent_submission_id") is None], and group_by_key(roots_only=True) — and counting one chain as one trip is correct. Taking the root's values is not. Every breakdown then describes the version that was rejected: by-SMT attributes the trip to whoever was named on the original even if the revision named someone else; department, funding source and cost centre all come from the returned row; and coverage percentages count the root's answers rather than the ones actually approved. The rule the reports should apply is: chain root for IDENTITY, the authoritative row for VALUES. This is not a cost rule. It is how the whole report family should read a chain, and the cost section only exposed it because a wrong sum is arithmetically obvious in a way a wrong department attribution is not.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Read 2026-09-01. Defect site: gpus-reports/summary_report.py:271 and :292 (group_by_key, roots_only=True), and prebooking_report.py:240 (if s.get("parent_submission_id") is not None: continue). Both then read searchable_values off the row they kept. The estate has exactly one revision chain, 81c74ae935750caf (August 2026): the root reached RETURNED_FOR_REVISION, the child reached APPROVED on 2026-08-26. Full key-by-key diff of the two rows, run live: AlreadyBooked, CostAirfare, CostGroundTransport, DepartDate, Department, ReturnDate, SMTApprover, TravelScopeall identical. Exactly one key differs: CostLodging 113 → 140.
acceptance test A single authoritative_row(chain) helper exists, is defined once, and every value-reading section calls it — no section derives its own. Demonstrated able to fail with a fixture chain whose child changes Department and SMTApprover: the by-department and by-SMT tables must report the CHILD's values while the trip count stays 1. Currently no such helper exists and the fixture would report the root's values, so the test fails on today's code in both summary_report and prebooking_report.
blocks Nothing shipped is wrong today — see the note below — but every future revision that changes an attributed field silently mis-states a breakdown, with no signal.
blocked_by
date_raised 2026-09-01
date_verified

(a) CORRECTED — the 2026-09-01 editions SHIPPED WITH THE DEFECT BUT ARE NOT WRONG, and the distinction matters

This row was raised on the understanding that R1's 1 September edition — which Kevin has already read — carries wrong figures for one trip's attributes. Checked field by field, and it does not.

August contains exactly one chain, and the diff between its root and its approved row is:

AlreadyBooked        No                     No                     same
CostAirfare          2000                   2000                   same
CostGroundTransport  200                    200                    same
DepartDate           2026-10-10             2026-10-10             same
Department           Information Technology Information Technology same
ReturnDate           2026-12-10             2026-12-10             same
SMTApprover          Kevin Toruno           Kevin Toruno           same
TravelScope          International          International          same
CostLodging          113                    140                    *** DIFFERS ***

The only field the traveller changed is CostLodging, and no report currently reads it. R1 reads Department, SMTApprover, FundingSource, CostCenter; R3 reads AlreadyBooked, Department, SMTApprover. Every one is identical across the chain. So both published editions are accurate, and no correction is owed to Kevin.

Recorded this way deliberately, and it is the more useful record: the defect shipped, it is real, and it did not manifest — because the one revision in the estate happened not to change any attributed field. That is luck, not correctness, and the row stays open on the code rather than closing on the outcome.

Issuing a correction notice for figures that are right would have been the worse error. A retraction to the report's only reader, on his second ever edition, teaches him the numbers are unreliable — and it would have been untrue. M-08 in the other direction: a figure asserted wrong from reading the code rather than the data is as unsound as one asserted right from memory.

The single differing field is the one the cost section reads. That is not a coincidence worth passing over

CostLodging went 113 → 140 in the revision. It is the only changed value in the estate's only chain, and it is a cost field — precisely what the new approved-cost section reads and what no existing section did.

Had that section been built on roots_only=True like every section beside it, it would have summed 113 for that trip instead of 140: the lodging figure from the submission that was rejected. Small in dollars, and it is the whole finding in one number.

Revisions change the thing that was wrong. A trip is returned for a reason, and the reason is usually a value — so the fields most likely to differ across a chain are exactly the fields a reader most wants attributed correctly. A defect that reads superseded values is therefore biased toward being wrong about whatever mattered, not randomly wrong.

(b) AUTHORITATIVE ROW — defined once, here, not derived per section

The definition varies by the chain's state, which is why it must be stated rather than reimplemented three times:

Chain state Authoritative row Why
Reached APPROVED the row that reached APPROVED it is what was authorised; earlier rows were superseded
Still in flight (IN_APPROVAL / BLOCKED) the latest submission in the chain it is what approvers are currently looking at
RETURNED_FOR_REVISION, no revision filed yet the root nothing has replaced it; the root is still the only account of the trip
DECLINED / WITHDRAWN the row that reached that state same reasoning as approved — it is the version the outcome was reached on

Identity is always the chain root, in every state. Only the value source moves. Trip counts, denominators and dedup must not change: 1/12 stays 1/12.

R3 shares this defect. R5 does not, and the reason is worth keeping

R3 (prebooking_report.py) HAS IT, in the same shape. :240 skips rows with a parent, then reads AlreadyBooked off the root and groups the yes-rows by the root's Department and SMTApprover. It is arguably sharper there than in R1: AlreadyBooked is the entire subject of the report. A traveller who booked before approval, was returned, and corrected the answer on the revision would be reported under the superseded answer — and the pre-booking rate is R3's headline number. Its 2026-09-01 edition is nonetheless correct, for the same reason R1's is: AlreadyBooked reads No on both rows of the only chain.

R5 (volume_report.py) DOES NOT HAVE IT, and not by accident. count_trips uses the root purely for identity, and R5 reads no searchable_values key at all — the only two mentions in the module are comments asserting exactly that (:55, :482). A report that attributes no values cannot attribute them to the wrong row. The strongest boundary in this family turns out to also be the one that made it immune, which is worth noting the next time that constraint feels expensive.

GOV-038 — 10.8.0.0/28 became a default allow-list by accretion, and nothing ever asked what it had accumulated

Raised 2026-09-08. Sits above VLN-053 and VLN-054 rather than merging with them: those two stay separate because their remediations differ. The question here is different in kind — not "is this service exposed" but "how did one source address acquire standing to reach eight of them, and who decided that?"

What the origin reaches

10.8.0.0/28 is the gpus-vpc-connector, the shared egress path for Cloud Run. Read off the two verbatim firewall-cmd --list-all captures of 2026-09-08:

Host Service Port Consumer in the repo Verdict
CEDAR SSH 22 backends SSH to collect metrics load-bearing
CEDAR Elasticsearch — the alert corpus 9200 soc-backend/app.py:78,408 via CEDAR_ES_URL load-bearing, and anonymous
CEDAR Kibana 5601 none found accreted
CEDAR Webmin (root-capable) 10000 none found accreted
MAPLE SSH 22 backends SSH to collect metrics load-bearing
MAPLE Prometheus 9090 soc-backend/app.py:298 via PROMETHEUS_CLOUD load-bearing
MAPLE Grafana 3000 none found accreted
MAPLE Webmin (root-capable) 10000 none found accreted

Eight services across two hosts from one source address, including two root-capable administration interfaces and the estate's complete security-alert corpus.

The finding is the mixture, not the count

An earlier framing of this — in the session that produced VLN-053 and VLN-054 — asserted that "none of these is a path any human uses" and that "every one of those rules could likely be deleted with no operational loss". The first half is true. The second half was wrong, and checking it is what produced the table above.

Three of the eight are genuinely load-bearing. The Cloud Run backends SSH into hosts to collect metrics, soc-backend reads CEDAR's Elasticsearch, and soc-backend queries MAPLE's Prometheus. Delete those rules and the SOC and status portals stop working.

Five have no consumer anywhere in the repository — Kibana, Grafana, and Webmin on both hosts. Nothing in soc-backend, status-backend or security-backend references port 5601, 3000 or 10000; the only textual hits are a docstring listing data sources.

That mixture is the governance problem. The rules are indistinguishable at the firewall: same source address, same accept verb, same zone. Nothing records which exist to serve a named consumer and which are residue. So a reviewer cannot narrow the origin without risking an outage, and an attacker inherits all eight for the cost of one foothold. The absence of "why" is the gap — not the absence of "what".

How an origin accretes standing

No single rule here looks wrong when it is added. A backend needs to read alerts, so 9200 is opened to the connector. Someone opens Kibana to look at the same data from the same place. Webmin gets opened because it is how the host was being administered that week. Each is a small, locally reasonable decision, justified against the service being added.

Nobody is ever presented with the accumulated total. firewall-cmd --list-all prints rules grouped by nothing in particular — the CEDAR capture lists its 16 rich rules in no discernible order — so the question "what can this one source address now reach?" is never on screen unless somebody deliberately assembles it. This register row is the first time it has been assembled for 10.8.0.0/28.

Who is on the other end

Five of the ten Cloud Run services in gpus-infra attach to this connector: gpus-forms-backend, gpus-forms-clamav-worker, gpus-security-backend, gpus-soc-backend, gpus-status-backend. All five are ingress: all (cf. VLN-018). The five static frontends do not attach and cannot reach any of it.

So the distance between the public internet and eight internal services, including two root-capable ones, is one server-side request-forgery or code-execution defect in any of five internet-facing services. This is not a claim that the internet can reach them — it cannot, and no such defect is alleged here.

The exposures are one command from widening

The GCP firewall and the host firewalls disagree, and the host firewall is the narrower. cedar-ingress permits tcp:9200 from 192.168.120.0/23 — the whole WDC LAN, containing 113 workstations per the IAR — while CEDAR's firewalld does not. Only the host layer withholds it.

So every exposure catalogued above is one systemctl stop firewalld from widening to the WDC LAN, with no cloud-side change and no alert. A host firewall is routinely stopped during troubleshooting. That is why aligning the cloud rule with the host rule is worth doing even though it changes nothing today (VLN-053 §7 Step 1).

The same gap exists on the cloud side of the firewall, and it bites harder

The connector rules are rich rules on the host, one port per rule, so they can be reasoned about individually. The GCP rules cannot. cedar-ingress is a single object granting 22, 5140, 5601, 9200, 10000, 9100 to all three of its source ranges at once; maple-ingress grants 22, 9090, 3000, 10000, 1514, 1515, 9100 the same way. A GCP firewall rule has one port list and one source list — it cannot express "this range gets only this port".

So the cloud layer has the accretion problem in a more dangerous form: not only is there no record of which grants are load-bearing, the grants are fused. Narrowing one source range necessarily narrows it for every port in the rule.

This turned a proposed "no-op" into a production outage. Removing 192.168.120.0/23 from cedar-ingress was written up as a safe alignment of the cloud rule with the host rule — true for 9200, 5601, 10000 and 9100, which the host firewall already refuses from that range. It is false for 5140, which the host firewall permits and the WDC estate actively uses: 9,192 of ~11,862 documents in cloud-logs-2026.09.08 (about 78%) came from 192.168.120.1, .2 and .3. The same removal on maple-ingress would have cut 1514/1515, the Wazuh agent channel for all four WDC hosts — ~77% of that day's alerts.

The fix is to split each rule so the WDC range keeps only the port it uses, and to create the narrow rule before narrowing the broad one. Detail lives in VLN-053 §7 Step 1.

The governance point, restated: a fused grant makes "what is this for?" unanswerable per-port, so nobody can narrow it safely without first measuring what flows through it. That measurement did not exist until this row was written, on either host.

Acceptance test

Two halves, both able to fail:

  1. Structural. A written statement of what 10.8.0.0/28 is permitted to reach on each host, with every rule tied to a named consumer — a service, a file, a line number — or marked for removal. It fails today: five of eight rules have no consumer anyone has named.
  2. Live, and it is the one that closes the row. After the accreted rules are removed, the SOC and status portals still render their Wazuh, Prometheus and host panels — verified in the portals, not inferred from the services being up. Per VLN-010, a silent stop looks exactly like success.

What this row does not claim

  • It does not claim the five accreted rules are unused. It claims no consumer was found in this repository. Something outside the repo — an operator's browser, a script on another host — could use them. That is what the acceptance test's first half is for: naming a consumer is a valid outcome.
  • It does not propose deleting the three load-bearing rules. They should be documented and secured, not removed. VLN-053 covers securing 9200.
  • It does not restate VLN-053 or VLN-054. Their remediations are specific and stay with them.

Step 1 of the split applied 2026-09-09 — the fix is now in the objects, not only recorded about them

Two narrow rules created in gpus-vpc, priority 800 ahead of the broad rules at 900, so they take effect immediately and the later narrowing is a non-event rather than a cutover:

Rule Ports Source Target tag
cedar-ingress-wdc-syslog tcp:5140 192.168.120.0/23 cedar-logging
maple-ingress-wdc-agents tcp:1514, tcp:1515 192.168.120.0/23 maple-monitoring

Each description names its consumers and its senders, which is this row's remedy applied to rules being created rather than only observed about rules that exist. Target tags match the existing objects exactly, so the new rules attach to the same instances.

They are not yet proven. The broad rules still permit the same traffic, so both feeds flow identically whether the narrow rules match or not — a signal that would be true either way, M-13. The narrow rules are unproven until step 3 removes 192.168.120.0/23 from the broad objects.

Pre-step-3 baseline, measured 2026-09-09 17:09–17:13 UTC. This is the comparison target, not proof the new rules work:

| Feed | Delta over 180 s | Per host, last 2 h |
|---|---|---|
| syslog `5140` | **+93** | `sky` 1926, `rain` 1583, `sun` 969 |
| Wazuh `1514` | **+4** | `sky` 244, `rain` 242, `wind` 216, `sun` 116 (plus `maple` 192, `cedar` 104, `oak` 60) |

All three syslog senders appear individually over a wide window, and **all
four WDC hosts are confirmed live Wazuh agents**. After step 3 these
figures must still move, per host, over a window of the same width — SKY
and RAIN are bursty and a ten-minute window fails for the wrong reason.

**SSH survives step 3 independently**, verified from the objects rather
than assumed: `allow-onprem-via-vpn` at priority 1000 grants `tcp:22` to
`192.168.120.0/23` and `192.168.124.0/24` in its own rule, so narrowing
`cedar-ingress` and `maple-ingress` cannot cut off access.

Three facts read out of the objects that change step 3

Recorded now rather than rediscovered later. cedar-ingress and maple-ingress are each one object at priority 900 carrying three source ranges — 10.8.0.0/28, 192.168.120.0/23, 172.16.0.0/24 — so removing a range strips it from every port in the object.

1. cedar-ingress grants 9200 to the whole WDC LAN, and narrowing closes VLN-053 as a side effect. The object's own port list is 22, 5140, 5601, 9200, 10000, 9100. So the cloud layer permits every host in 192.168.120.0/23 to reach CEDAR's unauthenticated Elasticsearch, and only firewalld withholds it. That is VLN-053's two-layer disagreement confirmed from the firewall object itself rather than inferred. It makes step 3 more valuable than it appeared: it is not only tidying, it removes a standing grant to an unauthenticated indexer.

2. Grafana 3000 and Prometheus 9090 from the WDC LAN are unestablished, and this must be answered BEFORE step 3, not during. maple-ingress carries both. If anyone reaches Grafana from the office, narrowing that rule breaks it with no warning and no obvious cause. Nothing in the repository establishes a WDC consumer either way. An open question in front of a change is a reason to stop, not a caveat to carry through it.

3. Port 10000 is now dead config in both rules and should be REMOVED at step 3, not narrowed. Webmin was disabled on MAPLE and CEDAR on 2026-09-09 (VLN-054), so both objects permit a port to a service that no longer exists. That is precisely the accretion this row describes — with the difference that this instance was created today by our own change, so removing it is cleanup rather than archaeology. A rule that outlives its service is how the next 10.8.0.0/28 starts.

Also noted: 9100 node_exporter is scraped from MAPLE in 172.16.0.0/24, so the WDC range is not needed for it either. And flow logging is disabled on both objects, which means there is no record of what has actually used these grants — the evidence that would answer question 2 was never being collected.

GOV-039 — every WDC syslog event is indexed four hours in the past, and the filters meant to parse them are dead code

Raised 2026-09-08. Three defects in one pipeline (/etc/logstash/conf.d on CEDAR). None loses data; all three corrupt or fail to enrich it, and the first one manufactured a false incident before it was understood.

(1) The collector parses EDT timestamps as UTC — the root cause

WDC hosts run EDT (UTC−4); cloud hosts run UTC. RFC3164 syslog carries a local timestamp with no zone. The Logstash syslog input parses it in the collector's own zone, so Sep 8 13:53:52 sent from SUN at 17:53 UTC is indexed as 13:53:52.000Zfour hours early.

Measured 2026-09-08 at 17:53:55Z, newest document per source:

Source docs last @timestamp offset
172.16.0.12 (MAPLE) 3,592 17:53:52.554Z current
192.168.120.3 (SUN) 5,554 13:53:52.000Z −4h
192.168.120.1 (SKY) 3,257 13:36:23.000Z −4h
192.168.120.2 (RAIN) 2,287 13:32:31.000Z −4h

Nothing is lost — every WDC event arrives and is indexed. But any time-bounded query silently excludes them: a now-10m window returns MAPLE only, and an hourly histogram shows WDC traffic apparently "stopping" exactly four hours before the current hour.

This defect fabricated a 4.5-hour outage on 2026-09-08

Acting on those queries, the author reported a syslog outage, diagnosed half-open TCP sockets after a tunnel flap, and recommended systemctl restart rsyslog on three production hosts. There was no outage. The restarts were unnecessary (harmless), and the "half-open sockets" reading of ESTABLISHED connections with empty queues was confirmation bias — that state equally describes healthy idle connections.

A clean N-hour cliff in an estate with a known timezone split should raise a clock hypothesis before an outage hypothesis. Millisecond precision is a second tell: WDC documents end in .000 (a reconstructed second-resolution timestamp), MAPLE's carry sub-second digits.

Fix — and it does NOT go where it first appears to. The obvious change is timezone => "America/New_York" on the date filter. That would do nothing (see (2)): the date filter's input field never exists, so it is a no-op. The shift is produced by the syslog input plugin, which does its own RFC3164 parse. The timezone belongs there:

input {
  syslog {
    port => 5140
    type => "syslog"
    timezone => "America/New_York"
  }
}

Constraint to record with the fix: this declares every sender on 5140 to be EDT. Today all of them are WDC hosts — MAPLE arrives on 5141/udp, not 5140 — but a future UTC sender to 5140 would then be misparsed by the same mechanism, in the opposite direction. The durable fix is for WDC rsyslog to emit RFC5424 with an explicit offset.

This fixes forward only. Documents already indexed keep their wrong timestamps; there is no reindex proposed here.

Verify by measurement, not by config read: after the change, a now-10m query on cloud-logs-* must return 192.168.120.1, .2 and .3 alongside 172.16.0.12. Absence of any one of the three is the failure mode — check them individually, since an aggregate hides a single silent host.

(2) grok fails on every WDC event, and the filter is redundant rather than merely broken

Every WDC document carries tags: ["_grokparsefailure", "%{syslog_host}", "%{syslog_program}"]. The literal interpolation strings are proof the fields do not exist — the mutate is writing placeholders.

The cause is that the message has already been parsed. The syslog {} input performs RFC3164 parsing itself and strips the header, so the custom grok is matching a header pattern against a body that no longer has one. It has never matched a WDC event.

The input already provides, in ECS form, everything the grok was reaching for — and more:

grok field already populated by the input
syslog_host host.hostname (sky)
syslog_program, syslog_pid process.name (named), process.pid (1273)
syslog_message message (body)
event.original (full raw line, retained)
log.syslog.severity, .facility, .priority

Nothing consumes the grok fields. Searched across the repository: syslog_host, syslog_program, syslog_message and syslog_timestamp appear in exactly one place — infrastructure/cloud-vms.md, which documents this same config. No portal, backend or dashboard reads them.

So the fix is removal, not repair. Delete the grok, the date and the mutate from the syslog filter block. That eliminates _grokparsefailure and the two placeholder tags on every WDC event, at zero cost, because the input already produces better fields. Removing a broken filter beats fixing one nothing depends on. cloud-vms.md:1537-1543 documents the old config and must be updated in the same change.

(3) WIND's rsyslog is SELinux-blocked, and WIND is missing from the cloud log store entirely

A repeating AVC denial on WIND:

avc: denied { name_connect } for pid=6871 comm=rs:main dest=5140
scontext=system_u:system_r:syslogd_t:s0
tcontext=system_u:object_r:unreserved_port_t:s0 permissive=0

rsyslogd cannot connect outbound to port 5140. WIND's own config forwards *.* @@172.16.0.13:5140, so WIND's logs have never reached CEDAR — confirmed independently: 192.168.120.4 does not appear in any cloud-logs-* aggregation. WIND is the one WDC host with no cloud-side log record.

WIND is both a collector and a forwarder, and neither role was in any inventory or diagram. It listens on 5140/tcp and 5140/udp with elasticsearch and logstash both active — it is the on-prem ELK stack, a second and older log destination predating CEDAR. SUN forwards to it as well as to CEDAR, and that leg is also failing (cannot connect to 192.168.120.4:5140: Connection timed out), which is a separate problem from the AVC.

Decide the architecture before applying a remedy. If the WIND→CEDAR leg is wanted, the standard fix is:

sudo semanage port -a -t syslogd_port_t -p tcp 5140

Checked, and it is safe here: WIND's Logstash runs as unconfined_service_t, so relabelling the port will not prevent it binding 5140 on restart. Do not set permissive mode. If the leg is not wanted, removing SUN's forward and WIND's forward is the fix, and no SELinux change is needed.

WIND enumerated 2026-09-09 — read-only, and it changes the decision

This section framed WIND as possibly orphaned and asked whether the leg is wanted. Enumerated from WIND's own side, not from sender configs:

Finding
Who sends SKY (2 connections), RAIN (3). SUN cannot — times out
What lands DNS, DHCP, auth syslog — nothing indexed since 2026-08-31
Why Elasticsearch at 999/1000 shards, every write rejected HTTP 400
Who reads Nobody. Kibana binds 127.0.0.1; 0 dashboards, 0 saved searches, 0 visualisations

WIND is collecting today and storing nothing. The shard ceiling is self-inflicted: 518 primaries plus 481 replicas a single-node cluster can never assign equals exactly 999.

Two corrections to this row. WIND is not orphaned, it is actively receiving. And its blocked forward targets CEDAR — the live production store, not a destination with no owner — so semanage would close this row's own coverage gap rather than restore unwanted traffic. Still not run, because the read question answers "nobody" and the shard problem is unresolved.

DHCP is the one thing not duplicated. RAIN's /etc/rsyslog.d/30-dhcpd.conf sends local7.* to a local file and to WIND, then & stop — so CEDAR has never received DHCP logs, confirmed across 24 hours of CEDAR syslog containing named and zero dhcpd. No data is lost: /var/log/dhcpd.log is current on both nameservers with daily rotation. What is dead is the only searchable, centralised copy.

Also settled: SKY's "unexplained" connection to WIND is /etc/rsyslog.d/50-forward-wind.conf using action(type="omfwd" ...), which contains no @ and so was invisible to the grep that declared it unexplained.

(4) host.ip is mapped as text, not as an IP type

Aggregating or range-querying by source address fails on host.ipFielddata is disabled on [host.ip] — and must use host.ip.keyword. event.original and message are likewise text with .keyword subfields, and both appear in _ignored, meaning values exceeded the keyword length limit and were dropped from the index for exact-match purposes. A range query such as "all sources in 192.168.120.0/23" is not expressible against a text field; it has to be enumerated term by term.

Acceptance test

  1. Structural: the syslog filter block contains no grok, date or mutate, and the syslog input carries an explicit timezone. Fails today.
  2. Live, and it closes the row: a now-10m query on cloud-logs-* returns all three WDC senders individually alongside MAPLE, and newly indexed WDC documents carry no _grokparsefailure tag. Both verified in the index, not in the config.

What this row does not claim

  • It does not claim data was lost. Every WDC event arrived and was indexed throughout. Only its @timestamp is wrong.
  • It does not propose reindexing the ~11,000 existing mis-dated documents.
  • It does not resolve whether the WIND leg is wanted (3). That is an architecture decision, not a defect finding.

APPLIED AND VERIFIED 2026-09-09 — parts (1) and (2) are closed

Applied to CEDAR at 14:10:47Z. The backup went to /root, not to conf.d: Logstash concatenates every file in that directory regardless of extension, so a .bak beside the config becomes live config — two syslog inputs on 5140 and two Elasticsearch outputs. Same trap as backups in /var/ossec/etc/rules.

--config.test_and_exit returned Configuration OK, which also validated 02-wazuh-alerts.conf. Logstash had been up since 2026-04-09, so this was the first restart in five months.

Verified against the index, not the config. The decisive evidence is a single post-fix document whose payload carries its own zone-qualified clock:

event.original : <14>Sep  9 10:11:30 sun grafana[1538]: ... t=2026-09-09T10:11:30.012301802-04:00
@timestamp     : 2026-09-09T14:11:30.000Z

10:11:30-04:00 is 14:11:30Z. The document agrees with the event's own embedded timestamp, which is a stronger test than "the timestamp looks current" — it cannot be satisfied by a clock that is merely close.

Check Result
WDC senders, now-20m, individually rain 172, sky 89, sun 83
_grokparsefailure in window 0
%{syslog_host} / %{syslog_program} tags 0 / 0
Documents still landing 4 h back 0
host.hostname, process.name, process.pid populated by the input
MAPLE 5141/udp feed across the restart 6,068,269 → 6,068,276 in 100 s
SOC portal wazuh panel ok: true, newest event current

Pre-fix the same now-10m query returned 0 documents from 5140. Existing documents keep their wrong timestamps; this was forward-only and nothing was reindexed.

The acceptance test in this row was wrong, and it would have failed the fix

The test above read: a now-10m query returns all three WDC senders individually. That criterion cannot pass, and it never could. Two separate reasons, both established today:

  1. SKY and RAIN do not log every ten minutes. Over 24 hours they produce 2,859 and 1,820 documents against SUN's 6,942, but in bursts. A ten-minute window routinely contains SUN alone. Their 24-hour presence is the honest liveness test; a ten-minute window is not.
  2. Their clocks are 15–16 minutes slow — see GOV-041. Their events land in the past even after this fix, so any window anchored on now and narrower than the skew excludes them by construction.

Both were mistaken for absence during verification. A per-host logger marker pushed through the real path settled it, and is the test that should have been written: it is independent of volume and of clock skew.

Corrected acceptance test, and it is met: for each of 192.168.120.1, .2 and .3, a marker emitted with logger on the host appears in cloud-logs-* with the correct host.hostname and process.name and no _grokparsefailure; and a 20-minute window returns all three individually.

One recorded constraint is false, and it is a latent defect

The fix comment and this row both state that declaring every 5140 sender to be EDT is safe because MAPLE arrives on 5141/udp and not on 5140. MAPLE is configured to send to 5140. /etc/rsyslog.conf:6 on MAPLE reads:

*.* @@172.16.0.13:5140

It is simply not connected — CEDAR shows established peers on 5140 from 192.168.120.1, .2 and .3 only, and MAPLE has no outbound socket to 5140 at all. So the premise holds by accident, not by design. If that forward ever succeeds, MAPLE runs UTC and its events will be parsed as EDT and indexed four hours in the future — the same defect in the opposite direction, and harder to spot because future-dated documents appear at the top of every descending sort.

Two consequences, recorded rather than fixed here:

  • MAPLE's own OS syslog is absent from cloud-logs-*. The only MAPLE documents present are Wazuh alerts arriving on 5141 (see GOV-042). This was read on 2026-09-08 as "MAPLE is current" and taken as evidence the pipeline was healthy for MAPLE. It is not the same claim.
  • The durable fix remains RFC5424 with an explicit offset from every sender, which removes the collector-zone assumption instead of re-pointing it.

GOV-040 — BIND on SKY and RAIN is not chrooted, and Puppet manages the unit that is not running

Raised 2026-09-08. Measured on both DNS servers:

Host named.service named-chroot.service
SKY active inactive, disabled
RAIN active inactive, disabled

BIND is not chrooted on either authoritative DNS server. Both run the plain distro unit.

The consequence is that a body of Puppet work manages a unit that has never run on these hosts — including the duplicated PIDFile=PIDFile= in the named-chroot unit and the ExecReload change, held at commit d088b4bf. Fixing a unit file that is disabled changes nothing about the running service, and a --noop that reports the fix applying cleanly would say nothing about whether DNS is affected.

Same shape as the estate's recurring failure: a control whose artifacts are all in order, managing something that is not there. Cf. VLN-010, where the rules loaded, the count was correct, and nothing was ever evaluated.

Not assessed here: whether chroot is wanted. If it is, the gap is that a documented hardening control is absent on both nameservers. If it is not, the gap is that Puppet carries and maintains manifests for a unit the estate has decided against. Either way the current state — maintained manifests for a disabled unit — is the wrong end state.

Acceptance test: a stated decision on whether BIND runs chrooted, and the running unit on both hosts matching it, verified with systemctl is-active and is-enabled rather than from the manifest.

Scope question answered 2026-09-09 — Puppet does not manage SKY or RAIN

This row asks whether d088b4bf's named-chroot management ever applied to the nameservers. It cannot have. Enumerated on both hosts: no puppet package in the rpmdb, no /opt/puppetlabs, /etc/puppetlabs, /etc/puppet, /var/lib/puppet or /opt/puppet, no puppet binary, all of puppet, puppet-agent and pxp-agent inactive, and no cron or timer referencing a converge. No file-distribution path from phoenix is visible from the host side either.

So the unit is not "managed but never run" on these hosts — it is not managed at all, and the register should stop treating a Puppet change as the route to fixing chroot on SKY and RAIN. Whatever configures these two hosts, it is not the gpus-dist flow. Establishing what does configure them is the open question this leaves.

Checked because the chrony fix in GOV-041 would have been silently reverted by a converge if they were managed. They are not, so it stands.

GOV-041 — SKY and RAIN have no time synchronisation, and both nameservers are running 15 minutes slow

Raised 2026-09-09. Found while verifying GOV-039, not from a failure or an alert. A logger marker pushed from each WDC host arrived correctly parsed but stamped roughly sixteen minutes in the past. The collector was not at fault; the sources are.

Measured against CEDAR at 14:18:00Z:

Host Reports UTC Skew System clock synchronized NTP service
sky (192.168.120.1) 14:02:06 −16 min no inactive
rain (192.168.120.2) 14:02:51 −15 min no inactive
sun (192.168.120.3) 14:18:44 +0 yes active
wind (192.168.120.4) 14:18:52 +0 yes active
maple (172.16.0.12) 14:18:56 +0 yes active
cedar (172.16.0.13) 14:19:01 +0 yes active

The split is exact: only the two nameservers. Both carry the correct timezone (America/New_York), so this is not the GOV-039 defect recurring — it is free-running hardware clocks with no discipline at all. Drift of this size is not a one-off excursion; it accumulates, and nothing bounds it.

The evidence was already in the record and was read as something else. The 2026-09-08 measurement published in GOV-039 shows, at 17:53:55Z:

Source last @timestamp lag behind SUN
SUN 13:53:52
SKY 13:36:23 17 min
RAIN 13:32:31 21 min

All three were attributed to the same four-hour timezone shift, and the residual 17 and 21 minutes were read as low log volume. They are the same defect being recorded here, visible a day earlier in data already collected. A uniform offset explains the four hours; it does not explain the extra seventeen minutes, and the leftover was not interrogated.

Why it matters beyond log ordering. These are the primary and secondary DNS/DHCP servers, and they are the two hosts that sign wdc.us.gl3:

  • DNSSEC signature validity is wall-clock bounded. Inception and expiry are absolute times. A signer whose clock is wrong emits signatures whose validity window is wrong by the same amount, and a resolver checking them uses its own clock. This compounds an already-fragile posture — the zone is an island of trust with no DS in the parent, and re-signing is manual.
  • DHCP lease times and log correlation across the estate are off by a quarter hour, in the direction that makes WDC events look older than cloud events for the same incident.
  • Any future Kerberos or AD integration fails at five minutes of skew, and this is three times that.

RESOLVED 2026-09-09 — and the root cause is not what this row assumed

Fixed on both hosts. The cause was a single invalid word in /etc/chrony.conf, not drift, not the power event, and not a missing time source. Line 5 read denyall; chrony 4.5 has no such directive, the syntax is deny all as two words. chronyd therefore exited 1 on every start attempt with Invalid directive at line 5 in file /etc/chrony.conf.

The row's framing above — "free-running hardware clocks with no discipline" — describes the symptom, and the sentence proposing to "enable chronyd and verify against a source they can actually reach" would not have worked. The daemon was already enabled. Starting it fails until the word is fixed. The source was reachable the whole time.

It has never run on these hosts. The package and config timestamps place the error at build time, weeks before the May reboot:

Host chrony-4.5 installed chrony.conf written Gap
SKY 2026-02-26 09:36:36 2026-02-26 10:05:13 29 min later
RAIN 2026-02-27 11:19:19 2026-02-27 11:25:32 6 min later

Only one chrony version has ever been in either rpmdb, so denyall was not a formerly-valid directive invalidated by an upgrade — that hypothesis was tested against dnf history and refused. The file was written onto an already-installed chrony 4.5 and was wrong the moment it was saved. Both files were byte-identical at 244 bytes, written a day apart, so it was one paired build step repeated on the second host.

So what corrected the clocks until May? Both hosts are VMware guests with vmtoolsd active, and VMware Tools syncs guest time at power-on. The 121-day drift matches the uptime since the 2026-05-11 boot exactly. The Feb–May interval drifted too and was silently corrected by that same reboot. The only mechanism that has ever set these clocks is a power cycle. The estate had no time discipline on its nameservers at any point since they were built; it had a reboot that happened often enough to hide it.

Applied and verified, RAIN first so SKY stayed authoritative throughout:

RAIN SKY
chronyd failed since 2026-05-11 13:15:52 EDT 2026-05-11 13:14:42 EDT
Offset before 916 s slow 961 s slow
Step 14:43:43Z → 14:58:59Z 14:51:03Z → 15:07:04Z
After 0.0000011 s off NTP 0.000254 s off NTP
Measured drift 59.132 ppm 57.826 ppm
named restarted no, PID preserved no, PID preserved

The correction was forward, because the clocks were slow. That is the benign direction for a signer: inception earlier than it should be is accepted, whereas a fast signer emits not-yet-valid signatures. named was not restarted on either host — ActiveEnterTimestamp still reads 2026-05-11 on both. The zone answers NOERROR from both nameservers, serial 2026090901, RRSIG inception 20260909060801 and expiry 20261009060801 unchanged and still valid against corrected time, so no re-signing was required. AIDE baselines promoted unconditionally on both, with a separate mv and never && mv.

SKY's step crossed the 15:00/15:01 cron minutes and RAIN's did not. Checked rather than assumed: no CROND, anacron or run-parts activity in the journal, /var/spool/anacron/cron.weekly mtime unchanged at 03:08:12, serial and RRSIG unchanged. Nothing re-ran.

One consequence to carry to the next re-sign. This morning's 03:08 weekly run signed the zone on the 16-minute-slow clock, so today's signatures carry an inception 16 minutes earlier than a correct clock would have produced. Benign, no action; the run on or about 2026-09-16 will be the first signed against true time.

Two things this fix did not resolve, both now dependencies:

  • No internal NTP server exists in the estate. MAPLE, CEDAR and SUN all time out on udp/123; the only reachable source is the public rhel.pool.ntp.org, across the VPN. So DNS now depends on the tunnel for time, and DNSSEC validity is wall-clock bounded. If the tunnel is down long enough the nameservers free-run again at ~5 s/day with nothing to notice. Recorded as a finding; an internal source is proposed, not built.
  • Nothing alerts on a dead time daemon. chronyd was enabled and failed for the entire life of these hosts and no check anywhere reported it. timedatectl says NTP service: inactive, which reads as "not configured" rather than "configured and broken" — the wording actively conceals the state.

Also observed, not a time defect. dnf-automatic-install.timer is active on both nameservers, applying unattended package upgrades daily at about 06:0x. It explains the unrelated files in today's AIDE report (/usr/share/man/man1/wget.1.gz). Unattended upgrades on the primary and secondary DNS may warrant a row of its own.

Puppet does not manage these hosts, so the fix cannot be reverted by a converge — see the amendment to GOV-040.

Acceptance test

Both met 2026-09-09.

  1. timedatectl on both hosts reports System clock synchronized: yes and an active NTP service. Add a third clause the original test lacked: systemctl is-active chronyd must read active, because enabled and failed was the actual state and neither timedatectl nor systemctl --failed phrasing distinguishes it from "never configured".
  2. A logger marker emitted on each host appears in cloud-logs-* with an @timestamp within 5 seconds of the emitting shell's own UTC clock, read from CEDAR in the same command. Measured, not inferred from timedatectl.

What this row does not claim

  • It does not claim any log was lost. Ordering and correlation are wrong; delivery is not affected.
  • It does not claim DNSSEC validation is currently failing. The zone is an island with no validating path from the parent, so the practical impact today is bounded. The exposure is recorded because drift is unbounded and the signing is manual.
  • It does not establish when the drift began. Nothing on either host records the clock's history, and the 2026-09-08 figures are the earliest measurement available.

GOV-042 — the documented Wazuh pipeline on CEDAR does not exist, and what runs instead duplicates the alert corpus into world-readable /tmp

Raised 2026-09-09. Found while validating the GOV-039 change, because --config.test_and_exit validates every file in conf.d and the second file had to be read to be sure the test result belonged to our change.

cloud-vms.md §11.2.1 documents a pipeline that is not deployed. The documented config and the live file share only the port number:

Documented (§11.2.1) Live on CEDAR
Input beats, type => "wazuh" udp, no type
Filter json + date on timestamp none
Output elasticsearch, wazuh-alerts-%{+YYYY.MM.dd}, guarded by if [type] == "wazuh" file, /tmp/wazuh-debug.log, unguarded

The live file is nine lines and reads as a debugging artefact left running in production.

The pipeline is single, and neither output is guarded. pipelines.yml defines one pipeline over /etc/logstash/conf.d/*.conf, so both files concatenate and every input reaches every output. Two consequences, both confirmed by measurement rather than by reading the config:

  1. Wazuh alerts are indexed into cloud-logs-*, unparsed. Today's index holds 7,863 documents with type: syslog and 2,019 without, and every untyped one has host.ip: 172.16.0.12. Their message is the entire raw line — <132>Sep 9 14:22:10 maple ossec: {…} — with the whole alert JSON as an uninterpreted string, and their @timestamp is Logstash's ingest time, not the event time.
  2. WDC syslog is written to /tmp/wazuh-debug.log. The file output is not restricted to the Wazuh input, so the debug file receives the WDC corpus as well as the Wazuh one. The name understates what it holds.

This is a second unauthenticated copy of the alert corpus, by a different mechanism than VLN-053. VLN-053 records the network exposure of CEDAR's indexer on 9200. This is the local one: a continuous, unrotated, world-readable write of alert content into /tmp, a directory subject to systemd-tmpfiles cleanup, on the host that holds the log store. Same data, second exposure. It should be carried as a security finding alongside VLN-053, not only here.

The duplication buys nothing, and yesterday's reasoning about it was wrong. It was recorded on 2026-09-08 that removing this file would have destroyed MAPLE's Wazuh feed. It would not have. The production path to wazuh-alerts-* does not pass through Logstash at all: Filebeat on MAPLE holds an established connection to 172.16.0.13:9200 and writes there directly. wazuh-alerts-* is current — newest document 14:21:01Z — and soc-backend/app.py queries wazuh-alerts-* at five call sites and cloud-logs-* at none. The portal has never read the 5141 copy.

That does not make yesterday's caution wrong as a decision — the feed's consumers had not been enumerated, and declining to remove an unexamined input on the SIEM was correct. It makes the stated reason wrong, and the reason is what would have been carried forward.

Acceptance test

  1. 02-wazuh-alerts.conf either matches cloud-vms.md §11.2.1 or the section is rewritten to match what is deployed. The two must not continue to disagree, and the direction is a decision, not a defect fix.
  2. No Logstash output writes alert content to a world-readable path.
  3. If the 5141 input is retained, both outputs carry an explicit conditional, so the shared pipeline stops cross-feeding.
  4. A count of untyped documents in cloud-logs-* over a 24-hour window is zero, or the duplication is documented as intentional with a named consumer.

What this row does not claim

  • It does not claim Wazuh alerting is impaired. The production path via Filebeat is healthy and the SOC portal renders from it.
  • It does not claim /tmp/wazuh-debug.log has been read by anyone. No access record was examined; the finding is the exposure, not an incident.
  • It does not propose which way to reconcile the documentation. Deleting the stray input and deleting §11.2.1 is one answer; deploying §11.2.1 as written is another. Both need the consumer question settled first.

GOV-043 — every WDC host runs a volatile journal, and the check that found it reports healthy on three of the four

Raised 2026-09-09. Found while recovering the reason chronyd failed in GOV-041. The reason was unrecoverable, and why it was unrecoverable is the larger finding.

systemd-journald runs Storage=auto on all four WDC hosts, which persists to disk only if /var/log/journal exists. It exists on none of them:

Host systemd-journal-flush /var/log/journal Volatile journal in RAM
SKY active absent 384 M
RAIN failed absent 385 M
SUN active absent 361 M
WIND active absent 385 M

Four of four. No WDC host has ever persisted a journal, and every reboot discards the entire local log history.

The unit state is the wrong signal, and it points at the wrong host

systemctl --failed surfaces RAIN and nothing else, which invites the reading that RAIN is broken and the other three are fine. They are not. With Storage=auto and no /var/log/journal, flushing is a successful no-op, so SKY, SUN and WIND report active because nothing persists. The three hosts that look healthy are in exactly the same state as the one that looks broken.

This is the recurring error of the week in a new place: a check that passes for a reason unrelated to what it is believed to test. Companion to M-13. The correct probe is test -d /var/log/journal, not the unit's state.

RAIN's own failure is separately unexplained and now unexplainable. It is recorded as failed since 2026-05-12 09:05:24 EDT, a day after the reboot rather than at boot, with Main PID: 1046 (code=exited, status=0/SUCCESS) — a successful main process on a failed unit. The volatile journal reaches back only to 2026-08-26, so the May evidence is gone, not hidden. Recorded as unrecoverable rather than guessed at.

Retention is about two weeks, and that is the real exposure

/run/log/journal is capped by RuntimeMaxUse and is currently holding roughly two weeks per host — 2026-08-26 to 2026-09-09 on RAIN. So the local record is fourteen days, volatile, and lost on reboot. There is no long-term local log on any WDC host.

That places the entire durable record of these four hosts in cloud-logs-* on CEDAR — the same store whose WDC timestamps were four hours wrong until GOV-039 was fixed today, whose senders are unauthenticated, and whose alert corpus is duplicated to world-readable /tmp under GOV-042. The three rows are not independent: the central store is the only copy, and it was not trustworthy.

Not fixed, and it should not be fixed blind

Creating /var/log/journal and running journalctl --flush would enable persistence in one step, and all four hosts have room — a dedicated /var/log volume with 25–28 GB free. It is deliberately not applied here:

  • It changes disk-write behaviour on the primary and secondary DNS servers.
  • SystemMaxUse should be set explicitly rather than inheriting the 10 % default, on hosts whose /var/log is shared with BIND query logging.
  • RAIN's unexplained failure should be understood before its cause is potentially recreated on three more hosts.

Acceptance test

  1. test -d /var/log/journal succeeds on all four WDC hosts. This, not the unit's state, is the check.
  2. journalctl --list-boots returns more than one boot on each host — proof of persistence across a reboot, which the current configuration cannot produce.
  3. SystemMaxUse is set explicitly in journald.conf on each host.

What this row does not claim

  • It does not claim any log reached no destination. WDC syslog forwards to CEDAR independently of journald's storage mode, and that path is working. What is lost is the local record and everything from before the last boot.
  • It does not explain RAIN's failure. The evidence expired before it was looked for.
  • It does not establish when /var/log/journal was last present, if ever. Nothing on the hosts records that, and their journals cannot say.

GOV-044 — the documented post-change AIDE step fires ten level-12 credential-dumping alerts, and costs 20 points of SOC score

Raised 2026-09-09. Found while confirming that disabling Webmin on MAPLE had not disturbed the SIEM. It had not; the SOC score had moved anyway, from 75 to 55, and the cause was our own change procedure.

Ten alerts of rule 100015, LOLBin: /etc/shadow access attempt — T1003 OS Credential Dumping, level 12, on SKY and RAIN within five seconds of each other. Every one is AIDE. The audit record names it without ambiguity:

comm="aide" exe="/usr/sbin/aide" AUID="dnsadmin" tty=pts3
syscall=195 (llistxattr) key="credential_access"
file: /etc/shadow

llistxattr lists extended attributes. AIDE did not read /etc/shadow; a file-integrity tool inspecting the file it exists to inspect is scored as credential dumping.

It fires on --update, not on --check. Across 14 days rule 100015 appears on exactly two days — 2026-09-02 (5) and 2026-09-09 (10) — and both are days an operator ran aide --update. The daily aide-check cron does not trigger it.

So Step 5 of the change procedure generates critical alerts by design. An operator following the documented checklist correctly produces a burst of level-12 credential-access findings and drops the posture score by 20 points. The five on 2026-09-02 went unremarked at the time, which is the more useful half of the finding: the alerts are already being ignored, and a rule that is routinely ignored cannot detect the thing it was written for.

Same class as the 100010 curl exclusion. The remedy is an exclusion scoped to exe="/usr/sbin/aide"not a blanket suppression of 100015, and not raising the threshold.

Acceptance test

  1. An aide --update on SKY or RAIN produces zero 100015 alerts.
  2. A non-AIDE process touching /etc/shadow still produces one. Prove the second clause deliberately; an exclusion that silences the rule entirely passes clause 1 and defeats the control.

What this row does not claim

  • It does not claim the rule is wrong. /etc/shadow access is worth alerting on. The defect is the missing exclusion for a known, scheduled, authorised reader.
  • It does not claim the score is meaningful. Whether a 20-point swing from ten alerts is the right weighting is a separate question about soc_score, not addressed here.
  • It does not establish how far back this goes. wazuh-alerts-* was queried over 14 days only.

GOV-045 — DHCP was absent from the central log store entirely, and the obvious one-line fix would have broken two other logs

Raised and FIXED 2026-09-09. Found while enumerating WIND. CEDAR held zero dhcpd documents across 24 hours while carrying 1,435 from named, so the estate's central store had DNS but no DHCP at all.

The two nameservers failed differently, and SKY was worse:

Host local7 went to Reached CEDAR
RAIN /var/log/dhcpd.log, then WIND, then & stop no
SKY /var/log/dhcpd.log, then & stop no — and not WIND either

SKY's DHCP reached no destination beyond the local file. RAIN's reached WIND, which has indexed nothing since 2026-08-31 (GOV-039 (3)). So the only searchable copy of DHCP lease history in the estate was a dead index, and half of it was never even sent there.

The & stop is load-bearing, and deleting it was the tempting wrong fix

rsyslog.conf includes /etc/rsyslog.d/*.conf at line 37 and its own default rules come after, at lines 47 and 66:

47: *.info;mail.none;authpriv.none;cron.none   /var/log/messages
66: local7.*                                   /var/log/boot.log

DHCP uses local7, which is the facility RHEL reserves for boot logging. Removing the stop would have pushed every DHCPDISCOVER and DHCPOFFER into /var/log/messages and /var/log/boot.log, unbounded. A one-character deletion described as a stray directive would have created two new problems while fixing one.

The fix is to add, not remove: insert the CEDAR forward above the stop, so local7 still terminates before reaching the default rules.

local7.*    /var/log/dhcpd.log
local7.* @@172.16.0.13:5140
& stop

Applied to both hosts with rsyslogd -N1 validation before each restart, backups to /root and not into the *.conf load path.

Verified by measurement, against a 24-hour baseline of zero:

Measure Result
dhcpd documents, total 114 → 162, +48 in 120 s
Last 45 min, per host rain 78, sky 30
Document shape process.name: dhcpd, facility: local7, current @timestamp, no _grokparsefailure

One documentation defect found alongside it

RAIN's 30-dhcpd.conf line 3 reads "Forward to Logstash on SKY (CIS 8.9 — centralized logging)" above an address of 192.168.120.4, which is WIND, not SKY. The comment has misnamed the destination for as long as it has existed, which is a plausible route by which this drifted unnoticed.

Acceptance test

  1. dhcpd documents from both 192.168.120.1 and 192.168.120.2 appear in cloud-logs-* over a window wide enough to absorb burstiness, and the count moves on a second read. Met.
  2. /var/log/messages and /var/log/boot.log on both hosts contain no DHCPDISCOVER/DHCPOFFER lines after the change — the clause that distinguishes this fix from the one that was almost made. Met 2026-09-09, all four counts zero:
sky   /var/log/messages:0   /var/log/boot.log:0
rain  /var/log/messages:0   /var/log/boot.log:0

Both clauses now pass, and they had to be tested separately. Clause 1 alone is satisfied by simply deleting the & stop, which was the tempting fix and would have flooded both files. An acceptance test that only asserts the presence of the thing you wanted cannot tell a fix from a regression that happens to include it. The clause asserting an absence is what makes the pair diagnostic.

What this row does not claim

  • It does not claim DHCP history was lost. /var/log/dhcpd.log is current on both hosts with daily rotation back through August. What was missing was the central, searchable copy.
  • It does not address WIND. RAIN still forwards DHCP there, into a cluster that rejects every write. That is GOV-039 (3) and an architecture question.

Method learnings

Constraints on future controls, not findings. These are deliberately unnumbered, status-less and unowned — they carry no GOV- id, no acceptance test and no date_verified, because there is nothing here to close. They are the conditions any future reconciliation control, enumerator or coverage gate has to satisfy to be trustworthy. Recorded 2026-08-17 from enumeration Stages 1, 1b, 1c, 2, 3 and 3b.

M-01 — LOADED ≠ REACHABLE

Three instances in this program: VLN-010's Wazuh rules, rooster's commented-out class (VLN-034), and the transfer crons (PRG-019) firing on schedule doing nothing. This is the estate's characteristic failure — a control present, scheduled and running while doing nothing.

M-02 — Permission-denied is not unreachability

An SSH timeout is a gap to route around; a 403 means the enumerator is blind and must halt. Recording a denial as UNREACHABLE produces a file that looks complete because it ran to the end.

M-03 — Exit codes carry no information for gcloud list verbs

All six access probes returned exit 0, denials included, with the error only on stderr. stderr capture is the sole discriminator. (Carried as VLN-017.)

M-04 — --quiet is required on read-only gcloud calls, and forbidden on mutating ones

GCP returns PERMISSION_DENIED for a disabled API as well as for an IAM denial, and gcloud converts that into an enable-and-retry prompt — a project mutation offered inside a read stage. On mutating verbs, --quiet auto-confirms.

M-05 — API_DISABLED is not an IAM denial

An IAM denial hides resources that may exist; a disabled API means that resource type has no enumeration path on that project. Two different unknowns that must not collapse into one.

M-06 — Recorded absence, never silent absence

A blocked project gets an explicit BLOCKED row with the probe result and date.

M-07 — Confidence and enforcement are different axes

Reading a manifest tells you what is declared, not whether it ever applied. (Carried as GOV-011.)

M-08 — Counts supplied from memory rather than a live read are provisional

Four were overturned during these sessions: the orphan count twice (GOV-018), the gpusa UNKNOWN count (PRG-022), and the unlisted Cloud Run service count (corrected on PRG-003, 2026-08-13). Every count in a register needs a source.

M-09 — Enumerate fully before classifying

The moment classification starts, the list looks finished and enumeration quietly stops.

M-10 — UNKNOWN must be the default and must survive to the end

The honest figure at Stage 2 was 50.0%. Every prior artifact implied near-complete knowledge.

GOV-022 — inventory.yaml cannot express a capability that spans components

Field Value
id GOV-022
title cloud_services has no part_of: / capability: field and the coverage check has no concept of a capability, so a system spanning several resources is registered only as loose siblings under a YAML comment. Nothing machine-readable ties them together.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Two multi-resource capabilities are now registered this way. The 2.5(c) routing + ClamAV scan pipeline is five cloud_services entries grouped only by a # ── comment header and a shared documented_in. The travel approval workflow (registered 2026-08-26) is four Pub/Sub entries plus desc extensions on three pre-existing entries — gpus_forms_backend, gpus_forms_routing_worker, gpus_forms_db — with nothing linking the seven. REFERENCE_FIELDS in check-component-coverage.py is ("powers", "powered_by", "fed_from", "hosted_on", "vms"); none expresses membership, and hosted_on expresses only host residency. This is the same structural gap the docs portal's Platform Services section was created to solve — that section's own preamble states the estate is host-organised and the approval workflow does not fit — solved for documentation, unsolved for inventory.
acceptance test A query over inventory.yaml alone can answer "which components make up the travel approval workflow?" and return all seven, without reading a comment or a doc. Whatever field is added must be validated — an unvalidated free-form key would read as structure while providing none, which is precisely why one was not added when the approval entries were written.
blocks
blocked_by
date_raised 2026-08-26
date_verified

Deliberately not fixed in passing

Raised while registering the approval workflow, and not solved there, on the stated ground that adding part_of: as a free-form key would satisfy nothing: REFERENCE_FIELDS would not check it, no validator would read it, and it would read to the next contributor as structure that exists. The schema doc, a validator, and a backfill of the five 2.5(c) entries are the real work. Related: VLN-048, which is the same theme in the enforcement layer rather than the schema.

GOV-023 — Live GCP resources that no inventory entry covers

Field Value
id GOV-023
title Enumerating live Pub/Sub while registering the approval workflow found resources present in gpus-infra and absent from inventory.yaml. The registration standard says every managed asset is in the inventory; these are not, and nothing detects it, because coverage is checked inventory→docs and never GCP→inventory.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence gcloud pubsub topics list --project=gpus-infra on 2026-08-26 returned six topics. After the approval registration, two remain unregistered: gpus-forms-submission-ready and gpus-forms-submission-ready-dlq — the 2.5(c) routing pipeline's own transport, named in the routing worker's boot line (topic=gpus-forms-submission-ready). Subscriptions are worse: of five live, gpus-forms-routing-worker-sub and gpus-forms-routing-dlq-alert-sub are unregistered, and both are named in gpus_forms_routing_worker's own desc. Separately, the gpus-reports cron fleet on MAPLE — six scheduled report jobs including the live R2 travel-approval report — has no inventory entry of any kind.
acceptance test A re-runnable reconciliation, GCP→inventory, over at least Pub/Sub topics and subscriptions in gpus-infra, reporting every live resource with no cloud_services entry. Must be able to fail: running it on the 2026-08-26 tree must report the four Pub/Sub resources named above. Direction matters — the existing check walks inventory and asks whether docs mention it, which by construction cannot see a resource that was never registered.
blocks
blocked_by
date_raised 2026-08-26
date_verified

Why the existing gate cannot find these

check-component-coverage.py iterates inventory.yaml and validates outward — schema, references, documented_in, servers.py, portal presence. Every one of those starts from an inventory entry. An asset that was never registered has no entry to start from and is therefore invisible to all seven validators simultaneously. The check answers "is what we wrote down consistent?" and never "is what we wrote down complete?" — and the standard it enforces claims the latter.

GOV-024 — Two Finance sources disagree on what the cost centre list is

Field Value
id GOV-024
title The 31-value cost centre list loaded into travel-request-001 v3 comes from Finance's "Greenpeace Inc - Accounts" export. A project spreadsheet Finance sent earlier carries a different set. Neither is marked authoritative, and the portal now serves one of them to every traveller.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned — needs a Finance owner, not an IT one
evidence Recorded at the point of loading, 2026-08-27. Present in the project spreadsheet and absent from the Accounts export now in the portal — 12000, 12010, 12100, 12160, 20610, 20620, 39013, and others. Present in the Accounts export and absent from the project spreadsheet — 13000 (OTHER ACQUISITION). The disagreement runs in both directions, so neither list is a subset of the other and "use the longer one" does not resolve it.
acceptance test A named Finance owner states which export is authoritative for staff-facing cost-centre selection, and the portal's Travel Cost Center pulldown matches it — verified by comparing the live pulldowns row against that source, not against this document.
blocks Nothing today. It will block any reconciliation of travel spend against the general ledger, because a submitter can only pick from the list the portal offers.
blocked_by An answer from Finance.
date_raised 2026-08-27
date_verified

What was NOT done, and why

The Accounts export was loaded verbatim, including the three mixed-case entries (Major Gifts and Foundations - General Expense Only, Content Core, Democracy Core), because the portal's stored value should match Finance's record rather than a normalised guess at it. The codes from the project spreadsheet were not merged in, and the casing was not normalised. Both would have produced a list that matches no Finance source at all — a third vocabulary, created by IT, that neither side recognises.

The cost of leaving it is real and should not be understated. The stored value lands in searchable_values and is what R1 would group by, so every submission filed before this is answered is grouped by a list of contested provenance. That is recoverable — the code prefix makes a remap mechanical — but it is recoverable work, not free.

GOV-025 — One org unit, two names, both served by the same form

Field Value
id GOV-025
title travel-request-001 now serves two different names for the same organisational unit on one page. The Travel Department pulldown offers People & Culture; the Travel Cost Center pulldown offers 20200 — HUMAN RESOURCES. A traveller in that unit picks one label at field 2 and the other at field 5.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Confirmed in the repo 2026-08-27. forms/pulldowns.yamlTravel Department contains People & Culture (the hr → People & Culture rename has already landed there). The same file's new Travel Cost Center, loaded verbatim from Finance's Accounts export, contains 20200 — HUMAN RESOURCES. Both are served on the same form, three fields apart.
acceptance test Either the two vocabularies agree, or the form states which is which. The test must be direction-neutral — it fails on any org unit whose Travel Department label has no corresponding Travel Cost Center name and vice versa, not merely on this one pair, since the rename programme will produce more of them.
blocks
blocked_by GOV-024 in practice — there is no point reconciling names against a list whose membership is disputed.
date_raised 2026-08-27
date_verified

Deliberately not reconciled in the load

Renaming HUMAN RESOURCES to People & Culture inside the cost centre pulldown would have made the form internally consistent and made the portal's value stop matching Finance's ledger, which is the one thing the cost centre value exists to do. The mismatch is the honest state and it belongs to the rename programme, not to a form load.

Worth noting which direction is safe when this is fixed: the department label is GPUS-facing and can be renamed freely; the cost centre string is a key into someone else's system and cannot.

M-12 — A staging directory that persists makes "redeploy" indistinguishable from "reinstall the same thing"

Observed 2026-08-27 on the routing worker's bc0c84b77da2ab deploy, and it is M-01's shape in the deploy path rather than the run path.

The MAPLE procedure stages to /tmp/routing-worker, then sudo cps from it. That directory survives between deploys. A re-run that starts in a shell already sitting there — which is what happens when someone returns to a half-finished deploy — skips the gsutil cp and copies last week's files back over themselves.

Every step reported success. The copy succeeded, restorecon succeeded, the restart succeeded, the unit came up active, both drain heartbeats logged RUNNING, and the installed file's mtime updated. Nothing changed. The deployed module still had zero occurrences of the symbol the deploy existed to ship.

One line revealed it: boot version= still reading the old hash. The VERSION drift guard — which exists precisely because "MAPLE is not a git client" — was the only signal that separated a successful deploy from a successful no-op, and it worked.

The constraint on future deploy procedures:

  • Verify the payload before installing it, not only after. cat VERSION and an mtime check cost nothing and sit before the first sudo. Added to the worker README.
  • A procedure that reads from a persistent scratch location must prove the location's contents, or stage to a fresh path per deploy. Idempotent-looking commands over stale inputs are silently idempotent in the wrong direction.
  • Health signals cannot detect this class. is-active, heartbeats and a clean journal all report on the process that is running, and the wrong process runs perfectly. Only an identity assertion — a hash the artifact carries and the log states — distinguishes them.

M-11 — Findings come from checking what looks correct, and from colleagues using the thing

Not one of this program's recent findings arrived because something broke. Recorded 2026-08-27 because it contradicts the model most monitoring is built on, and therefore changes what is worth building next.

Session of 2026-08-26/27 — seven findings, zero from a failure.

Finding Arrived from
VLN-047 A colleague asking how to revise a returned request
VLN-048 Reading the coverage gate while using it — it was passing
VLN-049 Reading the form's YAML for an unrelated field change
GOV-022 Registering an entity and noticing the schema could not express it
GOV-023 Enumerating live GCP and comparing it to the inventory
GOV-024 Comparing two Finance sources that were each internally consistent
GOV-025 Loading two pulldowns and reading them side by side

Every one was found by looking at something that was working. Nothing alerted. No build failed for the reason that mattered — the one build that did fail, failed for an unrelated network timeout (VLN-039). Two of the seven (VLN-048, VLN-049) were protected by green checks at the moment they were found: a coverage gate that passes every entity unconditionally, and a unit test requiring copy that had become false.

Three of the last four security findings came from a person, not a check. VLN-045 from a report about "the approver in the other forms"; VLN-046 confirmed by ticket #USFIN00367708; VLN-047 from Tanu Garg's report. The estate runs Wazuh, Prometheus, OpenVAS, Lynis, AIDE, fail2ban and a coverage gate. None of them found any of the three, and none of them could have: all three were correctness and truthfulness defects in what the portal told a staff member, and no monitoring system in this estate has an opinion about whether a sentence is true.

The constraint this puts on future controls — the reason it is here and not in a retrospective:

  • Do not size verification work by incident count. A component with no alerts and no failures has produced no evidence about itself. VLN-048 sat green for three weeks over two unrendered production services.
  • A green check is a claim, and claims need testing. Both VLN-048 and VLN-049 are checks that passed while asserting something untrue. Any new gate needs a demonstration that it can FAIL on a real defect — the acceptance test written for VLN-048 is that shape, and is the pattern to copy.
  • Treat user reports as the highest-yield detector available, and instrument accordingly. Three findings in four days came through a colleague noticing something odd. That is not a fallback for missing monitoring; on this evidence it is the most productive source the estate has for this defect class, and it has no intake path, no triage, and no way to be counted.
  • The corollary, stated so it is not read as complacency: this is a claim about where findings came from, not about where defects are. A defect nobody looked at and nobody tripped over produces no finding at all, and nothing here gives any information about how many of those exist.

M-13 — A signal that would be true either way is not evidence the change landed

Observed 2026-08-27, on the same routing-worker deploy that produced M-12. M-12 is the cause — a staging directory that let a redeploy reinstall the previous build. This is the detection failure, and it is the more portable of the two.

Four gates were checked after the 00:19:55 restart. Three were green, and all three were meaningless.

Gate Read Would it differ if the deploy had not happened?
approval_recipient_override=OFF OFF — correct No. Env file, unchanged by any deploy
Approval-drain heartbeat RUNNING No. Present in the old build too
DLQ-drain heartbeat RUNNING at iteration 1 No. Same
version= bc0c84b — the OLD hash Yes. The only one

Each green reading was a true statement about the running process. The running process was bc0c84b. Three true statements, zero information: every one of them describes a property the previous build also had, so observing them cannot distinguish "the new code is running" from "the new code was never installed". Only version= is derived from the artifact that was supposed to change.

The constraint on every future deploy check:

A signal that would read the same whether or not the change landed is not evidence that the change landed. Every deploy verification needs at least one assertion tied to the artifact's identity — a commit hash, a content digest, or a symbol that exists only in the new code and is asserted present.

Corollaries worth stating, because each was nearly missed here:

  • Liveness is not identity. is-active, heartbeats, a clean journal and an absent-error check all report on a process running correctly. The wrong process runs correctly. This is M-01 (LOADED ≠ REACHABLE) with the operand changed: RUNNING ≠ THE THING YOU SHIPPED.
  • Configuration read-backs prove nothing about code. The override mode is read from /etc/gpus-forms-routing.env, which no deploy step writes. It was going to say OFF regardless.
  • An identity assertion must name the expected value. cat VERSION printed a hash on the first attempt too. It is only a gate if something compares it to the hash intended — a bare read of a correct-looking value is the same trap one layer down.
  • Absence-based checks are the weakest form. "no unknown_kind in the journal" is satisfied by a build that never had that failure mode, by a build that is not processing anything, and by a build that is not running. Prefer asserting a new symbol present over asserting an old symptom absent.

Same family as the VLN-039 note in the vulnerability trackera green status means least exactly when a pipeline has just changed, there inverted to a red one. Both reduce to: a status summarises the last thing that happened, not the thing you were trying to make happen. VLN-048 and VLN-049 are the same class again — checks that passed while asserting something untrue.

Applied. The worker README's MAPLE block now fetches before it copies, with the identity assertion between the two rather than after the restart — see gpus-forms-routing-worker/README.md. Placing it before the first sudo is deliberate: a check that runs after the restart tells you what you broke, one that runs before it tells you not to.


M-14 — A correction is not complete until every co-located claim agrees with it

Observed 2026-08-31, on platform/travel-approval-workflow.md. Third instance of the same mechanism, and the first one where the sweep that missed it was performed by someone who had just written the correction.

What was on the page. abb6c61 (2026-08-27) corrected the v3 exercise instructions: naming Kevin Toruno as SMT approver does not mail nobody, it collapses step 1 to SKIPPED_DUPLICATE and dispatches straight to Shereyll Woodley's live queue. The correction was written as a !!! danger admonition headed "CORRECTED 2026-08-27 — this section previously said to pick Kevin Toruno. That was wrong." It was thorough, it showed the resolver output for both branches, and it was right.

Eight lines below it, in the "What to submit" table, line 810 read:

| SMT approver | `Kevin Toruno` | one human, already in the chain at step 3 |

Not a leftover fragment — a table row with a stated rationale, sitting under a heading that says what to submit, which is precisely what a reader skims to when they have decided to act. Whichever half a reader landed on is what they followed. The advice inverted three times, twice from the same author, and the reason it kept inverting is that the page argued both ways simultaneously and each half read as deliberate.

Why the correction felt complete when it was not. The author corrected the claim they had been thinking about — the prose that made the wrong argument — and the table was a different artifact on the same page, holding the same claim in a different shape. A grep for Kevin Toruno would have found it in under a second. Nobody ran one, because the correction had been written, and writing a correction produces a strong sense that the thing is now correct.

The constraint: a correction is not complete until every co-located claim agrees with it, and the sweep must cover the same PAGE, not just the same file. After writing a correction, grep the rendered page for the thing you just said was wrong — the value, the name, the number — and read every hit. A correction that leaves a contradicting neighbour standing has not reduced the error rate; it has made the page argue with itself, which is worse than the original error because both halves now read as considered.

The three instances, and what widened each time:

Correction What was left standing Scope that would have caught it
VLN-047 The outcome mail stopped denying the revise path SubmissionStatus's comment still argued the old copy was right Same file
VLN-049 /status and the outcome mail fixed The served instructions column, a test pinning it, and two stale comments — four surfaces, one of them in the database Same claim, across files and data
This The v3 exercise prose corrected A table row on the same page, eight lines below Same page

The scopes are not nested and that is the point. VLN-049's lesson was "the sweep grepped SOURCE; this lives in DATA" — a widening outward. This one narrows: the contradicting copy was in the same file the author had open, below the fold. Widening the sweep to other files and to the database does not help if the check is never run against the page in front of you.

Corollaries:

  • The dangerous residue is the confident kind. A bare stale value invites doubt. A stale value with a rationale — "one human, already in the chain at step 3" — tells the reader it has been thought about, exactly as approval_render.py's VLN-047 note records: "a well-argued justification defends a stale conclusion better than a bare one does." Prefer deleting a contradicting rationale to amending it.
  • Admonitions do not bind what is below them. A !!! danger block reads as authoritative to someone who reads it, and as decoration to someone scrolling to the table they came for. Correcting the prose and leaving the operative table is the worst split available: the warning is seen by the cautious reader and missed by the one about to act.
  • State the supersession in the surviving copy, not only in the correction. Line 810 now carries "this row said Kevin Toruno until 2026-08-31 and contradicted it" inside the table cell. The next reader who lands on the table alone gets the correction without needing to have read the admonition.
  • This is M-13 with the operand changed again. M-13: a signal that would read the same either way is not evidence. Here: a page that reads correct in the section you edited is not evidence the page is correct. Checking the part you changed is the thing that feels like verification and is not.

Cheapest available form of the check, applied the same day: after correcting travel-approval-workflow.md, a grep for Kevin Toruno returned six hits; four were legitimate (he is the step-3 approver), one was the correction itself, one was line 810. Three seconds. It also surfaced Sushma Raman in the candidate list — out sick, still in the roster, so a request naming her dispatches into a mailbox nobody is reading, which is worse than a name that would BLOCK visibly. The sweep found a second defect the correction had not been about. That is the usual return on it.


M-15 — A counter that resets is an identity assertion; a status that says "running" is not

Observed 2026-08-31, on the routing-worker v4 deploy. M-13 said a signal that would read the same either way is not evidence. This is the constructive half: what a restart proof actually looks like, and it is cheaper than the thing it replaces.

The four signals available after systemctl restart, and what each is worth:

Signal Reads Would it differ if the restart had not happened?
systemctl is-active active (running) No. True of the old process too
Clean journal, no errors clean No. The old process was healthy
boot approval_recipient_override=OFF OFF No. Read from a file no deploy writes
Heartbeat iterations= 1 YES

The heartbeat line at the moment of the deploy:

17:40:57 python3[3621123]: approval.drain.heartbeat RUNNING iterations=11225 …
17:43:03 python3[4075181]: approval.drain.heartbeat RUNNING iterations=1     …

Two things change together and neither can be faked by a process that did not restart: the PID, and a monotonic counter that went back to 1. A process cannot un-count. iterations=11225iterations=1 is the assertion; active (running) is a description.

Prefer a signal that RESETS over a signal that is TRUE. Uptime counters, iteration counts, "since boot" totals and PIDs all carry identity because their value is derived from the process's own life. Status words — active, enabled, healthy, OK — are properties the previous instance also had, and asserting them proves the service exists, not that it changed.

Why this is worth its own entry rather than a line under M-13. M-13 was written after a deploy where three of four gates read green on a process that had not been replaced, and its constraint is a prohibition: do not accept a signal that would read the same either way. Applied honestly that leaves you asking "so what DO I read?" — and on 2026-08-27 the answer was version=, which required the artifact to carry a version marker someone had built. The heartbeat needed no such construction. It was added for an unrelated reason — making drain silence falsifiable after the 2026-08-18 reachability defect — and it turned out to be a restart proof for free, because it counts.

So the durable form is a selection rule, not another marker to build: when verifying that a long-running process was replaced, look first for something already in the logs that counts, and read it before reading anything that merely affirms.

Corollaries:

  • A reset counter proves replacement; it does not prove WHICH build. Pair it with an identity marker (version=, a file digest). iterations=1 on the old binary is a restart of the wrong thing, and reads identically. The two assertions answer different questions and neither substitutes for the other.
  • Ordering matters, for the same reason as M-13. Read the counter and the version FIRST. Once you have seen active (running) you are already disposed to read everything after it as confirmation.
  • This generalises past deploys. "Did the config reload?", "did the cache clear?", "did the connection pool recycle?" — all have the same shape, and all have a status word available that answers a weaker question than the one being asked.
  • Where the estate already has one: approval.drain.heartbeat and approval.dlq_drain.heartbeat on the routing worker (iterations, messages, errors, all "since boot"). gpus-reports has no equivalent — its deploy proof is VERSION plus, as of this entry, a file digest. That is an identity assertion without a liveness-reset one, which is the right trade for a cron-invoked script that has no long-running process to restart.

M-16 — An assertion that prints an expectation and exits 0 is not a gate

Observed 2026-09-01, on the gpus-reports deploy of 3f63f33. The post-sudo verification block printed three expectations. All three were wrong, the installed state was right, and the script printed INSTALL COMPLETE.

window fix present (want 1 and 1):   ->  2 and 2
report_cron.sh bytes (want 13605):   ->  14878

Every mismatch was benign — grep -c counts lines and both symbols legitimately occupy two (volume_report.py:188 the def, :652 the bind; report_cron.sh:112 R1, :115 R3), and the byte count moved because e4fd73c rode along in the same deploy. That is what makes it worth recording: the check was harmless on the day, and it was still broken.

The defect is not the wrong numbers. It is that being wrong cost nothing. A block that says want X, observes Y, and continues has the shape of a gate and the behaviour of a comment. Its real output is a reader who learns that the numbers under ASSERTIONS do not have to match — and that reader skims the one line that would have caught a bad deploy.

Same class as || true on a security scan, and worse in context: this was the verification block of a runbook written to fix a drift problem, where the whole premise is that a signal reading green on the wrong state is more dangerous than no signal (M-13).

Two ways out, and they are not equal.

  • Make the expectation correct. Cheap, and it decays. All three values here were hand-computed against a payload that then changed underneath them — 13605 was accurate when written and stale two commits later.
  • Make the expectation DERIVED, and make the mismatch halt. ← the fix taken. The staged payload is present at verification time, so the assertion becomes every installed file is byte-identical to the file just staged (cmp -s in a loop, exit 1 on the first difference). Nothing is written by hand, so nothing goes stale; it names the offending file rather than saying something is wrong; and it is correct for a VERSION-only deploy, where a whole-directory digest that "MUST differ" is wrong by construction.

The corollary is the reusable part. If you are about to type an expected value, first ask whether the thing you are checking against is already on disk. In a deploy it almost always is — that is what staging is for.

And if a number must be printed without an expectation, print it with no claim attached. Both runbooks now keep the digest line as explicitly informational. A number that asserts nothing cannot be a gate that does not gate. What is not allowed is the middle: a stated expectation with no enforcement.

Where this was fixed: gpus-reports/README.md step 5b and gpus-forms-routing-worker/README.md step 5 — both of which carried the identical print-and-continue shape, written the day before, and neither of which had yet been wrong.

GOV-026 — Engineering-practice material has no home, so it accumulates inside registers

Field Value
id GOV-026
title There is no "How We Build" or engineering-practice document in the portal. Method material — how this estate verifies, what a control has to prove, where findings actually come from — lives as an unnumbered Method learnings section inside a findings register, and as prose scattered through individual VLN- rows.
register Governance / tracking gaps
status open
priority / severity not assessed — none supplied
owner unassigned
evidence Recorded 2026-08-27 when M-11 needed a home. find mkdocs-portal/docs -iname "*how-we*" -o -iname "*engineering*" -o -iname "*practice*" -o -iname "*ways-of-work*" returns nothing, and the portal nav has no such section (Guides, Governance, Priorities, Architecture, Infrastructure, Platform Services, Host Registries, Security, Compliance, Response Plans, Incidents, Change Management, Access Control, SSO & Identity). The eleven M-NN entries in this file are the estate's only collected engineering-practice material, and this file's own purpose statement is governance and tracking gaps — places records or checks do not do what they are documented to do. M-11 was filed here because it is the existing carrier, not because it is the right one.
acceptance test Method material is findable by someone who has not read the registers. Concretely: a reader looking for "how does this estate verify a control" finds it from the portal nav without knowing that gap-register.md exists. Migrating M-01M-11 is the obvious first move; the test is discoverability, not the migration.
blocks
blocked_by
date_raised 2026-08-27
date_verified

Why this is a gap and not a filing preference

The M-NN lessons are constraints on future controls — the section says so. They are the most reusable output this program has produced, and they are reachable only by someone who already knows to open a register of tracking gaps and scroll past a hundred rows. The material most likely to change how the next piece of work is done is filed where it is least likely to be read before that work starts.

Deliberately not solved by creating a page in passing. A "How We Build" doc with no owner and no review cadence would become the next artifact that is internally coherent and quietly stale — which is the failure mode half the rows above describe. It needs an owner first.

Incomplete rows

Four rows are missing a mandatory field. They are recorded as incomplete rather than completed with a plausible guess, per the register's own rule.

ID Missing What was supplied instead What would complete it
GOV-002 acceptance test Nothing A decision on whether the standard drops to two surfaces or the check rises to three
GOV-003 acceptance test Nothing A stated close condition — most likely the sentinel replaced by expiring entries in .coverage-exceptions.yaml
GOV-004 evidence "This is the phoebe mechanism, still open" — a characterisation, not a record A dated scheduler/job listing demonstrating no reconciliation control exists
GOV-006 acceptance test "Retain as a closed record of the defect class" — a disposition, not a close condition A stated observable condition; note GOV-009's test is not imported as a substitute

GOV-005 is partial rather than incomplete: its evidence names specific identity strings and an ADC date, but no read date or command output.

2026-08-17 load

Eight of the ten rows in the 2026-08-17 load are missing an acceptance test. Every one of the ten carries evidence with a named source and a date, so none is incomplete on evidence — a marked improvement on the 2026-08-13 load.

ID Missing What was supplied instead What would complete it
GOV-011 acceptance test "This must gate disposition, not merely annotate it" — a design requirement An observable condition, e.g. a stated catalog-retention threshold below which disposition is refused
GOV-012 acceptance test Nothing A stated close condition for the five ASSERTED-STALE nodes
GOV-013 acceptance test "Mirror /etc/puppetlabs/code as a filesystem capture, not a git clone" — a method, not a close condition The observable state that ends the row — presumably a verified capture, but the verification is unstated
GOV-014 acceptance test "Mirror before any teardown work touches those projects" — a sequencing rule A close condition; note the 39 GB repo is unassessed, so "mirrored" may not be sufficient
GOV-015 acceptance test Nothing A stated close condition for the twelve power devices
GOV-016 acceptance test A design constraint on a control that does not yet exist Properly a property of GOV-004's control design; closes when that control joins correctly
GOV-017 acceptance test Nothing A stated close condition; note VLN-013's widened test covers only the water rows, not fire, flower or the storage models
GOV-019 acceptance test Nothing A stated close condition; the immediate step is API rights for the querying account

GOV-010 and GOV-018 are complete on both mandatory fields.


Change log

Version Date Author Change
v1.32 2026-09-09 R. Chhetry / Claude GOV-045 fully closed — the second acceptance clause passes, and it was the one worth writing. /var/log/messages and /var/log/boot.log on SKY and RAIN both return zero DHCPDISCOVER/DHCPOFFER lines, so local7 still terminates at the & stop and the CEDAR forward was genuinely added above it rather than the stop being removed. Both clauses had to be tested separately: clause 1, DHCP arriving at CEDAR from both hosts, is satisfied equally well by simply deleting the & stop — the tempting fix, which would have flooded two log files without bound. An acceptance test that only asserts the presence of the thing you wanted cannot distinguish a fix from a regression that happens to include it. The clause asserting an absence is what makes the pair diagnostic, and it is the same shape as GOV-044's requirement to prove a non-AIDE reader still alerts after the exclusion. Recorded as a method point, not just a passing test.
v1.31 2026-09-09 R. Chhetry / Claude GOV-038 step 1 applied — the narrow firewall rules exist and carry consumer-named descriptions; step 3 is HELD on an open question. cedar-ingress-wdc-syslog (tcp:5140) and maple-ingress-wdc-agents (tcp:1514, tcp:1515) created in gpus-vpc from 192.168.120.0/23, priority 800 ahead of the broad rules at 900 so the later narrowing is a non-event rather than a cutover, target tags cedar-logging and maple-monitoring matching the existing objects exactly. Each description names consumers and senders — this row's remedy applied to rules being created, not only recorded about rules that exist. They are explicitly NOT proven: the broad rules still permit the same traffic, so both feeds flow identically whether the narrow rules match or not, which is M-13 and is stated rather than glossed. Three facts read out of the objects change step 3. (1) cedar-ingress is one object spanning 22/5140/5601/9200/10000/9100, so it grants 9200 to the entire WDC LAN — the cloud layer permits every WDC host to reach CEDAR's unauthenticated Elasticsearch and only firewalld withholds it, confirming VLN-053's layer disagreement from the object itself and making step 3 a security fix rather than tidying. (2) maple-ingress carries Grafana 3000 and Prometheus 9090 from the WDC LAN with no established consumer either way; narrowing would break office Grafana access silently, so this is answered before step 3, not during — an open question in front of a change is a reason to stop. (3) Port 10000 is now dead config in both objects after Webmin was disabled today, so it should be removed at step 3 rather than narrowed — the accretion this row describes, except we created this instance ourselves today, which makes it cleanup rather than archaeology. Also: 9100 is scraped from 172.16.0.0/24 and does not need the WDC range, and flow logging is disabled on both objects, so the evidence that would answer (2) was never being collected.
v1.30 2026-09-09 R. Chhetry / Claude GOV-045 raised and FIXED — DHCP was absent from the central store entirely, and the obvious one-line fix would have broken two other logs. CEDAR held zero dhcpd documents over 24 hours beside 1,435 from named. The two nameservers failed differently and SKY was worse: RAIN's local7 reached WIND then stopped, SKY's reached no destination beyond the local file. So the only searchable DHCP record was WIND's dead index, and half of it was never sent even there. The & stop is load-bearing, not stray: rsyslog.conf includes rsyslog.d/*.conf at line 37 with its own defaults at lines 47 and 66, and DHCP uses local7 — the facility RHEL reserves for boot logging — so deleting the stop would have flooded /var/log/messages and /var/log/boot.log unbounded. The fix is to add above the stop, not remove it. Applied to both with rsyslogd -N1 validation, backups to /root and not into the *.conf load path. Verified against a zero baseline: 114 → 162 documents, +48 in 120 s, rain 78 and sky 30 over 45 minutes, correct process.name, facility: local7, current timestamps, no _grokparsefailure. Also found: RAIN's comment reads "Forward to Logstash on SKY" above the address of WIND — a plausible route by which this drifted unnoticed. Acceptance clause 2 deliberately asserts the absence of DHCP lines in messages and boot.log, which is the clause distinguishing this fix from the one almost made.
v1.29 2026-09-09 R. Chhetry / Claude Session close-out: GOV-044 raised, GOV-039 (3) settled by enumeration, Webmin closed on both SIEM hosts, handoff published. Webmin disabled on MAPLE and reset-failed on both — 10000 closed, Wazuh ingestion moving across the change with five agents reporting, SOC portal wazuh.ok true. The portal's soc_score moved 75 → 55 anyway, and the cause was our own procedure: ten level-12 100015 credential-dumping alerts, all of them exe="/usr/sbin/aide" calling llistxattr on /etc/shadow during the documented Step 5 baseline update. It fires on --update and not on the daily --check, on exactly two days in fourteen, both operator-run. The five on 2026-09-02 went unremarked, which is the sharper half — a rule already being ignored cannot detect what it was written for. openvas: false was checked against a snapshot taken earlier the same day and is pre-existing, not caused by the change. GOV-039 (3): WIND enumerated read-only and the row's framing was wrong twice. It is not orphaned — SKY holds 2 connections and RAIN 3 — and its blocked forward targets CEDAR, the live store, so semanage would close this row's own coverage gap rather than restore unwanted traffic. But WIND has indexed nothing since 2026-08-31: its single-node cluster sits at 999/1000 shards, 518 primaries plus 481 replicas it can never assign, rejecting every write with HTTP 400 — a collector blocking its own ingestion. Nobody reads it: Kibana binds 127.0.0.1, zero dashboards, zero saved searches, zero visualisations. semanage still not run. DHCP is the one non-duplicated feed30-dhcpd.conf's & stop means CEDAR has never received dhcpd at all (24 h of CEDAR syslog: named present, dhcpd zero), though /var/log/dhcpd.log is current on both nameservers with daily rotation, so nothing is lost except the searchable copy. SKY's "unexplained" connection was action(type="omfwd") syntax containing no @, invisible to the grep that called it unexplained — the fourth empty-result-from-the-wrong-path this week. Handoff published at governance/session-handoff-2026-09-09.md routing five workstreams, including two dated items: *.us.gl3 expires 2026-09-17 where the action is confirming nothing broke on the 18th rather than renewing, and the world-readable CA private key on phoenix (/opt/puppet/dist/ca/, 0644) which signs every internal certificate and has no owner in any register.
v1.28 2026-09-09 R. Chhetry / Claude GOV-041 RESOLVED on both nameservers, and its own diagnosis was wrong; GOV-043 raised; GOV-040's scope question answered. The cause was not drift, not the power event and not a missing time source: /etc/chrony.conf line 5 read denyall, which chrony 4.5 does not accept — the syntax is deny all, two words — so chronyd exited 1 on every start. This row had proposed "enable chronyd and verify against a source they can actually reach"; that would not have worked, because the daemon was already enabled and the source was reachable throughout. It has never run on these hosts. chrony-4.5 was installed 2026-02-26 09:36:36 on SKY and 2026-02-27 11:19:19 on RAIN, and the configs were written 29 and 6 minutes later, so the file was wrong onto an already-installed 4.5; the upgrade-invalidated-a-valid-directive hypothesis was tested against dnf history and refused — one chrony version has ever been in either rpmdb. Both files byte-identical at 244 bytes, one paired build step. What corrected the clocks until May was a reboot: both are VMware guests with vmtoolsd active, and the 121-day drift matches uptime since 2026-05-11 exactly, so the estate never had time discipline on its nameservers, it had a power cycle frequent enough to hide the absence. Applied RAIN first: +916 s and +961 s forward steps (slow, therefore the benign direction for a signer), named not restarted on either, zone NOERROR from both, RRSIG 20260909060801/20261009060801 unchanged and still valid, AIDE promoted unconditionally with a separate mv. SKY's step crossed the 15:00 cron minutes and nothing re-ran — checked, not assumed. Two dependencies recorded rather than fixed: no internal NTP server exists (MAPLE, CEDAR, SUN all time out on udp/123), so DNS now depends on the VPN for time while DNSSEC validity is wall-clock bounded; and nothing alerts on a dead time daemontimedatectl's NTP service: inactive reads as "not configured" and actively conceals enabled+failed. GOV-043: Storage=auto with no /var/log/journal on any of the four WDC hosts, so none has ever persisted a journal and each retains ~14 volatile days lost on reboot. systemctl --failed names RAIN alone, and the other three report active precisely because flushing is a successful no-op — the healthy-looking hosts are in the identical state. The correct probe is test -d /var/log/journal. RAIN's own failure, 2026-05-12 09:05:24 EDT with a status=0/SUCCESS main PID, is unrecoverable: the volatile journal reaches back only to 2026-08-26. Not fixed — it would change write behaviour on the DNS pair and RAIN's cause should be understood first. GOV-040: enumerated on both hosts — no puppet package, no puppetlabs/puppet directory in any layout, no binary, no service, no converge cron — so d088b4bf is not "managed but never run" there, it is not managed at all, and the chrony fix cannot be reverted by a converge. Also observed: dnf-automatic-install.timer is active on both nameservers, which explains today's unrelated AIDE drift and may warrant its own row. change-management/procedure.md Step 5 corrected — it taught the aide --update && sudo mv idiom that skips promotion exactly when drift exists.
v1.27 2026-09-09 R. Chhetry / Claude GOV-039 (1) and (2) APPLIED AND VERIFIED on CEDAR; GOV-041 and GOV-042 raised; two of this register's own recorded claims corrected. The Logstash change went in at 14:10:47Z after a Configuration OK test, backup to /root and not to conf.d, first restart since 2026-04-09. Verified against the index by a document that checks itself: event.original carries t=2026-09-09T10:11:30.012-04:00 and @timestamp reads 14:11:30.000Z, so the fix agrees with the event's own zone-qualified clock rather than merely looking current. Zero _grokparsefailure, zero placeholder tags, zero documents still landing four hours back, MAPLE's 5141 feed moving across the restart, SOC portal wazuh.ok: true. This row's own acceptance test was wrong and would have failed the working fix — a now-10m window cannot return SKY and RAIN, because they log in bursts (2,859 and 1,820 per day against SUN's 6,942) and because their clocks are a quarter-hour slow. Replaced with a per-host logger marker pushed through the real path, which is independent of both. GOV-041: SKY and RAIN have NTP service: inactive and System clock synchronized: no, running 16 and 15 minutes slow while SUN, WIND, MAPLE and CEDAR are all disciplined and exact. The evidence was already in v1.26 — SKY and RAIN trailed SUN by 17 and 21 minutes in the 2026-09-08 table, attributed to low volume; a uniform offset explains the four hours and not the residual, and the residual was not interrogated. These are the two hosts that sign wdc.us.gl3, and DNSSEC validity is wall-clock bounded. GOV-042: cloud-vms.md §11.2.1 documents a beats/json/wazuh-alerts-* pipeline that is not deployed; the live file is a nine-line udp input writing to /tmp/wazuh-debug.log, and because pipelines.yml runs one pipeline over conf.d/*.conf with neither output guarded, Wazuh alerts are indexed unparsed into cloud-logs-* (2,019 untyped documents today, all 172.16.0.12) while WDC syslog is written into the debug file. A second unauthenticated copy of the alert corpus, local rather than network, companion to VLN-053. The 2026-09-08 claim that removing it would have destroyed MAPLE's Wazuh feed is withdrawn: Filebeat writes wazuh-alerts-* directly to :9200, soc-backend/app.py queries that index at five sites and cloud-logs-* at none. The caution was right; the reason given for it was not. Also recorded: MAPLE's /etc/rsyslog.conf:6 forwards to 5140 and is not connected, so the fix's "every 5140 sender is EDT" constraint holds by accident — if that forward ever succeeds, MAPLE is UTC and would index four hours into the future.
v1.22 2026-09-01 R. Chhetry / Claude GOV-035 RESCOPED and GOV-036 raised, after a 25-minute outage. The incident GOV-035 was raised from was memory exhaustion, not a boot window: Uncaught signal: 7, pid=1 (SIGBUS) on 00117-d5x at 16:47:54 and 16:48:43, exit(1) at 16:51:41, and no container completed create_app() between 14:55:58 and 17:10:49. It began BEFORE the deploy — 00118 failed identically, so rollback would not have helped — and the database was healthy throughout (0.3 s queries, 11 of 25 connections), so the Cloud SQL contention hypothesis was looking in the wrong place. The startup probe is now the follow-up, not the fix; the fix was 512Mi → 1Gi (00119-wls), nothing else changed. The boot-window analysis stands and is unchanged. FormsBackendDown recorded as a TRUE POSITIVE — it was dismissed as a boot-window false alarm and closed as noise, and it was in fact the only signal in the estate that detected a real outage. Rule needs no change; no other rule shares the exposure, because the other four key on rate()/increase() over values and go silent rather than firing. Its cost is recorded too: the Alertmanager path does not auto-resolve, so every deploy bills a manual close to US - IT Security — which is the probe's real argument, not "cleaner signals". GOV-036: 512Mi / minScale=0 / concurrency=80 were all correct when set on portals with no traffic and are now unexamined inheritances on a service carrying live approvals — surveyed across all ten services, plus db-f1-micro, --timeout 60, MAILTO=root and the memory:// limiter. Acceptance is a recorded answer to what workload is this sized for, not bigger numbers.
v1.21 2026-09-01 R. Chhetry / Claude M-16 added — an assertion that prints an expectation and exits 0 is not a gate. The 3f63f33 install printed want 1 and 1 against 2 and 2, and want 13605 against 14878, then printed INSTALL COMPLETE. All three mismatches were benign — grep -c counts lines and both symbols legitimately occupy two; the byte count moved because e4fd73c rode along — and that is the point: being wrong cost nothing. A block shaped like a gate that behaves like a comment teaches the reader that the numbers under ASSERTIONS need not match, which is exactly the line that would catch a bad deploy. Same class as || true on a security scan, and worse inside the fix for a drift problem. Fixed by derivation, not by correcting the numbers: the staged payload is on disk at verification time, so the check is now every installed file is byte-identical to the file just staged (cmp -s, exit 1 on first difference) — nothing hand-written to go stale, it names the offending file, and it stays correct for a VERSION-only deploy where a "MUST differ" directory digest is wrong by construction. Applied to both runbooks, gpus-reports/README.md step 5b and gpus-forms-routing-worker/README.md step 5, which carried the identical print-and-continue shape written the day before. Corollary recorded: before typing an expected value, ask whether the thing you are checking against is already on disk.
v1.20 2026-09-01 R. Chhetry / Claude GOV-035 amended — the "nothing else in the project failed" leg was hollow, and checking it was worth more than the conclusion it supported. Cloud Logging for 14:43–14:52 shows zero /metrics scrapes to gpus-status-backend, gpus-soc-backend, gpus-forms-frontend, gpus-mkdocs-portal or gpus-forms-clamav-worker — and zero requests of any kind to four of the five. Confirmed at source: /etc/prometheus/prometheus.yml on MAPLE carries exactly one Cloud Run job, gpus-forms-backend; every other job is a host node_exporter. One Cloud Run service of roughly ten is under metric monitoring, so the others' silence carries no information. The finding stands — its mechanism was established positively from per-instance latencies and the thirteen-request drain, not by elimination — but the elimination argument is withdrawn. Recorded as a rule: a quiet neighbour is only evidence if someone was listening to it. Sharpens the detection gap: this fault class is invisible on nine services including gpus-soc-backend, whose slow /api/soc is GOV-033's subject.
v1.19 2026-09-01 R. Chhetry / Claude GOV-035 raised — /metrics 504s at exactly 59.99 s, and the handler was never the problem. generate_latest() over four in-process collectors touches no database; what blocked was create_app() running load_all() at gunicorn worker boot, 24–34 s that afternoon. Gunicorn binds :8080 before its workers finish importing, so a container advertises readiness it does not have and Cloud Run queues requests into it until they age out. Timestamped to the millisecond: db.engine.created 14:51:59.556yaml_loader.done 14:52:23.856thirteen queued scrapes all completing 200 in 45 ms at 14:52:24.3. One of three instances was warm and answering in ~4 ms throughout. Three findings recorded separately: the endpoint is unauthenticated with allUsers holding run.invoker and ingress: all, on the same 8 gunicorn slots as live travel approvals, with containerConcurrency: 80 — and its 300/hour limiter is memory://, i.e. per-instance and empty on a fresh boot, weakest exactly when the queue is deepest. Reports do not contend with the handler (it has no database access at all); boots during the window were 3–4× slower than the 12:31 quiet boot, but that is correlation on three points and concurrent boots against a db-f1-micro explain it at least as well — cause not established, and the Cloud SQL metrics needed were never captured. min-instances=1 is mostly a red herring: a warm instance already existed and served fine; minScale does not prevent scale-out, which is what hurt. Second time the 0 cost default has made "cold start" the wrong first hypothesis. The mechanism fix is a startup probe. Not fixed in this pass.
v1.18 2026-09-01 R. Chhetry / Claude GOV-032 gains its second real data point, and it goes the other way. af7ea350 (2026-09-01) entered 0 for EventRegistrationFee while its Purpose text records that each NRO is recharged 1,600 Euros per participant attending physically — so the approval mail and the queue-45 ticket both show Event registration fee (USD): 0 above a total that excludes it. Not a defect: the question was answered as asked, and no wording change is proposed here — that is Shereyll's to decide. The v4 compromise's mechanics are unaffected (0 does say "none" unambiguously); what is now in question is its semantics — the label can be read as "a fee I pay" rather than "the cost of attending". Recorded explicitly as not an argument for conditional display: a conditionally-revealed amount box asked the same way collects the same 0. With Madison's correct compliance on 2026-08-31 the row now carries two submissions, one each way, which is the honest state and a better basis for revisiting the field than either alone. Also noted: this row's own acceptance query cannot see it, because 0 is a valid answer, and no report surfaces it.
v1.17 2026-09-01 R. Chhetry / Claude GOV-034 raised, and GOV-033 gains its follow-on PRG-029 — both lifted out of the travel workstream as it closes, because neither is a travel problem and both are about the estate being unable to tell anyone that something failed quietly. GOV-034: cron's MAILTO=root is the only unattended failure channel that exists without someone having built one, and it is inert on 7 of 7 hosts — no root: alias anywhere, Postfix inactive on six (CEDAR, OAK, SKY, RAIN, SUN, WIND) so their output is discarded outright, and MAPLE — the only host with an MTA — holding 439,418 bytes of unread job output. The two failure modes are recorded as different: a silent discard is worse than an unread file and looks identical from outside. Probed on all seven hosts, with the six hosts' actual cron-output disposition recorded as UNVERIFIED rather than assumed. It covers the AIDE checks, Lynis, both backup jobs and all eight reports. PRG-029 carries GOV-033's third recommendation — a GCS latest.pdf freshness check across all eight report types — as program work rather than a gap, following the GOV-028PRG-028 precedent exactly: the GOV row records what is missing, the PRG row carries the prevention-grade build. It is deliberately scoped not to depend on GOV-034, so report silence becomes detectable without first fixing estate-wide mail. No row closed, no status changed.
v1.16 2026-09-01 R. Chhetry / Claude GOV-033 raised — the executive monthly report did not send on 2026-09-01 and nothing said so. Aborted at 08:00:01 on a 30 s read timeout to /api/soc, which answers in 14.3 s measured the same morning — the timeout is sized at ~2× real latency, so this recurs monthly. Two defects, recorded separately because they need different fixes: _fetch_soc_data() returns {} for every failure, so a transport error and an empty payload are the same value at if not soc_data: sys.exit(1); and report_cron.sh's own log "ERROR: PDF generation failed" is unreachableset -euo pipefail kills the wrapper at the PDF_PATH=$(…) assignment, and that line has never printed in the log's history against five successful monthlies. Every candidate signal was enumerated on the host and every one fails: the wrapper line cannot fire, nothing parses the log, cron's MAILTO=root lands in a /var/spool/mail/root with no /etc/aliases entry that is now 439 KB of unread success mail, there is no Prometheus metric or textfile collector, and Wazuh coverage of the log is UNVERIFIED for want of an interactive sudo and is recorded as unknown. Fix recommended in three parts and not implemented — align the executive family with the degraded= render-and-send contract R2 and R5 already use (a behaviour change to a board-adjacent report, and Rajesh's call), size the timeout from measurement with retries, and add one GCS latest.pdf freshness check covering all eight report types. Only the third would have caught this failure rather than making the next one louder.
v1.15 2026-09-01 R. Chhetry / Claude GOV-032 count corrected, and forms/README.md:74 fixed at source. The row published in v1.14 said eight legacy forms carry div_lock; new-employee-notification alone has thirteen. Both numbers were eyeballed from a truncated grep listing. Counted, and cross-checked against the live fields table on 2026-09-01: five forms, nineteen fields, top form 14 — plus one in forms/_schema.yaml, which yaml_loader.py:174 excludes from the load glob and which is therefore not a form at all. M-08 in its own register, published as fact. The argument is unchanged and slightly stronger: activating div_lock would alter five forms nobody asked to change, and the fourteen-field HR form is the whole blast radius either way — still larger than building conditional display fresh, which touches only forms that opt in. forms/README.md:74 no longer reads "Conditional show/hide": it reads INERT — DOES NOTHING, followed by a note giving the full layer-by-layer path the key travels, the inventory of dead values, and an explicit instruction not to add new ones. The decision remains declined.
v1.14 2026-09-01 R. Chhetry / Claude GOV-030, GOV-031 and GOV-032 raised. GOV-030: fact D of the ASVS scope statement — and approval_access.py:28-30 restating it — argue the frozen dispatch snapshot is acceptable because revocation exists as an administrative re-dispatch, and instruct the assessor to audit that path. The path does not exist: VOIDED has one write site, inside the decision transaction, and the approval API is three routes with no withdraw, reassign or re-dispatch among them. Not a code defect — an assurance document claiming a compensating control the estate lacks, with a574e4cb live at 6 days as the concrete instance. 021_withdraw_test_chains.sql has described the absent withdraw honestly since 2026-08-19, so the two documents have disagreed since the day the second was written. GOV-031: R2 is the only mechanism that surfaces a stalled approval; it goes to Shereyll alone and excludes Kevin by documented decision. Both are right, and together they leave the terminal step uncovered — e5ff2b59 is Shereyll's own trip, DISPATCHED to Kevin for 12 days, reported daily to the one person who cannot act on it. No escalation exists anywhere in the estate. Recorded with a runnable query and an explicit note that adding Kevin to R2 is not the fix. GOV-032: a decision record, not a work item — conditional display was requested twice by Shereyll, compromised around twice (v3 funding fields, v4 registration fee), and declined 2026-09-01 by Rajesh and Shereyll together; do not build or scope it. Carries both requests as recorded, what shipped instead, the two-part constraint (no conditional display and no cross-field validation), the single observed data point (Madison complied correctly with an unenforceable rule), and what a third request of the class looks like. One correction to the recorded reasoning, from source: div_lock already exists through yaml, the fields table and the form-schema API and is documented in forms/README.md:74 as "Conditional show/hide" — it is inert, with zero references in forms-frontend/src/, and five live forms carry nineteen dead values (figure corrected in v1.15 — v1.14 published "eight … thirteen", both wrong). That does not reopen the decision; it makes the gap's true shape a missing renderer plus a documented-but-absent capability. No row closed, no status changed.
v1.11 2026-08-27 R. Chhetry / Claude GOV-028 CLOSED on live evidence. Deployed to MAPLE at 59a253c; both acceptance commands pass on the host — (a) a live run emits version=59a253c matching the deployed marker, (b) no commit touching a shipped file is newer than it. The deploy also cleared the drift the finding was about: count_trips went 0 → 1 and the first R5 run afterwards reported trips_mtd=11 revisions_mtd=1 over the 12-row/11-trip/1-revision August dataset. The stamp is confirmed at all four surfaces including build 59a253c in the PDF footer. The row was held open through two sessions while the guard existed only in the repo, because closing on a merged commit would have been the same error the finding describes. PRG-028 opened for the prevention-grade step — this makes drift visible, not impossible.
v1.10 2026-08-27 R. Chhetry / Claude GOV-029 raised — found while deploying the GOV-028 guard. /opt/gpus-reports is cloudadmin-owned while the modules inside are root:root, so the files can be replaced by rm+cp without sudo; the root ownership confers no protection it appears to. Probed non-destructively on the host. Noted as weakening GOV-028 — a drift guard answers whether the deployed code matches what was pushed, not whether the deploy put it there — and the one-line fix is deliberately kept out of the drift-guard deploy so a rollback can separate the two causes.
v1.9 2026-08-27 R. Chhetry / Claude GOV-028 gains a built mechanism and a failing acceptance test — and stays OPEN. The drift guard is implemented and tested in the repo, mirroring the routing worker's rather than inventing a second pattern: gitignored VERSION, a report_version() of the same shape, emitted on the first stdout line of all eight report types, in each degraded= line, in the mailer's start line, and in every PDF footer. The row is not closed: the finding is about a host, /opt/gpus-reports/ still has no VERSION and is still three commits behind, and both acceptance commands still fail. Closing on "the code is written" would be the same error the finding describes. Also recorded: what the guard does not solve (visibility, not prevention — and R5 has one reader), and why the stale staging objects are left in place under 021's reasoning, with the condition that made that precedent transferable.
v1.8 2026-08-27 R. Chhetry / Claude GOV-027 and GOV-028 raised, both found while running assigned verification rather than from a failure — M-11 again. GOV-027: the required-field gate checks key presence, not value, on every form. GOV-028: gpus-reports has no drift guard and MAPLE ran three commits behind for two days, shipping R5 without the trip-deduplication fix while logging degraded=False. GOV-028 is M-13 recurring in a second system that the lesson never reached. No row closed, no status changed.
v1.7 2026-08-27 R. Chhetry / Claude M-13 added to Method learningsa signal that would be true either way is not evidence the change landed. Recorded from the routing-worker redeploy of 2026-08-27, where three of four post-restart gates read green on a process that was still the previous build; only version= distinguished them. Companion to M-12, which records the cause (a persistent staging directory) rather than the detection failure. The worker README's MAPLE block was restructured in the same change so the identity assertion sits between the fetch and the first sudo. No GOV- row added, none closed, no status changed.
v1.4–v1.6 2026-08-22 – 2026-08-27 R. Chhetry / Claude ROWS MISSING — recorded rather than reconstructed. The header advanced to v1.6 across the loads that added GOV-022GOV-026 and M-11/M-12, but no change-log rows were written for them. They are not invented here from the diff; the commits 9c32983, 2a27293 and 754f67b are the record until someone re-derives them. Flagged so the log is not read as complete.
v1.3 2026-08-21 R. Chhetry / Claude GOV-009 CLOSED AND VERIFIED 2026-08-18 — the first done row in this register. 73 register rows confirmed rendering live on infra.greenpeace.us at tracker v1.11, gap-register v1.2, priorities v1.21: all six pages HTTP 200, nav entries present, all 51 internal anchors resolving after 5d177fc / build f02200d4. One acceptance clause amended: the VLN-009VLN-012 redirect required a resolving anchor, but VLN-012 is a table row and table rows receive no heading id, so no target exists or can exist without restructuring — amended because the test asserted a capability the structure never had, not because it failed. GOV-006, GOV-007 and VLN-022 are unblocked by this closure and are not re-statused; they await re-adjudication.
v1.2 2026-08-17 R. Chhetry / Claude Evidence-completion pass. The enumeration corpus at ~/estate-enum-2026-08 became reachable; artifact paths appended to the evidence cell of GOV-001, GOV-005, GOV-010, GOV-011, GOV-015, GOV-017, GOV-018 and GOV-019, each stating what was re-verified against the artifact. GOV-005 upgraded from partial to sourced — access-map.md states the ADC/CLI divergence directly. Two evidence dates flagged wrong: raw/stage2-A1-inventory-census.txt reads Captured 2026-08-13, not the 2026-08-17 recorded on GOV-001 and GOV-010; figures unaffected, dates corrected in the rows rather than in the backlog, which is source text. No acceptance test, status, severity or owner changed.
v1.1 2026-08-17 R. Chhetry / Claude 2026-08-17 enumeration load (recorded retrospectively — this row was omitted when v1.1 shipped). Added GOV-010GOV-019 from priorities/backlog-2026-08-17.md items B-15–B-24, and a new Method learnings section carrying M-01–M-10 unnumbered, status-less and unowned. GOV-001 rewritten per B-15 — the finding is scope, not competing claims — with the 2026-08-13 wording retained verbatim as a dated superseded note and explicitly not closed. GOV-004 gained GOV-016 as an attached design constraint: join on serial/MAC for devices and IP for DNS. Eight of the ten new rows lack an acceptance test and are marked INCOMPLETE; all ten carry dated evidence. No row is done.
v1.0 2026-08-13 R. Chhetry / Claude Register opened. Nine rows, GOV-001GOV-009, mapped from provisional G-1G-9. Evidence for GOV-001, GOV-002, GOV-003, GOV-007 and GOV-009 confirmed first-hand against the repo and recorded as file paths and line numbers; the GOV-001 citation was corrected (no information-asset-registry.md exists — the competing claim is at hostregistry/index.md:13). Four rows marked INCOMPLETE on a mandatory field. No row is done.