Skip to content

Session Log

Classification: CONFIDENTIAL — Internal Use Only Document: priorities/session-log.md · v1.6 · 2026-08-24 · GPUS-IT

Reverse-chronological log of work sessions on GPUS-IT infrastructure and forms portal. Latest session is always at the top.

2026-08-24 — VLN-014 reopened; R2 delivery confirmed; GOV-020 raised

VLN-014 is not done and has been returned to in progress. The entry below, dated 2026-08-21, is corrected — not withdrawn. Its evidence was sound; its timing was not.

  • VLN-014 — REGRESSED 2026-08-19, reopened 2026-08-24. Queried live on both hosts: SOA serial agrees (2026081702) but the NSEC3 salts do not — sky CB18BB4A142868CF against rain A178C4BCAFEA2FA0 — and RRSIG inception differs, sky 20260819063701 against rain 20260817170231. rain still holds exactly the values verified on 08-17; sky has moved on without it. On sky the source zone is unchanged since Aug 17 14:01 at serial 2026081701 while the signed file was rewritten Aug 19 03:37 at 2026081702re-signed without bumping the source serial, so -N INCREMENT reproduced the serial rain already had and rain correctly declined to transfer. The same mechanism the row itself documents, and the exact failure its remediation note warns against in its closing sentence.
  • The lesson is about timing, not evidence. The 08-17 verification was real and thorough. It was written up on 08-21 from that verification rather than from a fresh query, and the regression landed on 08-19 in between. Close a row only from a query run at the time of closing.
  • GOV-009 — unchanged, still done. Re-checked 2026-08-24: the live render path works and its acceptance test remains met. Its closure was sound and is not affected.
  • R2 delivery confirmed end-to-end. Both ends of the 2026-08-24 travel approvals mail: Postfix status=sent dsn=2.0.0 to smtp.gmail.com, and the message in the recipient's inbox. The last unproven claim in the R2 build. Cron has now run four times, 08-21 to 08-24, all degraded=False.
  • GOV-020 raised — two tools share the single working tree at ~/gpus-infra-portals, and a push from either publishes whatever the other has committed, unreviewed. e0b2bf5 reached origin that way on 2026-08-20. Recommendation: one tool per working tree.

2026-08-21 — VLN-014 and GOV-009 closed; a stale note corrected

The first two done rows in any of the three registers. Both were verified end-to-end against live state, which is what the status rule has always required and what nothing had previously met.

  • VLN-014 — DONE, verified 2026-08-17. sky and rain converged: SOA serial 2026081702, RRSIG inception 20260817170231, expiration 20260916170231, NSEC3 salt A178C4BCAFEA2FA0 — all identical on both, keytag 6660 unchanged. rain's salt flipped from 83C78EB3B3503D4C, which is the proof it took the new signing. Transfer confirmed in rain's log; delv fully validated on both; dsset md5 unchanged so no parent DS update; five untouched zones on sky verified identical; named PID unchanged, so a reload not a restart.
  • GOV-009 — DONE, verified 2026-08-18. 73 register rows rendering live on infra.greenpeace.us at tracker v1.11, gap-register v1.2, priorities v1.21. All six pages HTTP 200, nav entries present, all 51 internal anchors resolving after 5d177fc / build f02200d4.

The root cause of VLN-014 was recorded as its symptom

The row said "the re-sign never bumped it, so rain had no transfer trigger." That is what was observed. The actual mechanism: dnssec-signzone -N INCREMENT bumps the serial in the output file only and leaves the source untouched, so every re-sign from an unchanged source reproduces the identical served serial — and a secondary declining to transfer an unchanged serial is behaving correctly. rain was never faulty. The 2026-07-29 and 2026-08-12 signings both derived 2026072902 from source 2026072901; two signings were lost this way before detection. The verified procedure, with the source serial bumped first, is recorded on the row.

Why the note in e0b2bf5 was wrong, and what the gap actually is

Commit e0b2bf5 added a note to VLN-014 on 2026-08-20 stating it was "accepted for action" with "eight days to expiry". The remediation had already been performed on the hosts on 2026-08-17 — three days earlier.

This is not a tooling error. It is the gap the registers exist to close. The note was written from repo state, and repo state was the only state available: the remediation happened on sky and rain and was not recorded anywhere the register could see. Every check that could have been run was run, and all of them agreed the row was open, because the artifact said so and the artifact was stale.

That is the same shape as GOV-001 — a declared source of truth that covers less than the estate — and the same shape as the Tier A rows in the 2026-08-13 flag list, where a recorded action was read ever after as a verified outcome. Here it ran in the opposite direction: work done on the hosts, not written down, and therefore invisible. A register is only as current as the discipline of writing to it, and a remediation that closes a row needs to reach the row.

Also amended: the acceptance test for the VLN-009VLN-012 redirect asked for a resolving anchor. VLN-012 is a table row and table rows receive no heading id, so no target exists or can exist without restructuring. Amended because the test asserted a capability the structure never had, not because it failed. The prose redirect serves the intent.

Left deliberately unchanged

  • VLN-023 stays in progress. Three acceptance conditions, one reported, none re-observed. Correct as written.
  • VLN-024 stays verbatim, including the second-working-set framing and the per-vendor rotation requirement.
  • GOV-006, GOV-007, VLN-022 were each blocked_by: GOV-009. That blocker is now gone, but none is re-statused here — each has its own acceptance test, and two of the three are INCOMPLETE on that field. They need a deliberate re-adjudication pass, not an automatic promotion.

2026-08-20 — Owner decisions on the two CRITICALs and the DNSSEC clock

Three decisions recorded on their rows. None makes anything done, and no acceptance test was changed.

  • VLN-023 → in progress. A spare is reported assigned on vmstorage. The acceptance test has three conditions — array normal, spare assigned, scrubbing available — of which one is reported and none re-observed. Closing it means re-running syno_probe.py (read-only by construction) and confirming all three. A 26.79 TB rebuild takes time, and a volume mid-rebuild still has no redundancy. The action date was not stated and is not inferred.
  • VLN-024 → deferred to the RHEL→Rocky migration and the Terraform vendor move. Recorded as a dated disposition; status stays open. The migration does not by itself close the row: new Terraform-managed credentials are a second working set, not an invalidation of the first. The six values stay live at litle, convio, adp and target until rotated or disabled vendor-side, so the migration needs an explicit per-vendor step or this row survives it. The exposure window is now the migration timeline, on credentials plaintext since 2019-01-24 across 879 commits and five working copies.
  • VLN-014 → accepted for action. Status stays open; prioritisation is not remediation. Eight days to the 2026-08-28 RRSIG expiry on rain. The NSEC3 salts differ between sky and rain, indicating two independent signings rather than one that failed to transfer — bumping the serial alone may not converge them, and the acceptance test's "salt identical on both" is what catches that.

Corrections to the record

The entry below was written believing the working date was 2026-08-17. The device clock reads 2026-08-20. Dates in that entry taken from artifacts are sound; any taken from the session's own sense of "today" should be read with that in mind. This entry is dated from the device clock.

GOV-009's push half is done. main is in sync with origin/main and all six register commits are in history and pushed. The render check is now actually possible and has not been performed here — this workspace still cannot reach infra.greenpeace.us, so GOV-009 stays open on its verification half.

One anchor written on 2026-08-17 was broken, and someone else fixed it. Commit 5d177fc"docs(tracker): fix VLN-013 anchor — em-dash slugifies to single hyphen" — corrected the VLN-004VLN-013 link and took the tracker to v1.11. This is precisely what GOV-009 predicts: structural validation confirmed the target file existed and could not confirm the slug resolved. The other anchor written that day, the VLN-009VLN-012 redirect, should be checked on the live render for the same class of fault.

Unrelated work has moved the repo substantially since — 38 commits, mostly the forms travel-approval workflow. None touched the registers except 5d177fc.


2026-08-13 → 2026-08-17 — Estate enumeration Stages 1–3b; 38 backlog items loaded into the registers

What the enumeration found

Six stages ran between 2026-08-13 and 2026-08-17 — Stage 1 (surface enumeration), 1b (DNS cross-reference), 1c (Meraki and Synology), 2 (declared-vs-live reconciliation), 3 (Puppet classification) and 3b (repo provenance). The estate was enumerated across four locations plus five Meraki networks and three Synology units.

Headline numbers:

Measure Figure
Declared entries in inventory.yaml 82
Deduped live entities found 481
gpusa — declared vs live zero declared against 174 live
Classified UNKNOWN at Stage 2 50.0%
gpusa VMs moved UNKNOWN → DERIVED at Stage 3 24 of 25

82 against 481 is the finding of the week. It is also the explanation for something the 2026-08-13 build could not account for: why the component-coverage gate has been passing. The gate validates only what is declared, and nothing declares the largest location — it cannot fail on the 399 entities it cannot see. That reframes GOV-001 from a documentation inconsistency into a control failure.

50.0% UNKNOWN is the honest figure, and every prior artifact implied near-complete knowledge. The 24-of-25 movement to DERIVED at Stage 3 looks like that resolving, and mostly is not: GOV-011 records eleven certnames against thirty-nine node definitions, with thirteen RUNNING VMs — including phoenix itself — that have never produced a retained catalog. DERIVED is a confidence statement; it is not enforcement.

Shipped

  • security/vuln/tracker.md v1.8 → v1.9VLN-023VLN-036 (14 rows, 2 CRITICAL, 5 HIGH, 6 MEDIUM, 1 LOW).
  • governance/gap-register.md v1.0 → v1.1GOV-010GOV-019, plus a new Method learnings section (M-01–M-10, unnumbered and status-less).
  • priorities/gpus-it-priorities.md v1.19 → v1.20PRG-013PRG-026.
  • governance/register-id-mapping-2026-08-13.md v1.0 → v1.1 — B-/M- to real-id mapping appended.
  • Source text committed separately as priorities/backlog-2026-08-17.md v1.0 (commit 6f045bc).

Amendments to existing rows — four, and no others

  • GOV-001 rewritten per B-15: the finding is scope, not competing claims. The 2026-08-13 wording is retained in full as a dated superseded note, and is not closed by the rewrite — its own acceptance test remains unmet.
  • GOV-004 now carries GOV-016 as an attached design constraint: any reconciliation control must join on serial/MAC for devices and IP for DNS. Hostname is not a reliable join key here.
  • VLN-019 cross-referenced to PRG-024 as second-source corroboration — gpus-dist zones say 10.1.96.40, the live VM is .46, inventory.yaml says .46.
  • VLN-013 acceptance test widened to "no artifact asserts water is ESXi 6.7 or hosts VMs", covering VLN-004, inventory.yaml:146-152 and downstream generated output. This closes open item 2 from the 2026-08-13 session log.

No other existing row was modified. The 13 flagged Done rows from 2026-08-13 remain untouched — re-adjudication is still GOV-008.

The two CRITICALs

  • VLN-023vmstorage is running a degraded RAID 6 with no spare (disk_failure_number 1, spares [], scrubbing unavailable), with three healthy disks sitting in not_use. It is the NFS datastore for every WDC VM, mounted by fire, water and flower. Nothing in monitoring surfaced it and no artifact mentions it. VLN-026 adds that the unit has no backup or replication task, and GOV-019 adds that snapshot configuration is unknown — a 403, not an absence.
  • VLN-024 — six plaintext vendor passwords for payment, payroll and CRM vendors in modules/relay/files/etc/pushtab, last committed 2019-01-24, deployed to every storehouse node, in a repo with 879 commits and five known working copies. Neither the Rocky migration nor deleting the file closes it; only vendor-side rotation does. PRG-019 decommissions the code that reads the file and does not close this row.

Method learnings recorded

Ten constraints on future controls are now in the gap register, deliberately unnumbered and status-less because there is nothing in them to close. The one that keeps recurring is M-01, LOADED ≠ REACHABLE — now at three instances in this program: VLN-010's Wazuh rules, rooster's commented-out class (VLN-034), and pushFPR firing every 30 minutes after exiting 0 at line 7 (PRG-019). M-08 is the other one worth internalising: four counts supplied from memory were overturned during these sessions.

Evidence-completion pass — later the same day

The four enumeration folders were connected mid-session, making the corpus at ~/estate-enum-2026-08 reachable for the first time. 34 rows across the three registers gained a verified artifact path in their evidence cell — every one re-checked against the artifact rather than matched on filename.

Confirmed to the exact figure: the four DNSSEC facts behind VLN-014, including both NSEC3 salts; EXIT_CODE=0 on permission-denied probes (VLN-017, M-03); six plaintext pushtab rows at lines 1–6 (VLN-024); Snapshot list: FAILED err=403 (GOV-019); 82 inventory entries with the gpusa bucket at 0 (GOV-001 / GOV-010); 11 PuppetDB certnames (GOV-011); 2 gpusa Cloud SQL instances, 18 gpus-it buckets, disk ratios 7:3 and 15:9, 47 static IPs, 25 gpusa instances, and 15 active DHCP leases.

access-map.md was found. On 2026-08-13, VLN-016 was recorded as citing an artifact that did not exist in the portal repo. It exists in the enumeration corpus, and its coverage matrix confirms the identity inversion exactly.

HB-004 gained evidence from phoebe-capture/ — six candidate consumers with vhost configs and logs. The logs are unanalysed, so its acceptance test stays INCOMPLETE.

Three things did not reproduce, and were flagged rather than adjusted:

  1. raw/stage2-A1-inventory-census.txt reads Captured 2026-08-13, not the 2026-08-17 recorded on GOV-001 and GOV-010. The figures are unaffected; the dates were wrong by four days and are corrected in the rows, not in the backlog, which is source text.
  2. VLN-029's counts: dnssec .private = 1 as stated, but the index holds 19 htpasswd/ files against 18, and 643 files under ca/ against 81 — the 81 is presumably a private-key subset whose filter is not recorded.
  3. PRG-011's reservation count: the capture holds 126 host declarations against ~111 stated. The lease side, 15 active, confirms exactly.

Also noted: VLN-018's evidence date is the derivation, not the observation — task4-cloudrun-20260813.txt states its source JSON was generated 2026-08-12T20:28:28Z.

No acceptance test, status, severity or owner was changed by this pass, and nothing became done — better evidence is not a render. GOV-009 stands.

NOT verified

Nothing in this load is done. mkdocs build --strict still cannot run in the Cowork workspace — mkdocs-material is not installed and the workspace has no network — so the build aborts on Unrecognised theme name: 'material'. Structural validation is not a render. Everything loaded here stays authored until the render is confirmed live on infra.greenpeace.us. That is GOV-009.

PRG-019 is a disposition decision by R. Chhetry, not a verification, and is recorded open.

Commits are local only. Nothing was pushed.

This entry does not close GOV-007

GOV-007 (session-log staleness) stays in progress. Two entries four days apart do not establish a cadence, no acceptance test has been supplied for what "current" means, and the 2026-04-24 → 2026-08-13 gap is still unbackfilled.

Open for next session

  1. Push, then verify the renderGOV-009. Until then everything since 2026-08-13 is authored, not done.
  2. VLN-023 and VLN-024 are the two CRITICALs and neither has been started. VLN-024 needs vendor-side action that nobody else can do for it.
  3. Acceptance tests: all 14 PRG- rows and 8 of 10 GOV- rows in this load lack one. 12 of 14 VLN- rows in the 2026-08-13 load still do too.
  4. GOV-019 is cheap and unblocks a severity assessment — API rights for the querying account would let VLN-026 be scored.
  5. GOV-014's ordering constraint is time-sensitive — mirror the three CSR repos before any teardown touches gpus-it-infrastructure, which is already scheduled for deletion, and 39 GB of it has never been assessed.
  6. VLN-014's clock — RRSIGs on rain expire 2026-08-28, now 11 days out.
  7. VLN-010 D6 (24h volume) was due 2026-08-13 and still reads open.

2026-08-13 — Three IT registers opened; VLN-009 duplicate resolved

Log gap: 2026-04-24 → 2026-08-13

The entry below is the first since 2026-04-24. Over that span gpus-it-priorities.md advanced from v1.1 to v1.18 — 17 documented version bumps with no corresponding session record. This entry does not backfill them. The gap is tracked as GOV-007, which remains in progress, not closed: one entry does not close a ~16-week hole, and no acceptance test was supplied for what "current" means here.

Shipped

  • security/vuln/tracker.md v1.6 → v1.8. Commit 7ff38ae applied the approved reconciliation; a later commit added the new-schema register section.
  • governance/gap-register.md new, v1.0 — the governance/tracking-gap register, GOV-001GOV-009.
  • priorities/gpus-it-priorities.md v1.18 → v1.19 — new ## Program work register — schema v1 section, PRG-001PRG-012 plus HB-001HB-004. No pre-existing row was edited.
  • priorities/flag-list-unverified-done-2026-08-13.md new, v1.0 — 13 pre-existing Done rows flagged. List only; no rows changed.
  • governance/register-id-mapping-2026-08-13.md new, v1.0 — provisional S/G/P identifiers to real IDs.
  • Nav entries added for the three new documents.

Key decisions made

  • VLN-009 stays with the passwordless-sudo finding (CVSS 9.9, open). The chronyd/NTP entry was reassigned to VLN-012 with a provenance note. The 2026-08-10 "recorded, awaiting decision" admonition was converted into a dated record of what was decided and retained, not deleted, so a stale "VLN-009 = chronyd" reference still resolves.
  • VLN-011 is In Progress, not Done. it-support-request is closed end to end and ticket-verified, but new-employee-notification-contractor-intern is still held and contract-extension-notification is unreviewed. One of two affected forms verified is not done.
  • No CVSS was minted for VLN-011. The finding page states none; footnote 2 records why, mirroring VLN-010's footnote 1.
  • Three registers, three ID series: VLN- (security findings, existing tracker), GOV- (governance gaps, new document), PRG-/HB- (program work and host-level blockers, inside the existing priorities doc). GOV- was chosen over GAP- because GAP-1 already means something else in the forms workstream.
  • New schema sits alongside the old, in the same files. Existing rows keep their original vocabulary and wording; the schema and the status rule apply to new items only.
  • HB- rows are attached to hosts, not to the program — so the disposition pass is not made to look blocked on four unrelated questions.

Found by reading

  • inventory.yaml:149 carries the same water/ESXi-6.7 misattribution as VLN-004hypervisor: VMware ESXi 6.7 on a host the vSphere API shows running 8.0.3 (build 24022510). Found while evidencing PRG-010.
  • inventory.yaml:152 records vms: [ocean] on a host with zero VMs, and :149 reads "currently hosts Ocean (KACE SMA)"PRG-010's mis-documentation is in the inventory itself, not only in prose.
  • gpus-forms-clamav-worker IS listed in inventory.yaml (as cloud_services.gpus_forms_clamav_worker). The item was filed as "4 previously unlisted Cloud Run services"; it is 3. The other three (gpus-security-backend, gpus-soc-site, gpus-status-backend) are genuinely absent — zero matches each.
  • The GOV-001 citation was wrong. There is no information-asset-registry.md in the repo. The competing single-source-of-truth claim is at hostregistry/index.md:13, against three documents asserting inventory.yaml (inventory-schema.md:10, component-coverage-standard.md:58, adding-new-infrastructure-quickstart.md:9).
  • GOV-002 and GOV-003 confirmed first-hand at scripts/check-component-coverage.py:564 (PORTAL_DIRS = two portals) and :557 (PORTAL_RENDER_SENTINEL), against .coverage-exceptions.yaml, whose header states "Indefinite exceptions are not allowed by design" while its exception list is empty — so the only bypass in effect is the one kind the file forbids.
  • The host rows HB-001HB-004 were meant to attach to do not exist. gannet, emu, ostrich, catbird, phoebe — zero matches each in inventory.yaml. Only catbird and phoebe appear anywhere under docs/, as /32 allowlist entries at compliance/iar.md:381-382.

Verified working

Structural validation of tracker.md, 2026-08-13: table pipe-count consistency (no mismatches), internal link resolution (5 of 5 targets exist on disk), admonition indentation, and VLN- id uniqueness — 12 ids, no duplicates, i.e. the defect this session set out to fix is gone from the file.

Git verification that no pre-existing row was touched: the only deletion in the gpus-it-priorities.md diff is the v1.18 version header line. 466 insertions, 1 deletion.

NOT verified — read this before trusting anything above

Nothing shipped today is done. mkdocs build --strict could not run: mkdocs-material is not installed in the Cowork workspace and it has no network access, so the build aborts on Unrecognised theme name: 'material'. Structural validation is not a render. Specifically unverified:

  • that any of the three registers renders on infra.greenpeace.us;
  • that the two new anchor links in tracker.md (VLN-004VLN-013, and the VLN-009VLN-012 redirect) resolve in a built site;
  • that the new nav entries appear.

Tracked as GOV-009, which blocks GOV-006, GOV-007 and VLN-022.

Commits are local only. Nothing was pushed.

Open for next session

  1. Push, then verify the renderGOV-009. Until then every 2026-08-13 row is authored, not done.
  2. Decide VLN-013's scope. Its acceptance test names only VLN-004 and will not catch inventory.yaml:146-152. Either widen it or raise a separate row. Not widened unilaterally.
  3. Fill the acceptance tests. 13 of 16 program rows have none, 6 have no evidence; 4 governance rows are incomplete. One pass stating the close condition per row clears most of it.
  4. Assign owners. Every row in all three registers reads unassigned. None was supplied and none was inferred.
  5. VLN-014 has a clock on it — RRSIGs on rain expire 2026-08-28, 15 days out. The only dated external deadline in the new rows.
  6. VLN-010 D6 (24h volume) was due 2026-08-13 and is still recorded as open in the tracker. Not touched this session.

2026-04-24 — Phase 2 forms backend LIVE

Shipped

  • Commits 9511f5a and 8c1f018 pushed to main
  • forms-backend Cloud Run service deployed to revision 00027-sdv with Phase 2 stubs
  • Added forms API contract, frontend design spec, and Phase 2 deploy checklist to mkdocs portal
  • auth_v2 module shipped as additive layer alongside Phase 1 auth (no breakage)

Key decisions made

  • Went with the cream theme over dark for the forms portal (matches Greenpeace brand warmth, better legibility for long-form intake)
  • Moss-green CTA buttons instead of standard Greenpeace campaign green (differentiates internal tooling from public-facing activism sites)
  • Fork submission flow into long-form vs short-form paths rather than a single adaptive form (cleaner validation, simpler per-path audit)
  • auth_v2 is an additive module — Phase 1 /api/forms* routes stay exactly as they are; Phase 2 gets its own namespace
  • Role derivation lives on the token, not in a DB lookup (keeps the hot path stateless)

Bugs found and fixed

  • Org-wide scope bug in gpus-okta-auth.js — fixed in commit 6cf8c5d
  • Canonical email mismatch: Okta profile returns .us but tokens carry .org — normalized in auth layer
  • Repo had an older Phase 2 scaffold that predated the v1 work — removed before layering in the new stubs

Verified working

  • /health/phase2 returns 200 with build metadata
  • /api/me returns caller identity with derived roles
  • /api/config returns Phase 2 feature flags
  • /api/admin/forms gated correctly (admin-only)
  • Token-based role derivation confirmed for both submitter and admin roles
  • 10 environment variables set on Cloud Run revision 00027-sdv (verified via gcloud run revisions describe)

Open for next session

  • Deploy forms-frontend Cloud Run service
  • Wire Phase 2 routes through to the Phase 1 data layer
  • SQLi tabletop + red/blue drill — this gates Phase 3 and must happen before any new write paths land

Reference