Skip to content

GPUS-IT Priorities

Classification: CONFIDENTIAL — Internal Use Only Document: priorities/gpus-it-priorities.md · v1.23 · 2026-09-01 · GPUS-IT Owner: Rajesh Chhetry · Review cadence: weekly (Fridays) + on every new initiative


Purpose

Living tracker of active and queued initiatives for the GPUS-IT infrastructure program. This document supersedes scattered priority notes across memory and Cowork; the canonical order of work lives here.

Rules of engagement:

  • Every initiative lists its gating dependencies — don't start work that is gated on unfinished prerequisites.
  • Every new asset added by an initiative must ship with its matching IR runbook (rb-00N-*.md), DR procedure update in drp.md, and red/blue drill entry in tabletop-playbooks.md + blue-team-drills.md. Infra docs without runbooks are not considered "done".
  • Status values: Planned · In Progress · Blocked · Done · Deferred
  • When an item completes, move it to Completed with the completion date and delete it from the active queue.

Sequencing

Forms portal is LIVE in production as of 2026-07-23 (go-live submission 28bd1ebf; routing worker f9a114b on MAPLE; override OFF; all 6 ingest addresses verified delivering; HappyFox tickets opening). Phases 2.5(a)–(d) are DONE and the broken-access-control remediation is complete. The forms program has shifted from build to cutover + hardening.

The order is now: cutover forms, harden forms, then Meraki, then WDC foundation.

  1. Forms cutover (T1, the only hard external deadline). Legacy forms portal (forms.us.gl3) retires 2026-08-01 — parallel running until then. The staff notice (email + Slack) and legacy decommission must land before that date. NOTE: Aug 1 is a Saturday — flag the cutover date for reconsideration (an LDAP deprovisioning gap argues for cutting over sooner, not later).
  2. Forms hardening (T1). Detection pipeline (blocks the drill program), remaining security hardening, content/data fixes, dedicated Okta forms app, 2.5(e) HappyFox API path, and portal propagation. Documentation set and the security exercise program (SQLi tabletop + blue/red drills) are gated on the detection pipelines existing.
  3. Meraki (T2) — Sedita fix + Meraki SSO, then fold Meraki into status/SOC/MkDocs coverage.
  4. WDC foundation (T3) — ESXi + Synology NAS inventory, cleanup, monitoring, and documentation. Deferred intentionally: the hypervisor + NAS are stable enough today that finishing the forms cutover + hardening is higher-value.

Everything else (SOC Ticketing, Vendor Access, MySQL decom, status site automation) slots in after WDC foundation unless a specific gating dependency inverts the order.


Summary

Tier Initiative Status Target
T1 Forms portal LIVE in production — go-live 2026-07-23 (sub 28bd1ebf); routing worker f9a114b on MAPLE, override OFF, all 6 ingest addresses verified delivering, HappyFox tickets opening Done 2026-07-23
T1 Broken-access-control remediation — Phase-1 decrypt/list/submit endpoints removed; require_role fail-closed; resolve_user least-privilege; IDOR ownership checks on Phase-2 endpoints. NO evidence of exploitation (empty audit_log + empty users table, independently corroborated) Done 2026-07-23
T1 Forms 2.5(a)–(d) backend + routing worker — submit/attachment/finalize wired, subject-template audit-persist gate resolved, GAP-1 renderer fix + template reconciliation landed Done 2026-07-23
T1 Forms cutover → forms.greenpeace.us — staff notice (email+Slack), legacy forms.us.gl3 decommission, in-flight handling, redirect-or-dark. Only hard external deadline. LDAP deprovisioning gap argues for cutting over sooner. STAFF-NOTICE CONTENT REQUIREMENT added 2026-08-25: the notice MUST state that portal mail arrives from alerts@greenpeace.us and is Gmail-tagged External, and that this is the permanent steady state — not a temporary condition. Per the 2026-08-25 decision that no send-as alias will be registered (see gpus-reports/README.md and architecture/forms-phase2.5c-design.md §7). A reader told this is temporary reports it as phishing three months later, and a phishing report against a real portal address costs more to unwind than the sentence costs to write. The four staff guides already carry the wording; the notice must match them, not paraphrase. In Progress 2026-08-01 (Sat — reconsider)
T1 Detection pipelineaudit_log→SOC + Cloud Run→Wazuh log sinks (neither exists); Wazuh rules 100026–100029 inert by starvation, not config. Blocks the exercise program Planned 2026-08
T1 Security hardening — post-go-live — retire legacy auth.py authz primitive; narrow backend-SA project-level over-grants; Cloud Armor/WAF (arch change + cost case); /metrics + /health/deep exposure; CORS *; per-instance rate limiting; MAC field validation Planned 2026-08
T1 Content / data fixes — ~~<% = Note => in Finance Termination template rendering literally in prod~~ DONE 2026-08-10 (template b1a51710e91b8b18{{ Note }}; it is the Facilities template, which Finance receives because it is bundled onto the Facilities destination line — audit §5.7); ~~live-DB legacy-tag sweep~~ DONE 2026-08-10: 1 of 48 live templates carried a legacy tag, and it was that one — estate now clean; DB↔repo drift root cause (HIGH); Insurance checkbox unanswerable-as-No; HR Termination authoring TODO; footer sweep; DSAR/erasure path; attachments retention lock Planned 2026-08
T1 Credential-rotation control never closed out (compliance drift) — the quarterly forms-portal-credential-rotation-quarterly control emits records but is never back-filled: 3 ACTION-NEEDED carried since Q2 (KMS rotation timestamps, Cloud SQL backup verification, HappyFox credential rotation) + an MFA-enforcement flag; possibly two consecutive quarters with no live console verification. Same class as the contract-vs-code and DB↔repo drift findings — artifacts asserting a state nobody verified. PRE-ANNOUNCEMENT GATE: confirm Okta MFA enforcement before the staff notice goes out — forms authz now rests entirely on Okta identity and the announcement leads with "sign in with Okta". Planned MFA: before cutover notice
T1 Okta — dedicated forms app + group — requested from Conan 2026-07-23; then verify SPA redirect URI, aud validation, preview-tenant test, client-ID cutover In Progress 2026-08
T1 2.5(e) HappyFox API — queue names from API, ticket-audit migration, per-queue double-ticket deconfliction; API preferred over email-ingest (greenpeace.us DMARC p=none; help.greenpeace.org no DMARC → forged ticket plausible); ingest sender restrictions unverified Planned 2026-08
T1 Portal propagationinventory.yaml reflects live forms state; coverage + validate_portal_presence pass; status + SOC show forms LIVE (derived, not hardcoded); verify live pages render it Planned 2026-08
T1 Okta cleanup — remove localhost redirect URIs (separate from the dedicated forms app above) Planned 2026-08
T1 Forms portal Phase 1.5 — legacy data migration Planned 2026-08
T1 Forms portal Phase 1.6 — ON CONFLICT refactor Planned 2026-08
T1.8 Forms documentation set (mkdocs) — threat model; OWASP ASVS L2; PCI-DSS applicability (out of scope, no CHD); OWASP→NIST 800-53→MITRE mapping; forms IR runbook + DRP entry; BAC security-finding record; governance finding; 2.5(d) as-built; field-exposure matrix regen; mitre-attack/threat-vectors/pentest-schedule/calendar updates. Depends on the assessment Planned 2026-08
T1.9 Forms security exercises — SQLi tabletop (60m), blue-team detection drill (90m), red-team simulation (90m). GATED on detection pipelines Blocked 2026-09
T1 α ClamAV scan worker — Commit 3 (hardening + hygiene, 8 items) Filed 2026-08
T2 Meraki cleanup — Sedita site + Meraki SSO Planned 2026-06
T2 Meraki integration — status/SOC/MkDocs coverage Planned 2026-06
T3 WDC foundation — ESXi inventory & cleanup Planned 2026-07
T3 WDC foundation — Synology NAS inventory & cleanup Planned 2026-07
T3 GCP terraform tree has no VCS — dedicated session: secrets review of tfvars/tfstate, .gitignore design, git init, push to private remote (CSR), import state (supersedes WDC VPN/route drift framing) Filed 2026-06
T3 VPN cold-start packet loss — diagnosis + fix (Mac → GCP private subnet path drops on cold start; 2nd incident in 7 days) Filed 2026-06
T3 IAP-as-break-glass — IAP-to-MAPLE 4003 confirmed host-level 2026-06-10 (not IAP edge); fix VM-side so IAP works as break-glass path when VPN/SSH down Filed 2026-06
T3 Cloud Scheduler missed-tick on cold-start (sweep.clean missed 05:45 UTC 2026-05-21; investigate scheduler retry config vs worker min-instances tradeoff) Filed 2026-06
T3 SKY portal-backup cron died ~2026-04-17portals/ snapshots stop there (server backups current). Restore the cron; backfill the gap Filed 2026-08
T3 2.5(e) PRECONDITION — HappyFox ingest/API deconfliction. G4.4 proved email_template legs landing in HappyFox-watched inboxes (e.g. gpus-it-support@) auto-open tickets via email-to-ticket (#USITS00357259, 2026-06-12; worker dispatched no happyfox action — verified). Once 2.5(e) dispatches via API, forms with both action types against the same queue will DOUBLE-TICKET. Deconflict per-form email recipients vs API destinations before 2.5(e). The deconfliction points in OPPOSITE directions per form and there is no single rule: on it-support-request remove the EMAIL leg (the HF leg is the intended path); on employee-termination-notification remove the FIVE HF legs (the 3 email legs already cover all 5 owning teams — audit §5.10/§5.11). Miss the second and every termination opens 5 duplicate tickets across 5 queues. contract-extension-notification (3 HF + 1 email) is UNREVIEWED — do not assume either direction. ~~Context: org moved to API because HappyFox blocks greenpeace.us email~~ CORRECTED 2026-08-10: that claim is STALE or was never true. #USITS00365024 and #USITS00365073 both opened in queue 45 from mail the relay rewrote to alerts@greenpeace.us. HappyFox accepted it. The claim helped justify the move to API dispatch and was carried for months untested. The DMARC argument for API (row above) is unaffected and still stands. Filed feeds 2.5(e)
T4 SOC Ticketing tab Planned 2026-08
T4 Vendor Access Portal (replace SFTP) Planned 2026-08
T5 Status site automation — Phase A (live Cloud Run list) Planned 2026-09
T5 MySQL decommission Planned 2026-09
T5 Status site automation — Phase B (BigQuery billing export) Planned 2026-10
T5 Status site automation — Phase C (per-service cost attribution) Planned 2026-11
T5 T5-EXPANDED — Portal static-debt retirement (P0–P4, per 2026-06-10 audit; forms Gate 5 now landed → unqueued; P0 truth-fixes may interleave with forms hardening) Filed 2026-Q3

Recently completed forms/α items (see Completed table): forms 2.5(a)–(d), routing worker (Gate 4/5), BAC remediation, T1.7 SPA submit-gate, α ClamAV Commits 1 + 2.


T1 — Forms: cutover & post-go-live hardening

Forms portal is LIVE in production as of 2026-07-23. Go-live submission 28bd1ebf; routing worker f9a114b on MAPLE with override OFF; all 6 ingest addresses verified delivering; HappyFox tickets opening on real submissions. Phases 2.5(a)–(d) DONE (see the phase records further down this section, kept as history). The active T1 work is now cutover (the only hard external deadline) plus the hardening workstreams below. The phase-history sub-sections (Phase 2, 2.5(a)–(e), UX gaps, etc.) follow as completed records.

Cutover to forms.greenpeace.us — the only hard external deadline

  • Status: In Progress. Legacy forms portal (forms.us.gl3) retires 2026-08-01; new and legacy run in parallel until then.
  • ⚠ Date flag: 2026-08-01 is a Saturday. Flag for reconsideration — a weekend decommission means no staffed coverage if a rollback or in-flight-submission issue surfaces. Weigh against the LDAP argument below, which pushes the other way (sooner).
  • Items:
    • Staff notice (email + Slack). Operational, independent of every other workstream — has no technical dependency and must go out before Aug 1. Do not let it wait on hardening.
    • Legacy decommission of forms.us.gl3 — plan the teardown; decide redirect-or-dark for the old hostname.
    • In-flight submission handling — drain/route anything submitted to the legacy portal during the overlap window.
    • Duplicate-ticket ambiguity during overlap — both portals feed the same ingest addresses, so a form submitted on both (or mid-migration) can double-open tickets. Deconflict before/during the overlap (ties into 2.5(e) per-queue deconfliction below).
    • LDAP deprovisioning gap — argues for cutting over SOONER. The Okta→LDAP sync keeps breaking, so a disabled account may still authenticate to the legacy portal. Every day of parallel running is a day a deprovisioned user retains legacy access. This is the security case for an earlier cutover, in tension with the Saturday-date concern above.

Detection pipeline — blocks the drill program

  • Status: Planned. Gating dependency for the T1.9 exercise program — there is no point running blue-team detection drills against detections that cannot fire.
  • Items:
    • audit_log → SOC pipeline — does not exist. Forms audit events never reach SOC. (Overlaps the long-standing T1.6 observability workstream; this is the concrete blocker for drills.)
    • Cloud Run → Wazuh log sink — does not exist. Wazuh rules 100026–100029 are inert by starvation, not by misconfiguration — the rules exist but no forms/Cloud Run events are being fed to them.
    • Verify Wazuh rule bodies + ossec.conf on MAPLE — confirm the rule definitions and agent config are actually what we think before wiring the feed.
    • Routing-worker systemd unit has an empty SyslogIdentifier — fix so the worker's logs are attributable in the journal/sink (prerequisite for the Cloud Run→Wazuh feed to be useful).

Security hardening — post-go-live (remaining)

The broken-access-control remediation is DONE (see the dedicated record below). These are the remaining hardening items surfaced during/after go-live:

  • Retire legacy auth.py authz primitive (Step 4). Migrate /api/admin/reload to auth_v2; the users table becomes login-audit only. (Completes the auth-module consolidation begun in Phase 2.1 item 5.)
  • Backend service-account over-grants. The backend SA holds project-level roles/cloudkms.cryptoKeyEncrypterDecrypter + project-wide roles/storage.objectAdmin. Narrow to the specific key and bucket scope (least privilege).
  • Cloud Armor / WAF — NOTE this is an architecture change, not a policy toggle. forms.greenpeace.us is a direct Cloud Run domain mapping today; adding Cloud Armor requires standing up a new external HTTP(S) load balancer + serverless NEG in front. Treat as an architecture change + cost case, not a quick enablement.
  • /metrics is unauthenticated — leaks per-form submission volumes. Gate it.
  • /health/deep exposes DB + KMS status publicly — reduce to an authenticated/internal probe.
  • CORS origins="*" — tighten to the known SPA origin(s).
  • Rate limiting is memory:// per-instance — the effective limit multiplies by instance count (so the real ceiling is N× the configured value). Move to a shared backend or account for instance count.
  • MAC field accepts malformed input — add validation.

Broken-access-control remediation — DONE (2026-07-23)

  • Status: COMPLETE. Recorded here as a closed security finding; the full finding record (finding + remediation + no-exploitation conclusion) is a T1.8 documentation deliverable.
  • What was fixed:
    • Phase-1 decrypt / list / submit endpoints removed.
    • require_role made fail-closed.
    • resolve_user moved to least-privilege.
    • IDOR ownership checks added on the Phase-2 endpoints.
  • No evidence of exploitation — corroborated independently by two facts: the audit_log is empty of any decrypt/view events, and the users table is empty. Two independent signals both consistent with "never exploited."

Content / data fixes

  • <% = Note => legacy placeholder surviving in the Finance Termination template — renders literally into production tickets. Highest-visibility content bug post-go-live.
  • Live-DB sweep for other legacy tags — a repo sweep provably misses drift (see DB↔repo root-cause item), so the sweep must run against the live database, not the repo.
  • DB↔repo drift root cause — HIGH IMPORTANCE. If the repo is not the authoritative source for templates/forms, then the gap-2 sweep, the field-exposure matrix, and the CI leak-gate are all validating the wrong artifact. Root-cause which store is authoritative before trusting any repo-based check.
  • Insurance checkbox: required with only "Yes" — unanswerable as "No" (a required field the user cannot legitimately decline). Fix, and sweep for the same pattern across other forms.
  • HR Termination template contains an unfinished authoring TODO line — remove/complete before it renders to a recipient.
  • forms.us.gl3forms.greenpeace.us footer sweep — replace stale legacy-hostname footers.
  • No DSAR / subject-erasure path — there is no implemented data-subject-access-request or erasure flow; the retention purge executor is unlocated; the attachments bucket retention (2555d) is still UNLOCKED. Privacy/retention gap to close.

Okta — dedicated forms app + group

  • Status: In Progress — dedicated forms app + group requested from Conan 2026-07-23.
  • Then: verify the SPA redirect URI, aud (audience) validation, run a preview-tenant test, and perform the client-ID cutover to the new app.
  • Related but separate: the existing "remove localhost redirect URIs" cleanup (Phase 2.1 / Okta cleanup section) is a different task from standing up the dedicated app.

2.5(e) — HappyFox API integration

  • Status: Planned. Native HappyFox API ticket creation, distinct from today's live email-ingest path.
  • Items:
    • Queue names from the API (stop hardcoding queue identifiers).
    • Ticket-audit migration — persist HappyFox ticket IDs / outcomes into the audit trail.
    • PER-QUEUE double-ticket deconfliction — a single form routing to N queues risks doubling a ticket across all N once both the email-ingest and API paths are active. This is the same hazard as the cutover overlap duplicate-ticket item; solve once, per queue. (See the T3 "2.5(e) PRECONDITION" row for the G4.4 evidence.)
    • API vs email-ingest decision — security prefers API. Rationale: greenpeace.us is DMARC p=none (vs .org's p=reject), and help.greenpeace.org has no DMARC record at all, so a forged portal-looking email ticket is plausible on the ingest path. The authenticated API path removes that spoofing surface.
    • HappyFox ingest sender restrictions unverified — confirm whether the ingest inboxes actually restrict senders (if not, the spoofing risk above is live today).

Portal propagation

  • Status: Planned (residual of forms Gate 5). The standing rule applies: portal content must be LIVE or DERIVED, never hardcoded, and must render from sources, not merely have endpoints exist.
  • Items:
    • inventory.yaml reflects the live forms state.
    • Coverage + validate_portal_presence pass.
    • status + soc portals show forms LIVEderived, not hardcoded.
    • Verify the live pages actually render it (the audit standard: endpoints existing is not enough; the frontend must display the derived value).

Phase 2 React + Okta (forms frontend SPA)

  • Status: Phase 2 frontend SPA: COMPLETE 2026-04-27. Live on forms.greenpeace.us. Backend rev gpus-forms-backend-00033-bfx, frontend rev gpus-forms-frontend-00003-74f. 28 forms render with editorial typography. Auth pipeline (Okta OIDC PKCE → JWKS → /api/forms) fully verified end-to-end.
  • Goal: Ship gpus-forms-frontend as React SPA with Okta OIDC PKCE auth, calling gpus-forms-backend with a validated JWT bearer.
  • Gating: None — gpus-forms-backend is live on Cloud Run, forms.greenpeace.us resolves, Okta Production cutover complete (2026-04-23).
  • Deliverables:
    • forms-frontend/ in gpus-infra-portals repo
    • JWT validation middleware on gpus-forms-backend (and on status / security / soc backends — Phase 2 is org-wide, not forms-only)
    • Cloud Run deploy + Cloud Build trigger wired
  • Asset docs required: Update iar.md with forms-frontend service; no new IR runbook needed (covered by existing portal runbooks); blue-team drill entry for "forged/expired JWT rejected" in blue-team-drills.md.
  • Phase 2 follow-up: FieldRenderer pulldown lookup — COMPLETE 2026-04-28 (commit 5512d8e). Phase 1's /api/forms/<id> serializer now returns id/label/pulldown_id (matching contract.ts). New endpoint GET /api/pulldowns/<name> added with PulldownResponse shape. Backend rev gpus-forms-backend-00034-8tn.

Phase 2.5 — Phase 2 backend wire-up

  • Discovered 2026-04-28: the Phase 2 SPA has been live but Phase 2's three submission endpoints (POST /api/submissions, POST .../attachments, POST .../submit) are unwired stubs in routes_phase2.py. They return fake UUIDs and a hardcoded 'TBD@greenpeace.us' literal. Submitting a form via the SPA appears to work to the user but persists nothing — no DB write, no GCS upload, no email/HappyFox call. The well-formed routing model exists (actions table populated by the MySQL migrator with per-form action_type + destination + template_id) but no submit-time code reads it. There is also no mailer module in the backend at all.
  • Implementation scope:
    1. Wire create_submission (routes_phase2.py:102) to write to the submissions table with KMS envelope encryption for non-searchable fields, returning real UUIDs. Pattern reference: routes/submissions.py:37 (Phase 1's working submit handler).
      • STATUS: COMPLETE 2026-04-30
        • Commit 1 SHIPPED 2026-04-29: 07cb75c (auth_v2 User.username via preferred_username claim, rev gpus-forms-backend-00036-89b).
        • Commit 2 SHIPPED 2026-04-30: 10bcd0d (design doc to mkdocs-portal/docs/architecture/forms-phase2.5a-design.md).
        • Commit 3 SHIPPED 2026-04-30: 83f85cb (routes_phase2.py create_submission wire-up, rev gpus-forms-backend-00038-5kg).
        • Verified end-to-end 2026-04-30: API contract response shape, DB persistence (encrypted fields + audit log), MAPLE-side inspection of test submission 0bac9326-4968-477b-8d10-d7d6f457e2a8.
        • Phase 2 SPA submissions now ACTUALLY persist (was theater since Phase 2 cutover).
    2. Wire upload_attachments (routes_phase2.py:128) to GCS bucket gpus-forms-attachments with ClamAV scan and attachment row inserts.
      • STATUS: COMPLETE 2026-05-08 (β closeout — see architecture/forms-phase2.5b-cleanup-closeout.md)
        • 2.5(b) handler shipped 2026-05-08: e17dacb (rev gpus-forms-backend-00041-mcf), three-layer verification PASS on test submission 999cf0cc-… (GCS object + attachments row + audit_log row, all consistent to the microsecond).
        • 2.5(b.cleanup) shipped 2026-05-08: 28964c0 (rev gpus-forms-backend-00043-d9q), 4-source MIME/size truth (prod env vars, Config, module-level shadows, _schema.yaml hint) collapsed to Config.ATTACHMENT_MAX_BYTES + Config.ATTACHMENT_ALLOWED_MIME as single source. Production env vars MAX_UPLOAD_BYTES + ALLOWED_MIME_TYPES removed.
        • Schema migration forms-backend/schema/002_add_submission_deleted_action.sql documents submission_deleted audit_action enum value (added in production during β step 8 to enable orphan-intent cleanup).
        • ClamAV scan handled in Phase 2.5(b.2) — see dedicated section below, COMPLETED 2026-05-19. At β time all attachments landed with clamav_status='pending'; partial index idx_attachments_clamav was already in place for the scanner.
        • Verification gap accepted: end-to-end positive case on rev 00043-d9q blocked by T1.7 SPA bug (silent attachment drop). Pre-cleanup positive cases (5c57b2e6 docx, 62d2fc73 xlsx) plus mechanical-substitution diff correctness accepted as evidence of backend behavior. Rationale in closeout doc.
    3. Wire finalize_submission (routes_phase2.py:151) to:
      • STATUS: COMPLETE — SHIPPED TO PRODUCTION 2026-07-23 (go-live). The routing worker is live: rev f9a114b on MAPLE, override OFF, all 6 ingest addresses verified delivering, HappyFox tickets opening on real submissions (go-live submission 28bd1ebf). Gates 3, 4, and 5 all landed. Locked decisions (from v0.2) as shipped:
        • Transport/shape: B-iii — MAPLE-resident routing worker, Pub/Sub pull subscriber, sends via localhost:25 Postfix reusing the report_mailer.py pattern. Zero new secret. First non-serverless service in the rebuild → carries its own systemd unit + monitoring + IR/DR/drill obligations (see the routing-worker SyslogIdentifier note under Detection).
        • DB role: reuse forms_app (no new role; RLS USING(TRUE) makes the GATE-4 state-coverage work N/A).
        • Attachments: signed-URL delivery in body (inline MIME deferred).
        • Interim body: minimal plaintext, NOT gated behind 2.5(d).
        • Audit: lean — migration 011_routing_audit_actions.sql (4 enum values) + 012 (submission_finalized), both live.
      • Gate close-out (2026-06-10 → 2026-07-23): Gate 3 (C1 GET, C2 routing_result finalize marker, C3 full predicate at finalize, C4 α publishes unconditionally, C5 = migration 012), Gate 4 (routing worker), and Gate 5 (propagation: inventory + coverage-standard portal-propagation section + status Services section + SOC alerting + ops triad + α backfill) all DONE. Residual Gate-5 propagation verification is carried forward as the Portal propagation workstream below (confirm live pages actually render forms LIVE, derived not hardcoded).
      • SELECT actions FROM actions WHERE form_id = <...> ORDER BY action_order
      • For each action, branch by action_type:
        • happyfox_template → render via templates table, POST to HappyFox API (secrets already in Secret Manager: gpus-forms-happyfox-api-key, gpus-forms-happyfox-auth-code)
        • email_template / email_raw → render template, SMTP send via Postfix on MAPLE (currently no SMTP client in backend)
      • Aggregate routing_result into submissions.routing_result JSONB.
    4. Build template render layer (Jinja2 against templates(id, body)).
    5. Build HappyFox API client (configures the integration that SOC dashboard memory #20 noted as "not configured").
    6. Decide policy on email_template / email_raw actions: keep, drop, or redirect through HappyFox? (Per Rajesh: HappyFox API is the operational current path because Google's stricter email policies broke the email-to-HappyFox flow.)
  • Estimated effort: ~~multi-session work~~ DONE. Sequence ran (a)→(b)→(c)→(d) with the routing/HappyFox path shipped at go-live 2026-07-23.
  • Gating: completed the "forms portal actually works" milestone. The 2.5(e) HappyFox API client (native ticket creation, distinct from today's email-ingest path) is now its own workstream below. The SQLi tabletop drill is unblocked functionally but is now GATED on the detection pipeline existing (no point drilling detection that can't fire) — see T1.9.

Phase 2.5(b.2) — α ClamAV scan worker

  • Status: Filed 2026-05-19. COMPLETED 2026-05-19 (5 migrations 004-008 shipped — enum + SQL user + table grants + tight scanner RLS + USING widen + sequence USAGE; pipeline verified end-to-end: 07c3c9e0 fixture + 4a0f5bd5 real upload both clean with audit rows; worker rev gpus-forms-clamav-worker-00002-dwx on Cloud Run, 2Gi, min-instances=0).
  • What shipped: new Cloud Run service gpus-forms-clamav-worker (Pub/Sub-push on gpus-forms-attachments OBJECT_FINALIZE → claim → clamscan → verdict to attachments + audit_log). Scoped Terraform (SA + IAM + topic + DLQ + OIDC push sub; VPN drift untouched, see T3 row). CSR Cloud Build triggers (push + weekly sigrefresh). Design doc alpha-clamav-worker.md v1.2.
  • α.1 / Commit-2 follow-ups: shipped 2026-05-20 in Commit 2 — see ### Phase 2.5(b.2) — Commit 2 immediately below.

Phase 2.5(b.2) — Commit 2 (Slack + GCS tag + DLQ + sweep)

  • Status: COMPLETED 2026-05-20 (worker rev gpus-forms-clamav-worker-00010-gr7 on :c9c2016; migrations 009 + 010 shipped; Terraform 6 resources applied in clamav-worker.tf; EICAR test verified the full infected pipeline end-to-end at 16:45 UTC).
  • What shipped:
    • Slack-post for infected verdicts — placeholder-tolerant fetch from Secret Manager (gpus-forms-clamav-slack-webhook), lazy module-global cache, cold-start dance documented; verified class=wired on first call during EICAR.
    • GCS quarantine metadata tag (quarantined=true, quarantine_reason=<sig>, quarantined_at=<iso8601>) — best-effort with try/except per C4 ordering (tag first, Slack second; audit row is the system of record).
    • /dlq-alert route + DLQ push subscription on gpus-forms-attachment-uploaded-dlq — closes the design §10 gap surfaced at Commit 1 GATE 4. Always returns 200 (no DLQ-of-DLQ).
    • /sweep-stuck route + Cloud Scheduler tick every 15 min — closes α.1 stuck-scanning recovery. Atomic UPDATE … RETURNING, one audit row per flipped id with revision_at_sweep from K_REVISION for deploy correlation, Slack only on flips > 0.
    • Cloud Scheduler API enabled on gpus-infra; clamav-scheduler@gpus-infra.iam SA wired with run.invoker on the worker (resource-scoped) + cloudbuild.builds.editor on the project.
    • Weekly sigrefresh scheduler job (Sun 02:00 UTC) fires the existing gpus-forms-clamav-worker-sigrefresh build trigger — closes the README's "weekly auto-refresh — NOT yet wired" follow-up from Commit 1.
  • Verification: EICAR test at 16:45 UTC. Full pipeline green:
    • event.receivedclaim.okclamscanINFECTED signature=Eicar-Test-Signaturequarantine.tag.okslack.webhook.resolved class=wired (first fetch — lazy init confirmed) → Slack post landed in #us-soc-alerts → audit row attachment_scanned_infected written.
    • All 5 verification checks (worker log, GCS metadata, attachments row, audit_log row, Slack channel visual) passed.
    • Bonus signal: the */15 sweep tick fired naturally at 16:45:03 UTC during the test window, logging sweep.clean count=0 — gate 2d wiring proven live end-to-end without manual stimulus.
    • Test fixture (1 submission + 1 attachment + 1 audit row + 1 GCS object) cleaned up post-validation; attachments now back at pre-test state (clean=2, pending=3).
  • α.1 / Commit 3 follow-ups (hardening / hygiene — NOT Commit-2 blockers):
    • Tighten cloudbuild.builds.editor to per-trigger IAM on just gpus-forms-clamav-worker-sigrefresh (currently project-wide — over-broad blast radius).
    • Rename SLACK_PLACEHOLDER_PREFIX_OKSLACK_VALID_PREFIX (constant name misleads; logic is correct).
    • Normalize _sweep_stuck audit INSERT to use the _AUDIT_ACTION dict (currently hardcodes the enum value as a string literal; PG casts implicitly so it works, but inconsistent with _record).
    • Sigrefresh build trigger needs --included-files=gpus-forms-clamav-worker/** filter — currently fires on every push to main regardless of path, racing with the main trigger and burning a build slot.
    • forms-backend/schema/ 002 ordinal collision (002_rls.sql + 002_add_submission_deleted_action.sql) — rename the second to a higher ordinal.
    • Backfill drain: 3 remaining pending attachments (real user uploads pre-dating Commit-1 worker deployment) still need scanning — carried over from Commit 1; safe to drain via the established byte-correct re-fire pattern, OR let /sweep-stuck pick them up if their uploaded_at is in the stuck window.
    • Commit the ~/terraform/gpus-infra/terraform/clamav-worker.tf Commit-2 additions to that repo (applied to GCP but the .tf change is local-only).
  • T3 candidates filed during this arc (separate from Commit 3):
    • VPN cold-start packet loss — 2nd incident in 7 days; routing to private-IP Cloud SQL (10.34.0.0/24, Private Services Access) failed mid-session.
    • IAP-to-MAPLE 4003 backend-fail — only matters as a backup access path when direct SSH is also down, but should still resolve.

Phase 2.5(a) design doc correction

  • Status: CLOSED 2026-05-07 via 2de57c9. Four corrections applied to mkdocs-portal/docs/architecture/forms-phase2.5a-design.md to match shipped code (commit 83f85cb):

    • _audit_v2 helper signature (kwargs + body)
    • allow-list deny audit action ("auth_failure" not "submission_denied")
    • SubmissionField attachment (session.flush + submission_id, not relationship)
    • second _audit_v2 call site (consistency fix)

    Plus a process-note appendix appended to the doc capturing the read-first-discipline lesson. - Discovered: 2026-04-30 during Phase 2.5(a) Commit 3 implementation. - Context: The design doc at mkdocs-portal/docs/architecture/forms-phase2.5a-design.md was committed in 10bcd0d and contained 3 bugs in the helper specs that Code caught before Commit 3 shipped: 1. _audit_v2 helper signature — design said actor=actor kwarg; correct is actor_username + 4 other Phase 1 fields (target_id, target_type, details, request_id, success). 2. Allow-list deny audit action — design said "submission_denied"; that value isn't in the audit_action ENUM. Correct is "auth_failure" with details={"reason": "not_in_allow_list"} and success=False — matches Phase 1's existing pattern. 3. SubmissionField attachment pattern — design used row.submission = submission (assumes ORM relationship that doesn't exist). Correct is Phase 1's session.flush() + submission_id=submission.id pattern.

Forms Phase 2 UX gaps

  • Discovered: 2026-04-29 PM during Phase 2.5(a) browser regression check.
  • Status: T1.5 sub-section is now FULLY CLOSED — items a/b/c/d all closed. Section can stay in the doc as a closed-finding record or be moved to a "Recently completed" archive at section author's discretion. Item (a) CLOSED 2026-05-07 via f31ec5f. Items (b), (c), (d) CLOSED 2026-04-30 via 2026-04-30 γ phase commits.
  • Priority order within sub-section: a (pulldown — user-visible blocker for form submission accuracy) → d (font — visible polish) → b, c (cosmetic).
  • Items:

    a. Pulldown regression — yes/no booleansCLOSED 2026-05-07 via f31ec5f.

    Root cause: Flask `<string>` route converter (the default) does not match forward slash. Pulldown names containing `/` (e.g. "Grant Funded ? (Yes/No)", "Bargain Type (In/Out)", "Shipping Label/Box Required?") failed to route — Cloud Run URL processing decoded `%2F` back to `/` before route matching, splitting the path into multiple segments that no route consumed. Result: 404 → SPA renders "Could not load options".
    
    Fix: change the `@bp.get` decorator from `"/pulldowns/<name>"` to `"/pulldowns/<path:name>"` — single character change.
    
    Diagnosis chain used DB inspection (via MAPLE access pattern established 2026-04-30): hypothesis 1 (missing pulldown rows) was falsified by data showing all referenced `pulldown_name`s exist in the `pulldowns` table. Pattern noticed in failing data: all 7 failing names contained `/`.
    

    b. COST CENTER duplicate fieldCLOSED 2026-04-30 via 6d9231a (resolved as a side effect of the NO_DATA_TYPES filter — duplicate was a divider/instructions row with that label).

    c. NOTESDIVIDER label leakCLOSED 2026-04-30 via 6d9231a (forms-backend routes/forms.py filters NO_DATA_TYPES from form-detail serializer).

    d. Font darkness / contrastCLOSED 2026-04-30 via f5c79e0 + e58dabc (forms-frontend .field-label color #4a4a44 — WCAG AA at ~9.6:1 contrast over --surface-warm).

Phase 2.1 — Phase 1 Production cutover completion

  • Goal: Finish the Okta Preview → Production cutover that only partially landed in Phase 1's auth/config layer. Phase 1 was originally built against the Okta Preview tenant; the Production cutover on 2026-04-23 left several stale references and shim layers in place. The 5 items below were surfaced 2026-04-27 during the forms-frontend smoke test and consolidate the remaining cleanup. None are user-visible bugs today; they are maintenance traps.
  • Gating: None — Phase 2 SPA is live; cleanup can land at any time.
  • Items (priority order):
    1. forms-backend/auth.py — replace audience=config.OKTA_AUDIENCE with audience=config.OKTA_CLIENT_ID. Remove the OKTA_AUDIENCE Cloud Run env var workaround. Estimated 1 commit, ~5 min.
    2. forms-backend/config.py — change OKTA_ISSUER = f"https://{os.environ.get('OKTA_DOMAIN', 'greenpeaceeu.oktapreview.com')}" to read OKTA_ISSUER directly with default https://greenpeaceeu.okta.com, matching auth_v2.py's pattern. Remove OKTA_DOMAIN env var afterward. Estimated 1 commit, ~10 min.
    3. forms-backend security headers — remove CSP/HSTS/X-Frame-Options/etc. from Flask middleware. nginx (forms-frontend/nginx.conf) is now the sole source for response-time security headers. The duplicate headers don't break anything but create maintenance traps. Estimated 1 commit, ~15 min.
    4. forms-backend CSP — any CSP that must remain in Flask should remove stale greenpeaceeu.oktapreview.com references. Replace with greenpeaceeu.okta.com. Stale since the 2026-04-23 Preview→Prod cutover. Estimated 1 commit, ~10 min.
    5. Consolidate forms-backend/auth.py and auth_v2.py into a single module once both stable. The _key_for_kid retry pattern shipped 2026-04-27 in auth.py (commit d0a7867) carries forward. The v2 module's signin/role pattern carries forward. Defer until after a few weeks of stable production traffic. Estimated 1 PR-sized change.

T1.9 — Forms security exercise program (SQLi tabletop + blue/red drills)

  • Status: BLOCKED — gated on the Detection pipeline (audit_log→SOC and Cloud Run→Wazuh log sinks must exist first; Wazuh rules 100026–100029 are inert by starvation today). No point running a blue-team detection drill against detections that cannot fire.
  • Goal: Exercise detection and response for SQL injection against the live forms backend.
  • Components: SQLi tabletop (60m)blue-team detection drill (90m) against the live endpoint → red-team simulation (90m) adversarial test.
  • Deliverables: Updates to tabletop-playbooks.md and blue-team-drills.md; Wazuh (+ any WAF) rule tuning if gaps surface; executive summary.
  • Owner: Rajesh.

Phase 3 HappyFox integration (forms backend) — SUPERSEDED by 2.5(e)

  • Status: Superseded. The auto-ticket routing shipped at go-live via the routing worker's happyfox_template / email-ingest legs. Remaining native-API work is tracked as 2.5(e) — HappyFox API integration above (queue names from API, ticket-audit migration, per-queue double-ticket deconfliction, API-vs-ingest decision). Kept here for lineage.
  • Historical goal: Wire form submissions to auto-create HappyFox tickets per form's routing config; retry + DLQ on failed ticket creation; ticket ID written back to submission record.
  • Asset docs required (carried into T1.8): iar.md (HappyFox as integrated system); forms IR runbook "HappyFox API outage" + "ticket creation failure → DLQ drain"; blue-team drill entry for "spoofed HappyFox webhook".

Forms portal Phase 1.5 — legacy data migration

  • Goal: Migrate existing legacy-form submissions (where applicable) into the new schema.
  • Gating: Phase 3 HappyFox integration live (so migrated records route correctly on any post-migration edits).
  • Scope: One-time batch import; validate row counts, checksum fields, encrypted columns round-trip correctly; retain legacy source read-only for 90 days post-migration.

Forms portal Phase 1.6 — ON CONFLICT refactor

  • Goal: Refactor upsert logic in gpus-forms-backend to use proper ON CONFLICT clauses rather than check-then-insert race conditions.
  • Gating: Phase 1.5 complete (don't refactor insert paths mid-migration).
  • Scope: Submission insert, audit_log append, pulldown cache refresh.

Okta cleanup — remove localhost redirect URIs

  • Goal: Remove dev-convenience localhost redirect URIs from the Okta Production app (0oavvg1y33wTWFsmP417) once Cloud Run deploy of forms-frontend verifies.
  • Gating: Phase 2 Cloud Run deploy verified — met (forms live in production). Unblocked; can land at any time. Coordinate with the dedicated-forms-app cutover above so the URI cleanup targets the right app.
  • Done-when: Production app redirect URI list contains only https://*.greenpeace.us/* entries. Preview tenant (greenpeaceeu.oktapreview.com, client 0oadhpjktd5UfCMDm0x7) retained as dev fallback.

T1.6 — Forms portal SOC / observability integration

STATUS: NOT STARTED — gap surfaced 2026-05-08 during Phase 2.5(b) scoping discussion

Priority: P2 — should ship after Phase 2.5 functional work completes (2.5b/c/d/e), before SOC Tickets workstream (T4) starts. T4 will assume forms portal events flow into SOC; this workstream makes that true.

Context:

The forms portal (forms.greenpeace.us, gpus-forms-backend, gpus-forms-frontend, Cloud SQL gpus-forms-db) has been operationally invisible to SOC since Phase 1 cutover. While the rest of GPUS infrastructure (WDC servers SKY/RAIN/SUN/WIND, GCP VMs OAK/MAPLE/CEDAR) emits structured events into Wazuh + ELK + Prometheus and is visible across the 15 tabs of soc.greenpeace.us, forms portal emits zero events into that pipeline.

Existing forms-portal capabilities:

  • KMS envelope encryption for sensitive form fields (Phase 2.5a)
  • Audit log table in Cloud SQL with actor_username, actor_ip, target_id, request_id, success
  • IAM-protected Cloud SQL access via service accounts
  • Allow-list authorization on form submission

Observability gaps:

  • No SOC dashboard tab — the 15 tabs at soc.greenpeace.us don't include "Forms"
  • No Wazuh rule coverage for forms-portal events
  • No Cloud Logging → CEDAR/MAPLE pipeline for forms events (audit log lives only in Postgres, not exported)
  • No Prometheus metrics emitted from forms-backend
  • No alert routing for forms-portal anomalies (failed auth bursts, validation errors, malicious upload attempts)
  • No IR runbook (rb-00N-forms-*.md) for forms-portal incidents
  • No DR procedure for forms-portal in drp.md
  • No red/blue drill in tabletop-playbooks.md

Compliance framing:

  • PCI DSS: not applicable (no cardholder data flows through forms portal)
  • NIST 800-53 / 800-171: applicable as best-practice — control families AC (Access Control), AU (Audit & Accountability), SC (System & Communications Protection), SI (System & Information Integrity). Audit log + KMS encryption partially satisfy AU/SC; AC + SI need observability work.
  • MITRE ATT&CK: not a compliance framework, but rest of GPUS infrastructure maps detected events to MITRE techniques. Forms portal events should reach the same taxonomy.
  • OWASP ASVS: directly relevant to the SPA + API surface; needs explicit verification pass.
  • CIS: applies to host hardening (not Cloud Run services directly); covered for forms portal's underlying platform.
  • GPUS internal IRP/DRP framework: forms portal needs runbook + DRP + drill coverage matching what the rest of infrastructure has.

Sub-items (provisional scope — refine when work starts):

a. Audit log → Cloud Logging structured events. Forms-backend audit_log table inserts should also emit structured Cloud Logging entries with proper severity. Pipeline: forms-backend → Cloud Logging → log sink → CEDAR (Elastic) for indexing.

b. Wazuh rule additions for forms-portal events. New rule IDs in the 100020+ range (memory entry on Wazuh ruleset). Cover auth_failure (HIGH severity), attachment_rejected (MEDIUM — potential abuse), submission_created (LOW — informational), attachment_uploaded (LOW). Emit MITRE technique tags where applicable.

c. Prometheus metrics from forms-backend. Standard four (request count, latency, error rate, in-flight) plus domain-specific (submissions_per_minute, auth_failure_rate, upload_size_p99, etc.). Scrape via MAPLE.

d. New Forms tab on soc.greenpeace.us. Tab structure: submission volume, auth-failure rate, attachment activity, slowest queries, recent rejections, threat hunting view (filter audit_log by anomaly patterns).

e. IR runbook rb-006-forms-portal-incident.md. Cover scenarios: compromised submitter account, mass-upload abuse, infected attachment in GCS (forward-look at ClamAV scenario), Cloud SQL unavailability, KMS key rotation, DEK compromise.

f. DR procedure for forms-portal in drp.md. Cover Cloud SQL point-in-time recovery, GCS bucket recovery, Cloud Run rollback pattern, Okta tenant outage fallback.

g. Red/blue drill in tabletop-playbooks.md + blue-team-drills.md. Scenario: malicious attachment uploaded by compromised submitter. Validate detection pipeline end-to-end.

h. OWASP ASVS verification pass. Walk the standard against the forms portal surface; document gaps.

Estimated scope: 3-5 sessions if done as a focused workstream. Could be done incrementally if items a-d are prioritized first (functional observability) and e-h follow (process + verification).

Dependencies:

  • Phase 2.5(b)/(c)/(d)/(e) ideally complete first (gives stable surface to instrument)
  • Wazuh rule slot range coordination (memory entry ranges)
  • Cooperation with T4 SOC Tickets workstream (forms events should auto-ticket)

T1.7 — Forms portal frontend: client-side validation must block submit, not just attachment

STATUS: FILED 2026-05-08. COMPLETED 2026-05-14 (commit f175810 shipped 2026-05-12; browser-verified end-to-end 2026-05-14: oversized blocks submit, wrong-MIME blocks submit, clear re-enables, positive case sub 9ac940c8).

Severity: Medium (silent data loss, recipient-visible).

Scope: Frontend only (forms-frontend SPA).

Bug: When client-side validation rejects an attachment (oversized, wrong MIME, or other triggers), the SPA hides the attachment but does not disable the submit button. Submission proceeds to the backend without the attachment. Recipient team sees a complete-looking submission with no attachment and assumes it was intentional.

Known instances (both cleaned up):

  • d20d2ac8-c19d-48df-a0a4-4f833a750e9b — 2026-05-08 17:15 UTC, csv test, cleaned in β step 8 (audit_log id=10).
  • e93efbc3-4c6e-4b7a-adc9-169d4aedec70 — 2026-05-08 18:09 UTC, docx test post-env-var-removal, cleaned in β closeout (audit_log id=12).

Probable triggers (uncharacterized):

  • Server-side rejection signal (oversized, wrong MIME)
  • Stale /api/config cache showing the old MIME allowlist (frontend may have cached the pre-cleanup 4-MIME list with text/csv and without docx/xlsx)
  • Possibly other (see closeout doc architecture/forms-phase2.5b-cleanup-closeout.md)

Fix shape (TBD by frontend session):

  • Submit button should be disabled while any attachment field has a validation error
  • OR submit handler should check for unresolved attachment errors before POSTing to /submit
  • OR both

Relationship to other T-priorities:

  • Filed below T1.6 (forms portal SOC/observability integration)
  • Independent of α (ClamAV) and 2.5(c)/(d)/(e) tracks
  • Should be addressed before any meaningful production user testing — silent attachment drop produces forensically-confusing audit trails (submission_created with no attachment_uploaded) and recipient-team confusion

Estimated: 1 session.


T1.8 — Forms documentation set (mkdocs)

STATUS: Planned — depends on the security assessment. The assessment (threat model + ASVS) is the upstream artifact several of these records cite; sequence the assessment first, then the derived docs. All new/updated docs live in the mkdocs portal.

Assessment (do first):

  • Threat model for the forms portal surface (SPA + API + routing worker + Cloud SQL + GCS + HappyFox path).
  • OWASP ASVS L2 assessment — walk the standard against the live surface; document gaps.
  • PCI-DSS applicability assessment — expected outcome out of scope (no cardholder data flows through forms); record the determination so it's on file, not assumed.
  • Control mapping: OWASP → NIST 800-53 → MITRE ATT&CK, matching the taxonomy the rest of the estate already uses.

Records / runbooks (derive from the assessment):

  • Forms IR runbook (rb-006-forms-portal-incident.md per T1.6) + DRP entry in drp.md with Cloud Run + Cloud SQL RTO/RPO.
  • Security finding record — the broken-access-control finding, its remediation, and the no-exploitation conclusion (empty audit_log + empty users table, independently corroborated).
  • Governance finding — the API contract documented the Phase-1 decrypt/list/submit endpoints as removed while the code kept serving them; nothing checks that the contract's claims match the code. Record the finding and propose a contract-vs-code conformance check.
  • 2.5(d) as-built — document the subject-template audit-persist behavior as shipped.
  • Regenerate the field-exposure matrix (against the authoritative store — see the DB↔repo drift item; regenerating against the repo is invalid if the repo isn't authoritative).
  • Update mitre-attack.md, threat-vectors.md, pentest-schedule.md, calendar.md.

T2 — Meraki

Meraki cleanup — Sedita site + Meraki SSO

  • Goal: Close out the Meraki P2 follow-ups identified after the P1 inventory (org 395909, 5 networks, 32 devices — completed 2026-04).
  • Gating: T1 forms tier complete.
  • Scope:
    • Sedita site misconfiguration — resolve (specifics to be confirmed at start-of-work)
    • Meraki SSO — currently broken, wire to Okta Production
  • Deliverables: Fixes verified end-to-end (Okta login → Meraki dashboard for an admin test user); Sedita site returns to expected operational state.

Meraki integration — status/SOC/MkDocs coverage

  • Goal: Fold Meraki into the same documentation and monitoring posture as the rest of the estate — Meraki currently lacks matching coverage.
  • Gating: Meraki cleanup above complete.
  • Deliverables:
    • Status site: Meraki org card (device count, online/offline, firmware currency)
    • SOC site: Meraki alerts surfaced (security events, config changes, WAN uplink loss)
    • MkDocs: new architecture/meraki-network.md describing org, networks, devices, admin model, SSO posture
    • iar.md entries for Meraki org + each site
    • Syslog from Meraki → WIND (so Wazuh indexes Meraki events into CEDAR)
    • IR runbook: rb-006-meraki-compromise.md (admin account takeover, rogue config push, AP impersonation)
    • DR procedure: Meraki config backup/restore procedure added to drp.md (Meraki backs up config in-cloud, but document how to roll back + how to replace a bricked device)
    • Red/blue drill: tabletop "Meraki admin credentials leaked" in tabletop-playbooks.md; blue-team detection drill for "unexpected config change outside change window" in blue-team-drills.md

T3 — WDC foundation

SKY portal-backup cron died (~2026-04-17) — ADJACENT (non-forms)

  • Status: Filed. Not a forms item, but an open backup-coverage gap worth surfacing alongside the WDC work.
  • What: The gpus-portal-backup.sh cron on SKY (nightly 02:30 → GCS, shipped 2026-03; see Completed) stopped running ~2026-04-17. portals/ snapshots stop at that date. Server backups are current — only the portal-content snapshot stream is affected.
  • Do: Root-cause why the cron stopped, restore it, and backfill the snapshot gap (≈Apr 17 → now). Note this compounds the T5-EXPANDED P0 finding that soc-site + forms-frontend were already unchecked by portal-backup coverage — restoring the cron should also extend coverage to those.

ESXi inventory & cleanup

  • Goal: Bring the ESXi hypervisor — currently undocumented — under the same documentation, monitoring, and IR posture as everything else.
  • Gating: T1 + T2 complete.
  • Deliverables:
    • Inventory: ESXi version, licensing, VMs hosted, networking, hardware health, management plane exposure
    • Confirm or sever ESXi ↔ NAS coupling (decision: is the NAS a datastore, a backup target, or both?)
    • Cleanup: disable unused accounts, rotate admin credentials, enable syslog → WIND
    • Reconfigure: NTP, DNS, timezone, email alerts → gpus-it-security@greenpeace.org
    • Monitoring: Prometheus scrape via vmware_exporter, Grafana dashboard, Wazuh agent on guest VMs where feasible
    • Status site: ESXi card (version, uptime, VM count, datastore usage)
    • SOC site: ESXi in asset coverage map
    • MkDocs: architecture/wdc-hypervisor.md
    • wdc-hostregistry.csv entry; iar.md entry
    • IR runbook: rb-007-esxi-compromise.md (hypervisor takeover, guest escape, management plane breach)
    • DR procedure: ESXi host failure recovery in drp.md
    • Red/blue drill: tabletop "ESXi vCenter creds leaked" in tabletop-playbooks.md; blue-team detection drill "unexpected VM clone / snapshot export" in blue-team-drills.md
  • Known risk: ESXi 6.7 is already flagged as EOL in tracker.md (VLN-004). Inventory may surface the need to accelerate hypervisor replacement — if so, that becomes its own T-tier item.

Synology NAS inventory & cleanup

  • Goal: Same as ESXi above, for the Synology NAS.
  • Gating: ESXi inventory complete (likely coupled — NAS may be serving as an ESXi datastore, which affects cleanup sequencing).
  • Deliverables:
    • Inventory: model, firmware, volumes, shares, users, backup targets, relationship to ESXi
    • Cleanup: disable unused accounts, rotate admin credentials, enable SNMP + syslog → WIND
    • Reconfigure: NTP, DNS, timezone, email alerts
    • Monitoring: Prometheus scrape via SNMP, Grafana dashboard
    • Status site: Synology card (volume health, SMART, firmware)
    • SOC site: in asset coverage map
    • MkDocs: architecture/wdc-nas.md
    • wdc-hostregistry.csv entry; iar.md entry (classification, owner, retention, criticality)
    • IR runbook: rb-008-nas-compromise.md (ransomware on shares, credential theft, firmware tampering)
    • DR procedure: NAS failure + volume rebuild in drp.md
    • Red/blue drill: tabletop "Synology admin portal exposed" in tabletop-playbooks.md; blue-team detection drill "mass file encryption on shares" in blue-team-drills.md

GCP terraform tree has no VCS (supersedes "WDC VPN/route TF state drift")

  • Goal: Initialize version control on ~/terraform/gpus-infra/terraform/ (currently untracked on rchhetry's Mac, single point of failure) so terraform changes can be reviewed, rolled back, and "what's the source of truth" has an answer.
  • Discovered: 2026-05-21, during α ClamAV Commit 2 close-out (A2). git -C ~/terraform/gpus-infra/terraform status returned fatal: not a git repository, and exhaustive .git search across /Users/rchhetry confirmed no repo contains these .tf files. The previously-filed "WDC VPN/route Terraform state drift" item assumed a remote repo existed to drift from — the framing was wrong; the no-VCS problem is the parent.
  • Acute risk mitigation (in effect 2026-05-21): ~/Downloads/terraform-snapshots/2026-05-21/ holds md5-verified copies of all 11 .tf files (no tfvars/tfstate/tfplan copied — those need secrets review first). Local-disk redundancy only; not version control.
  • Inherited drift (still real, blocked until VCS exists): terraform plan against live state shows google_compute_vpn_tunnel.wdc_tunnel local_traffic_selector forcing replacement; google_compute_route.onprem_mgmt + google_compute_route.onprem_prod must be replaced as dependents; google_compute_instance.{cedar,maple,openvas} in-place updates. Replacing tunnel + routes would tear down the WDC↔GCP site-to-site VPN (SKY/RAIN DNS-DHCP, SUN Prometheus, WIND ELK). Interim mitigation: all clamav-worker Terraform applied -target-scoped so the drift is never actioned.
  • Deliverables for the dedicated session:
    • Secrets review of terraform.tfvars + terraform.tfstate* + tfplan: enumerate values, identify what must NOT enter VCS.
    • Design .gitignore: minimum terraform.tfvars, *.tfstate*, tfplan, .terraform/, plus anything from secrets review.
    • git init in ~/terraform/gpus-infra/terraform/; first commit of the 11 .tf files (and any safe-to-commit ancillary files).
    • Decide remote: Cloud Source Repositories (matches existing convention for gpus-infra-portals) vs private GitHub.
    • Push initial commit; document the remote in the worker README.
    • With VCS in place: drift triage — root-cause the VPN tunnel local_traffic_selector mismatch, decide reconcile direction (update Terraform to match live, OR plan a maintenance-window apply that recreates tunnel/routes), if recreation: scheduled change window with WDC-connectivity-loss comms.
    • Remove the -target workaround note from the clamav-worker README once unscoped apply is safe.
  • Gating: independent of WDC inventory work; should run before any further unscoped terraform apply. Best-suited to a dedicated session — touching secrets + remote setup + drift reconcile + maintenance window planning needs full attention.

T4 — Queued

SOC Ticketing tab

  • Goal: Replace ad-hoc alert triage with tracked tickets on soc.greenpeace.us.
  • Gating: T3 complete (stable asset inventory before we wire ticketing to it).
  • Sources: Wazuh (level ≥ 10), Prometheus alertmanager, AIDE change alerts, Fail2ban bans, OpenVAS critical/high.
  • Dedup: 5 min window on (source, rule_id, host).
  • SLA: Critical = 15 min ack / 4 hr resolve · High = 1 hr ack / 24 hr resolve · breach → Slack #soc-alerts + email.
  • Asset docs required: iar.md update; IR runbook rb-009-soc-ticketing-outage.md; blue-team drill for "silent alert drop" (ingestion pipeline broken but tickets still showing green).

Vendor Access Portal (replace SFTP)

  • Goal: Zero-trust replacement for the current SFTP vendor drop.
  • Gating: SOC Ticketing in place (so vendor-portal anomalies ticket correctly from day one).
  • Controls: Vendor IP whitelist via Cloud Armor, signed expiring URLs (max 72h), full audit trail, per-vendor bucket prefixes, ClamAV scan before internal consumption.
  • Asset docs required: iar.md; IR runbook rb-010-vendor-portal-abuse.md (stolen signed URL, vendor account compromise); blue-team drill for "vendor credential used from unexpected geography".

T5 — Backlog

Status site automation — Phase A

  • Goal: Eliminate hardcoded values in status-site/index.html.
  • Scope: Cloud Run service list via gcloud run services list at render time; server count via servers.py; cost block labelled "last updated YYYY-MM-DD" (still manual this phase, but honest about staleness).
  • Estimated effort: 2h.

MySQL decommission

  • Goal: Retire the legacy MySQL instance. Remaining dependencies to be confirmed during inventory.
  • Gating: Legacy migration path confirmed (see Forms Phase 1.5 outcome — likely overlap).
  • Asset docs required: iar.md removal; drp.md update to drop MySQL recovery procedure; final backup captured and sealed in Coldline GCS with 7yr retention lock before shutdown.
  • 2026-06-10: legacy in_formfeed MySQL found on PUBLIC IP 34.171.123.238 — ownership unconfirmed. Decommission precondition: drift check via from_mysql.py --dry-run; authorized-networks check needed.

Status site automation — Phase B

  • BigQuery billing export, live current-month spend, trend, forecast.
  • Requires GPI budget approval for BigQuery storage + query cost (est. <$5/mo).

Status site automation — Phase C

  • Per-service cost attribution, budget alerts.
  • Depends on Phase B (BigQuery export) landing.

T5-EXPANDED — Portal static-debt retirement

  • Filed: 2026-06-10, per read-only audit of both portals (status + SOC).
  • Sequencing: forms 2.5(c) Gate 5 has now landed (go-live 2026-07-23), so this is no longer gated behind it. P0 truth-fixes may interleave with the forms cutover/hardening workstreams; the remainder queues behind them.
  • Phases:
    • P0 — truth-fixes:
      • SOC posture fail-open bug: unreachable host renders green "Compliant"; hardcoded auditd/SELinux/firewall columns.
      • Stale-wrong Exec risk register — DRP/IRP marked "not documented" but both exist.
      • "All 8 services" undercount.
      • Reports last_generated always None.
      • Portal-backup coverage gap: soc-site + forms-frontend unchecked.
    • P1 — wire already-served data: server cards from /api/status; discarded /api/carbon; /api/reports fields; Governance link-out.
    • P2 — author missing canonical sources: defense-in-depth.md, threat-model.md, risk-register.yaml (status & SOC registers currently DISAGREE), compliance-scores YAML — then build-time render runbooks/redblue/compliance (fixes SOC missing rb-006/007).
    • P3 — new collectors: VPN, DNS serial, Prometheus range charts, posture score, FLEET→inventory, Cloud Run table from inventory (overlaps forms Gate 5).
    • P4: BigQuery billing (= old Status-site Phase B, GPI-gated), control matrix, git-log audit trail.
  • Standing rule (recorded 2026-06-10): all portal content must be LIVE or DERIVED — never hardcoded. The coverage standard is to be extended to require that elements RENDER FROM sources, not merely that endpoints exist — the audit found backends over-serving and frontends discarding (e.g. /api/carbon fetched then thrown away).

Cross-cutting / lessons learned

Cloud Run env vars + Cloud Build deploys

Cloud Run env vars set out-of-band via gcloud run services update --update-env-vars are wiped on every Cloud Build deploy IF the cloudbuild.yaml uses --set-env-vars (destructive replace) instead of --update-env-vars (merge). Discovered 2026-04-27 in forms-backend/cloudbuild.yaml — fixed in commit 1ca4461. Other three backends (status, soc, security) don't have this bug because their cloudbuild.yaml files don't pass any env-var flag.

Lesson: any new backend cloudbuild.yaml should either omit env-var flags entirely (preserve) or use --update-env-vars (merge). Never --set-env-vars unless the deploy is intentionally the source of truth for ALL env vars.

Cloud SQL access for developer-side diagnostics

gpus-forms-db is private-IP only (10.34.0.3). Reaching it requires presence inside the gpus-infra VPC. Tested 2026-04-28:

  • Laptop direct: blocked (no VPN to forms-db's service-peering range)
  • Cloud Shell + cloud-sql-python-connector: blocked (timeout to private IP from Cloud Shell's managed network)
  • Cloud Shell + cloud-sql-proxy --private-ip: blocked (proxy bound locally fine but the dial to 10.34.0.3:3307 timed out)

Workable paths for developer-side diagnostic queries:

  • SSH into MAPLE/OAK/CEDAR (all in gpus-infra VPC), run script there. Caveat: that VM's service account principal needs Postgres-side grants (CLOUD_IAM_USER + SELECT).
  • Add VPC peering between Cloud Shell's project network and gpus-infra (administrative work, deferred).

Lesson: Phase 2.5 implementation work needs MAPLE-based or peered DB access established as a prerequisite. inspect_actions.py (forms-backend/migrate/, committed 9eecc35) is ready to run from any in-VPC environment.

Memory entries describing "current bugs" age fast

Diagnostic memos written during one session describe state at that moment, not current state. The 2026-04-22 count-drift memo described a real bug in auth.py — fixed one day later in commit dd810e0 — but the memo persisted in memory and led to a Phase 2.5(a) design draft that proposed re-fixing the already-fixed bug. Code surfaced the staleness by reading current auth.py before any edits.

Lesson: Read-first discipline includes reading current code, not just current memory. When a memory entry describes a bug, verify the bug still exists by reading the affected module's current state (and grep recent commits for likely fix language).

Design doc accuracy under read-first discipline (2026-04-30)

Three design-doc bugs were caught by Code's "STOP and tell me if anything doesn't fit" gate before Commit 3 shipped:

  • _audit helper signature mismatch with actual AuditLog model
  • Invented audit_action enum value not in schema
  • Assumed SQLAlchemy relationship that wasn't declared

All three would have crashed the handler at runtime (or worse, the third would have silently produced submissions with NO field rows).

Lesson: design docs that "mirror Phase 1 patterns" must be drafted from current reads of those patterns, not from memory of earlier reads. The three failures had a common root: the doc described what was remembered of Phase 1's behavior, not what Phase 1 actually does at the line number the design doc claims to mirror. Verifying the source pattern at draft time would have caught all three.

Read-first discipline doesn't end at "read the file once during investigation." It applies again at every implementation moment that references that file's content.

MAPLE access for ad-hoc DB queries (2026-04-30)

Established and verified working pattern for ad-hoc Cloud SQL inspection from MAPLE (Phase 2.5 implementation work depends on this for diagnostic queries):

  • SSH user: cloudadmin (NOT monitadmin, which is SUN/WIND only).
  • MAPLE has cloud-sql-proxy v2 + psql pre-installed.
  • MAPLE does NOT have python3.11 or git — Python connector path requires sudo dnf install.
  • Postgres user maple-agent@gpus-infra.iam exists with SELECT grants on submissions, submission_fields, audit_log (no GRANT needed).

Standard one-liner:

ssh cloudadmin@maple "cloud-sql-proxy --auto-iam-authn --private-ip \
  gpus-infra:us-central1:gpus-forms-db &" && sleep 5 && \
psql "host=127.0.0.1 port=5432 dbname=gpus_forms user=maple-agent@gpus-infra.iam sslmode=disable" \
  -c "<query>"

inspect_actions.py at forms-backend/migrate/ can run from MAPLE only if python3.11 + git are installed first. Defer Python install until there's a real need beyond what proxy + psql handles.

Hypothesis falsification via empirical data (2026-05-07)

α phase pulldown regression demonstrated a clean hypothesis-disconfirmation chain:

  1. Symptom: yes/no pulldowns failing "Could not load options"
  2. Initial hypothesis: missing pulldown rows in DB
  3. Hypothesis falsified: DB query showed all 7 failing pulldown_names exist in pulldowns table with proper ["Yes", "No"] values
  4. Pattern in failing data: all 7 names contained /
  5. New hypothesis: Flask string converter doesn't match /; URL decoding splits the path
  6. Fix: change <name> to <path:name> in route decorator
  7. Verified: dropdown opens with Yes/No options

Lesson: when a fix idea seems obvious ("add the missing pulldown rows"), check the data first. The obvious fix shipped without the falsification step would have been a no-op (rows already exist) and the bug would persist with diagnostic time wasted plus user trust diminished.

This pattern is reusable: when an outage looks like "data is missing," verify by query before shipping the inverse ("add the data"). The reverse pattern — when data exists but isn't being read — points at a different layer (routing, auth, encoding, serialization).

Operational learnings — 2026-06-10

  • L2TP stale tunnel, 3rd occurrence: tunnel reports "up" while the path is dead. Health checks must probe internal IPs, not tunnel state.
  • FortiClient utun6 route confound: FortiClient's interface can shadow routes and confuse VPN path diagnosis — rule it out before blaming the site-to-site tunnel.
  • IAP 4003 = host-level: the IAP-to-MAPLE 4003 backend-fail is host-level, not IAP edge. Promotes the T3 item to IAP-as-break-glass framing (see Summary).
  • Legacy in_formfeed MySQL on PUBLIC IP 34.171.123.238: ownership unconfirmed. Decommission item filed under T5 MySQL decom with drift-check precondition (from_mysql.py --dry-run); needs authorized-networks check.


Program work register — schema v1 (opened 2026-08-13)

This section uses the schema below. Nothing above it does, and nothing above it was changed. The tiered tables, the T1–T5 sections, the Completed table and the change log keep their original Planned / In Progress / Blocked / Done / Deferred / Filed vocabulary and their original wording. No pre-existing row in this document was edited, reworded, re-statused or removed — including the rows marked Done that predate the status rule. Those are flagged, not changed, in Flag List — Unverified Done Rows; re-adjudication is a later pass tracked as GOV-008.

Schema: id · title · register · status · priority · owner · evidence · acceptance test · blocks · blocked_by · date_raised · date_verified.

Status vocabulary for this section only: open · in progress · blocked · done. done means verified end-to-end via the real production path against live state — authored, present, loaded or validated is not done. No row in this section is done.

evidence records what was read, where, and on what date. acceptance test states in advance the observable condition that closes the row. A row missing either is marked ⚠ INCOMPLETE on that field and is not filled with a plausible guess. Most rows in this section are incomplete on the acceptance test; see Incomplete rows.

Two ID series

Series Meaning
PRG- Program work items — belong to the program queue
HB- Host-level blockers — belong to a host, not to the program

HB- rows are the four items that block a specific host's disposition. They are deliberately not filed as program items: filing them there would make the whole disposition pass appear blocked on four unrelated questions, when in fact each blocks exactly one host.

⚠ The host rows these four attach to do not exist yet. gannet, emu, ostrich, catbird and phoebe appear nowhere in inventory.yaml (checked 2026-08-13, zero matches each). Only catbird and phoebe appear anywhere under docs/ at all, and only as /32 allowlist entries at compliance/iar.md:381-382. There is therefore no host row to attach an HB- blocker to. Each HB- row below names its host in blocks and is parked until PRG-001/PRG-002 create the host rows. This is the one place the instruction "attach them to the host rows they block" could not be carried out as written, and the reason is a missing anchor, not a judgement call.

Index — program work items

ID Title Status Priority Blocked by
PRG-001 Stage 1c — enumerate Meraki org 395909 and all three Synology units open not assessed
PRG-002 Stage 2 — diff enumerated.yaml vs inventory.yaml, then disposition pass blocked not assessed PRG-001
PRG-003 Document the previously unlisted Cloud Run services open not assessed
PRG-004 Identify 2 undocumented Cloud SQL instances in gpusa open not assessed
PRG-005 Scope 18 buckets in gpus-it against the "empties and retires entirely" commitment open not assessed
PRG-006 Unattached disks and 47 static IPs — billing impact open not assessed
PRG-007 35 orphan DNS names + 18 orphan Puppet node definitions open not assessed
PRG-008 phoenix into inventory.yaml and under a liveness check open not assessed GOV-004 (derived)
PRG-009 desert, river, star — three powered-off VMs on flower, disposition unknown open not assessed
PRG-010 water hosts zero VMs and is mis-documented as hosting ocean open not assessed
PRG-011 ~111 workstation DHCP reservations against 15 active leases open low — "not urgent" (as supplied)
PRG-012 Billing export — highest-value cost action, not yet pulled open highest-value cost action (as supplied)

Index — host-level blockers

ID Title Blocks Status Priority
HB-001 Tamr support on Rocky 9 — vendor question gannet open not assessed
HB-002 GPI consumers of us.gl3? emu / ostrich rebuild open first on the critical path (as supplied)
HB-003 sssd/nslcd check across gpusa catbird delete open not assessed
HB-004 phoebe's real consumers phoebe disposition open not assessed

PRG-001 — Stage 1c: enumerate Meraki org 395909 and all three Synology units

Field Value
id PRG-001
title Stage 1c — enumerate Meraki org 395909 and all three Synology units
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence Partial. Supplied: estate total 455, which "currently excludes network edge and storage layer entirely". The figure is specific but its source artifact and read date were not given. To complete: the Stage 1 output the 455 comes from, with its date.
acceptance test INCOMPLETE — none supplied. "Enumerate X" states the work, not the observable condition that closes it. Not guessed. To complete: the enumeration artifact containing a stated count of Meraki devices for org 395909 and all three Synology units.
blocks PRG-002
blocked_by
date_raised 2026-08-13
date_verified

Related: VLN-021 covers the Meraki org's admin access levels; this row covers its device inventory. Same org (395909), different question.


PRG-002 — Stage 2: diff enumerated.yaml against inventory.yaml, then disposition pass

Field Value
id PRG-002
title Stage 2 — diff enumerated.yaml against inventory.yaml, then run the disposition pass
register Program work items
status blocked
priority not assessed — none supplied
owner unassigned
evidence INCOMPLETE — none supplied. No artifact, path or date was given for this row. Not guessed.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks HB-001, HB-002, HB-003, HB-004 — this pass is what creates the host rows those four attach to
blocked_by PRG-001 (stated)
date_raised 2026-08-13
date_verified

PRG-003 — Document the previously unlisted Cloud Run services

Field Value
id PRG-003
title Document the previously unlisted Cloud Run services: gpus-security-backend, gpus-soc-site, gpus-status-backend — and reconcile the status of gpus-forms-clamav-worker
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence Confirmed first-hand 2026-08-13 against inventory.yaml (repo root): gpus-security-backend 0 matches, gpus-soc-site 0 matches, gpus-status-backend 0 matches — all three genuinely absent. gpus-forms-clamav-worker is present, as cloud_services.gpus_forms_clamav_worker (underscore form). The cloud_services block holds nine entries, all forms/ClamAV-related.
acceptance test INCOMPLETE — none supplied. Not guessed. Note a bare "present in inventory.yaml" would be a weak test here: GOV-003 establishes that the cloud-services-render sentinel satisfies portal presence for cloud_services entities unconditionally, so an entry can be added and pass coverage without any portal actually rendering it.
blocks
blocked_by
date_raised 2026-08-13
date_verified

Correction to the item as supplied. It was filed as "the 4 previously unlisted Cloud Run services" including gpus-forms-clamav-worker. That service is listed. The count is 3 unlisted, not 4. The item is recorded with the corrected count rather than the supplied one, because the supplied count is checkable and wrong; nothing was merged or reworded beyond that.


PRG-004 — Identify 2 undocumented Cloud SQL instances in gpusa

Field Value
id PRG-004
title Identify 2 undocumented Cloud SQL instances in gpusa; retention and backup obligations unknown
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence INCOMPLETE — none supplied. The count (2) and the project (gpusa) are given, but no listing, command output or read date. Not guessed. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/gpusa-it-infrastructure-306400__sql__rchhetry-at-greenpeace.org.json. Confirmed: 2 Cloud SQL instances in gpusa.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

"Retention and backup obligations unknown" is the material part: an undocumented database is also an unclassified one. Note VLN-017gcloud list verbs exit 0 on permission denial — means any enumeration of gpusa that did not check stderr may have under-reported.


PRG-005 — Scope 18 buckets in gpus-it against the "empties and retires entirely" commitment

Field Value
id PRG-005
title Scope the 18 buckets in gpus-it against the "empties and retires entirely" commitment
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence Partial. Count (18) and project (gpus-it) supplied; no listing or read date. The "empties and retires entirely" commitment is referenced but its source document was not cited. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/gpus-it-infrastructure__buckets__rchhetry-at-greenpeace.org.json. Confirmed: 18 buckets in gpus-it.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

PRG-006 — Unattached disks and 47 static IPs; billing impact

Field Value
id PRG-006
title Unattached disks (gpus-infra 7 disks : 3 instances; gpus-it 15 : 9) and 47 static IPs — billing impact
register Program work items
status open
priority not assessed — none supplied. (Billing impact is stated as a consequence, not as a priority.)
owner unassigned
evidence Partial. Ratios and counts supplied and specific — gpus-infra 7 disks against 3 instances, gpus-it 15 against 9, and 47 static IPs — but no listing, command output or read date. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): All four figures confirmed. raw/gpus-infra__disks__*.json = 7 disks against raw/gpus-infra__instances__*.json = 3 instances; raw/gpus-it-infrastructure__disks__*.json = 15 against 9 instances; and the three *__addresses__*.json captures total 47 static IPs across gpus-infra, gpus-it-infrastructure and gpusa-it-infrastructure-306400.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

The billing impact here is quantifiable but not yet quantified — PRG-012 (billing export) is what would size it. No dependency is asserted between them because none was stated.


PRG-007 — 35 orphan DNS names and 18 orphan Puppet node definitions

Field Value
id PRG-007
title 35 orphan DNS names in cloud.us.gl3 plus 18 orphan Puppet node definitions — likely one uncleaned decommission wave
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence Partial. Counts supplied (35 DNS names, 18 Puppet node definitions) with the zone named; no dump, command output or read date. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task1-arecord-crossref-20260813.txt (captured 2026-08-13) for the DNS side, and raw/puppet__nodes-parsed__phoenix-ssh-rchhetry.json for the node definitions. See GOV-018 — the 35 figure this row inherits was superseded twice and now stands at 50.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

"Likely one uncleaned decommission wave" is recorded as supplied — it is a hypothesis, not a finding, and the row does not treat it as established. Distinct from VLN-020, which covers five node definitions that can never match a certname; that is a naming defect, this is orphaned scope. The two may overlap and neither row assumes it.


PRG-008 — phoenix into inventory.yaml and under a liveness check

Field Value
id PRG-008
title Bring phoenix into inventory.yaml and under a liveness check
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence Partial — supplemented first-hand. No source or date was supplied. Confirmed 2026-08-13: phoenix has zero matches in inventory.yaml, consistent with the item as filed.
acceptance test phoenix present in inventory.yaml and covered by a liveness check. (Derived from the item's own stated target state, not invented — but see the blocker below, which means the second half is not currently satisfiable.)
blocks
blocked_by GOV-004derived, not supplied. GOV-004 records that no GL5 liveness monitoring exists; the second half of this acceptance test cannot be met until it does. Flagged as a derived dependency so it can be rejected on review.
date_raised 2026-08-13
date_verified

PRG-009 — desert, river, star: three powered-off VMs on flower, disposition unknown

Field Value
id PRG-009
title desert, river, star — three VMs discovered on flower, powered off, no DNS, no DHCP presence; disposition unknown
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence Partial. Four specific observations supplied — discovered on flower, powered off, no DNS, no DHCP presence — but no source artifact or read date. Confirmed first-hand 2026-08-13 that none of desert, river or star appears in inventory.yaml (zero matches each). Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-esxi-inventory-20260813.txt, the same authenticated vSphere capture that carries the VLN-013 build numbers.
acceptance test INCOMPLETE — none supplied. "Disposition unknown" states the open question, not the condition that closes it. Not guessed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

Note the host these sit on: flower is the hypervisor VLN-013 identifies as running the 2018 GA 6.7 build that has never been patched.


PRG-010 — water hosts zero VMs and is mis-documented as hosting ocean

Field Value
id PRG-010
title water hosts zero VMs and is mis-documented as hosting ocean; capture CPU, RAM and datastore capacity to test consolidating fire and flower onto water
register Program work items
status open
priority not assessed — none supplied
owner unassigned
evidence Confirmed first-hand 2026-08-13 in inventory.yaml:146-152. The water entry reads desc: First on-prem ESXi host — currently hosts Ocean (KACE SMA) (:149) and vms: [ocean] (:152) — the mis-documentation is in the inventory itself, not only in prose. ocean is separately defined at inventory.yaml:988 with fqdn: ocean.wdc.us.gl3. The same entry also carries hardware_model: "" # TODO: walk-around fill-in (:151), i.e. the CPU/RAM/datastore capture this row calls for has an existing empty slot waiting for it.
acceptance test CPU, RAM and datastore capacity captured for water, sufficient to test consolidation of fire and flower onto it. (Stated in the item; recorded as given.)
blocks
blocked_by
date_raised 2026-08-13
date_verified

Second occurrence of the VLN-013 misattribution, found while evidencing this row. inventory.yaml:149 also records hypervisor: VMware ESXi 6.7 for water. Per VLN-013's authenticated vSphere API evidence, water is on 8.0.3 (build 24022510). VLN-013's acceptance test names only VLN-004 and will not catch this line. Raised here rather than silently widening VLN-013's acceptance test; see the open question in the session log entry for 2026-08-13.


PRG-011 — ~111 workstation DHCP reservations against 15 active leases

Field Value
id PRG-011
title ~111 registered workstation DHCP reservations against 15 active leases — ghost-asset exposure
register Program work items
status open
priority low — recorded from the supplied qualifier "not urgent". This is the only priority signal given for any PRG- row and is carried as stated, not converted to a tier.
owner unassigned
evidence Partial. Counts supplied (~111 reservations, 15 active leases); the ~ is carried through rather than resolved. No source artifact or read date. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task3-dhcp-parsed-20260813.txt and raw/dhcp-static-reservations-sky-20260813.txt. The lease side confirms exactly: the sky census reads active 15 against backup 162, free 88, TOTAL 265 distinct IPs. ⚠ Discrepancy — figure did not reproduce. The static-reservation capture contains 126 host declarations on sky, not the ~111 stated. 126 includes infrastructure hosts — the first stanza is host sky itself — so ~111 is presumably workstations after filtering, but the filter is not recorded. The stated figure is left unchanged pending that definition.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

PRG-012 — Billing export: highest-value cost action, not yet pulled

Field Value
id PRG-012
title Billing export — flagged as the highest-value cost action, not yet pulled
register Program work items
status open
priority highest-value cost action — recorded as supplied. Carried verbatim rather than mapped to a tier, since no tier was given.
owner unassigned
evidence INCOMPLETE — none supplied. No artifact, path or date. "Not yet pulled" is a state, not evidence of who established it or when. Not guessed.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

Note the existing backlog rows Status site automation — Phase B (BigQuery billing export) and Phase C (per-service cost attribution) sit at T5 with 2026-10 / 2026-11 targets. This row is filed as the highest-value cost action. That tension is recorded, not resolved — reconciling the two would mean re-statusing a pre-existing row, which this pass does not do.


Host-level blockers

Each row below blocks exactly one host. None is a program blocker. The host rows they attach to do not exist yet — see the warning at the top of this section — so blocks names the host and its pending disposition rather than a row id.

HB-001 — Tamr support on Rocky 9

Field Value
id HB-001
title Tamr support on Rocky 9 — vendor question
register Host-level blocker
status open
priority not assessed — none supplied
owner unassigned
evidence INCOMPLETE — none supplied. Not guessed.
acceptance test INCOMPLETE — none supplied. A vendor answer is the implied output, but the condition that makes it sufficient was not stated. Not guessed.
blocks gannet — host row does not exist; created by PRG-002
blocked_by
date_raised 2026-08-13
date_verified

HB-002 — GPI consumers of us.gl3?

Field Value
id HB-002
title GPI consumers of us.gl3?
register Host-level blocker
status open
priority first on the critical path — recorded as supplied
owner unassigned
evidence INCOMPLETE — none supplied. Not guessed.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks emu / ostrich rebuild — host rows do not exist; created by PRG-002
blocked_by
date_raised 2026-08-13
date_verified

This is the highest-priority signal supplied anywhere in the program list, and it is attached to a host row that does not yet exist. Both emu and ostrich are the DNS masters carrying the VLN-015 NOPASSWD sudo finding.

HB-003 — sssd/nslcd check across gpusa

Field Value
id HB-003
title sssd/nslcd check across gpusa
register Host-level blocker
status open
priority not assessed — none supplied
owner unassigned
evidence INCOMPLETE — none supplied. Not guessed.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks catbird delete — host row does not exist; created by PRG-002. catbird currently appears only as a /32 allowlist entry at compliance/iar.md:382, marked "Remove after Phase 4 cutover"
blocked_by
date_raised 2026-08-13
date_verified

HB-004 — phoebe's real consumers

Field Value
id HB-004
title phoebe's real consumers
register Host-level blocker
status open
priority not assessed — none supplied
owner unassigned
evidence INCOMPLETE — none supplied. Not guessed. Evidence now available — phoebe-capture/ under ~/estate-enum-2026-08, captured 2026-08-13. This row was INCOMPLETE on evidence; the directory holds Apache vhost configs and access/error logs for six candidate consumers: budget.us.gl3, forms.us.gl3, lam.cloud.us.gl3, phoebe.cloud.us.gl3, redirects.cloud.us.gl3 and sharperlight.us.gl3, plus messages-tail.txt. forms.us.gl3 carries the largest logs — four weekly rotations through 2026-07 totalling ~5 MB — and lam.cloud.us.gl3 has rotations running back to 2025-12. This is raw material, not an answer: the logs have not been analysed, so which of the six are live consumers versus redirect stubs is still open, and the acceptance test remains INCOMPLETE.
acceptance test INCOMPLETE — none supplied. Not guessed.
blocks phoebe disposition — host row does not exist; created by PRG-002. phoebe currently appears only as a /32 allowlist entry at compliance/iar.md:381, marked "Remove after Phase 4 cutover"
blocked_by
date_raised 2026-08-13
date_verified

GOV-004 names this host as the origin of the missing-reconciliation-control class ("the phoebe mechanism"). No blocks relation is asserted between them because none was stated.


Incomplete rows — program register

Recorded as incomplete rather than completed with a plausible guess.

ID Evidence Acceptance test
PRG-001 ⚠ partial — 455 figure, no source or date missing
PRG-002 missing missing
PRG-003 ✅ confirmed first-hand missing
PRG-004 missing missing
PRG-005 ⚠ partial — counts, no source or date missing
PRG-006 ⚠ partial — counts, no source or date missing
PRG-007 ⚠ partial — counts, no source or date missing
PRG-008 ⚠ partial — supplemented first-hand ✅ derived from the item's stated target
PRG-009 ⚠ partial — supplemented first-hand missing
PRG-010 ✅ confirmed first-hand ✅ as supplied
PRG-011 ⚠ partial — counts, no source or date missing
PRG-012 missing missing
HB-001 missing missing
HB-002 missing missing
HB-003 missing missing
HB-004 missing missing

13 of 16 rows lack an acceptance test. 6 of 16 lack evidence entirely. The pattern is consistent and worth naming: the program list was supplied as a list of observations and questions, which is what it is good at. Acceptance tests are decisions about what "finished" means, and those had not been taken for most of these items. Every one is recoverable — the fastest route is one pass stating the close condition per row.


Program work register — 2026-08-17 load (PRG-013 – PRG-026)

Continues the schema v1 section above. Same schema, same status vocabulary, same rule: done means verified end-to-end via the real production path against live state. No row in this load is done, including PRG-019, which records a disposition decision rather than a verification.

Source: priorities/backlog-2026-08-17.md v1.0, items B-25 – B-38. Owner is R. Chhetry throughout except PRG-015, where the source text names Rob MacMillan as co-decider. No severity or priority was stated for any row in this load, so every row reads not assessed; none is inferred. No acceptance test was stated for any of the fourteen, so all fourteen are marked INCOMPLETE on that field.

Index

ID Title Status Priority Notes
PRG-013 Consolidate fire and flower onto water open not assessed RAM headroom thin
PRG-014 WS2025 golden image family is a dependency, not a future item open not assessed second image lane required
PRG-015 magpie contradiction — keep vs. project committed to retire entirely open not assessed decision: R. Chhetry + Rob MacMillan
PRG-016 Seven unmanaged gpus-it hosts unexplained open not assessed
PRG-017 Two undocumented Meraki networks plus three unassigned APs open not assessed
PRG-018 Meraki WDC edge is an MX95 HA pair; the asset registry says MX100 open not assessed registry is wrong
PRG-019 Vendor transfer stack — DECOMMISSION open not assessed decision, not a verification
PRG-020 Stage 3 result — 24 of 25 gpusa VMs moved UNKNOWN → DERIVED open not assessed gated by GOV-011
PRG-021 The estate spans two GCP projects for Puppet purposes open not assessed
PRG-022 gpusa UNKNOWN count is 154, not 146 open not assessed
PRG-023 gpus-dist staleness profile — 95% over three years old open not assessed
PRG-024 duck IP contradiction confirmed from a second source open not assessed corroborates VLN-019
PRG-025 Enumeration surfaces missing from Stage 1 open not assessed
PRG-026 Seven entities declared with no identifying data, never observed live open not assessed

PRG-013 — Consolidate fire and flower onto water

Field Value
id PRG-013
title Consolidate fire and flower onto water
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Authenticated vSphere API, 2026-08-13. water: Xeon Gold 5416S, PowerEdge R660xs, 63.5 GB RAM at 3.8%, 1.66 TB datastore at 0.1%, zero VMs. fire: Xeon E5640 (2010 silicon), R610, 48 GB at 70%, hosts all four core servers. flower: Xeon E5-2697 v3, R630, 128 GB at 11.6%, local datastore 80.6% full. Combined RAM in use across fire and flower is ~48 GB against water's 63.5 GB. All three already mount the same vmstorage NFS datastore, so compute moves without storage moving. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): capacity figures at raw/task3-esxi-capacity-20260813.txt, inventory at raw/task2-esxi-inventory-20260813.txt.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Retires two EOL 6.7 hosts onto supported hardware already owned. Recorded as supplied: RAM headroom is thin — per-VM allocations are needed before committing. ~48 GB into 63.5 GB leaves little margin, and the figure is current usage rather than allocation.

Depends in practice on VLN-023: the shared vmstorage NFS datastore that makes this move cheap is the same array currently running degraded with no spare. Consolidating three hosts onto one storage path does not change that array's state, but it does concentrate what depends on it.


PRG-014 — WS2025 golden image family is a dependency, not a future item

Field Value
id PRG-014
title WS2025 golden image family is a dependency, not a future item — a second image lane is required
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 2 Part B, 2026-08-17. inventory.yaml declares duck, grebe and nfs-gw as REPLACE targeting Windows Server 2025, while the golden-image program currently plans a single Rocky 9 family.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

The framing is the point and is carried as supplied: this is a dependency of the current program, not a later phase. Three declared REPLACE targets have no image lane to be replaced onto.


PRG-015 — magpie contradiction

Field Value
id PRG-015
title magpie declared disposition: keep inside a project committed to empty and retire entirely
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry and Rob MacMillan — the source row names this as a decision for both
evidence Stage 2 Part B, 2026-08-17. inventory.yaml declares disposition: keep, role_status: confirmed-data-team, os: rhel-8, os_eol: 2029-05-31, live RUNNING — in gpus-it-infrastructure, which is committed to empty and retire entirely.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Recorded as supplied: either the commitment has an unwritten exception or magpie needs a destination. Both branches are decisions, not investigations — which is why this row names two owners and no acceptance test rather than a research task. Note PRG-021: magpie is also one of the two hosts that make the estate span two GCP projects for Puppet purposes.


PRG-016 — Seven unmanaged gpus-it hosts unexplained

Field Value
id PRG-016
title Seven unmanaged gpus-it hosts unexplained — running, never Puppet-managed, unaccounted for in either repo
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stages 3 and 3b, 2026-08-17. duck, falcon, grebe, gull, kestrel, nfs-gw, woodpecker — running, never Puppet-managed, and nothing in either repo accounts for them. falcon, gull and kestrel appear only as DNS records; grebe, nfs-gw and woodpecker appear nowhere. falcon's public record (35.223.86.92, 2026-07-27) is the newest content in gpus-dist.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Three of the seven — duck, grebe, nfs-gw — are the same hosts PRG-014 records as REPLACE targets needing a WS2025 lane, and duck carries the IP contradiction in VLN-019 / PRG-024. That falcon's DNS record is the newest content in an otherwise 2017-era repo (PRG-023) is recorded as supplied, without inference about what it means.


PRG-017 — Two undocumented Meraki networks plus three unassigned APs

Field Value
id PRG-017
title Two undocumented Meraki networks — DC Apartments and Major Gifts — plus three unassigned APs
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Meraki API, 2026-08-13. DC Apartments (4 APs — MR33×3, MR52×1) and Major Gifts (systems-manager only), plus three unassigned APs (MR32×2, MR33×1). The working brief says three sites; there are five networks and 32 devices, 29 assigned.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

This closes out the Meraki half of PRG-001 (Stage 1c) as an enumeration result, though PRG-001's own acceptance test remains unstated and that row is not re-statused here.


PRG-018 — Meraki WDC edge is an MX95 HA pair; the registry says MX100

Field Value
id PRG-018
title Meraki WDC edge is an MX95 HA pair with warm spare enabled — the asset registry says MX100 and is wrong
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Meraki API, 2026-08-13. MX95 HA pair, warm spare enabled, primarySerial Q2XN-V4XE-UQKX, firmware wired-26-1-5. The asset registry says MX100. WAN1 virtual IP 38.140.146.68 matches the documented VPN endpoint; the physical MXs hold .66 and .67.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Another instance of the GOV-017 pattern: the entity is present in the registry and its declared attributes are wrong. The virtual-vs-physical IP detail matters for any firewall or peer configuration derived from the registry.


PRG-019 — Vendor transfer stack: DECOMMISSION

Field Value
id PRG-019
title Vendor transfer stack — DECOMMISSION (relay module, transfer crons, storehouse hosts)
register Program work items
status open — a disposition decision is recorded; the decommission itself is not started and certainly not verified
priority / severity not assessed — none supplied
owner R. Chhetry
evidence R. Chhetry decision, 2026-08-17: stale code, not required, removed with the move to Rocky Linux. Covers the relay module, the transfer crons and the storehouse hosts.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
evidenced_by VLN-035stated in the source row. dataflow.us.gl3 resolving to a host that does not exist supports the decommission.
date_raised 2026-08-17
date_verified

This is a decision, not a verification — do not record as done

Carried verbatim from the source. A disposition decision by the owner is an input to work, not evidence that work happened. Under the status rule, done would require the stack removed and verified against live state. Status is open.

Premise corrections recorded with the decision. These correct the picture the decision was taken against and are kept because they change what the decommission actually has to remove:

  • There are 27 cron resources but only 13 potentially-active transfers — 13 are declared twice (ensure => present under if($live == true), ensure => absent in the else branch) and pushPSI is unconditionally absent.
  • Three scripts do not do what their names say. pushFacter performs no transfer at all and only downcases filenames locally. pushFPR exits 0 at line 7 ("Turning this script off for now") and has fired every 30 minutes doing nothing. pushPSI is explicitly deactivated.
  • Scripts live in the Puppet control repo at modules/relay/files/sbin/ with -rooster and -whistler variants; sourceselect => first picks the hostname variant and whistler is the RUNNING one. Endpoint routing comes from pushtab.

pushFPR is the second instance of LOADED ≠ REACHABLE in this row alone — see M-01. Note that this disposition does not close VLN-024: the plaintext credentials in pushtab survive deletion of the code that reads them.


PRG-020 — Stage 3 result: 24 of 25 gpusa VMs moved UNKNOWN → DERIVED

Field Value
id PRG-020
title Stage 3 classification result — 24 of 25 gpusa VMs moved UNKNOWN → DERIVED
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3, 2026-08-17. 24 of 25 gpusa VMs moved UNKNOWN → DERIVED. Only quail remains UNKNOWN (TERMINATED, no node definition). Across all 44 node definitions: 38 DERIVED, 5 ASSERTED-STALE, 1 ASSERTED. kingfisher resolved to DERIVED via the testing modulepath, where postgresql is defined — puppetEnvironment => testing. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/gpusa-it-infrastructure-306400__instances__rchhetry-at-greenpeace.org.jsonconfirmed: 25 instances in gpusa, the denominator of the 24-of-25 claim. Classification output at raw/stage3-noderows.json and raw/stage3-class-resolution.json, with the testing-modulepath resolution that reclassified kingfisher at raw/stage3-class-resolution-testing.json.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by GOV-011derived. DERIVED is not enforced; the source row states this row is subject to that gap.
date_raised 2026-08-17
date_verified

DERIVED is a hypothesis, not evidence

This row looks like progress and is recorded as a classification result, not a verification one. Per GOV-011, thirteen RUNNING VMs with node definitions have never produced a retained catalog — including phoenix itself — so a DERIVED statement on such a host is a hypothesis. kingfisher resolving only via the testing modulepath is worth watching for the same reason.


PRG-021 — The estate spans two GCP projects for Puppet purposes

Field Value
id PRG-021
title The estate spans two GCP projects for Puppet purposes — comparing node definitions against gpusa alone over-reports orphans by two
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3, 2026-08-17. magpie and kingfisher have node definitions and live in gpus-it-infrastructure, not gpusa. Comparing node definitions against gpusa VMs alone over-reports orphans by two.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

This is the fact that corrected the middle orphan figure in GOV-018 (52 → 50). Both hosts are already carried elsewhere: magpie in PRG-015, kingfisher in PRG-020.


PRG-022 — gpusa UNKNOWN count is 154, not 146

Field Value
id PRG-022
title gpusa UNKNOWN count is 154, not 146
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Partial. Supplied: the count is 154, not 146, as read 2026-08-17. The date is given; no source artifact or query was named. To complete: the enumeration output the 154 was read from. ⚠ Still unsourced after searching the enumeration corpus 2026-08-17. No artifact stating the 154 figure was located under ~/estate-enum-2026-08. The row remains partial: the date is recorded, the source is not.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

One of the four counts M-08 records as overturned during these sessions.


PRG-023 — gpus-dist staleness profile

Field Value
id PRG-023
title gpus-dist staleness profile — 900 of 946 files are more than three years old
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3b per-file provenance, 2026-08-17. 900 of 946 files (95%) are more than three years old; 800 were last touched in 2017. The sole current content is DNS zones — zones/greenpeaceusa.org*.zone at 2026-07-27 is the newest file in the repo. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage3b-dist-provenance.txt and raw/stage3b-present-dates.json, with the file index at raw/stage3b-dist-index.json.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Recorded as supplied: DNS is the consistent exception to staleness across both repos. That pattern matters for GOV-014's mirror-before-teardown constraint — the one live thing in an otherwise dead repo is the thing most likely to be missed if the repo is written off as stale.


PRG-024 — duck IP contradiction confirmed from a second independent source

Field Value
id PRG-024
title duck IP contradiction confirmed from a second independent source
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 3b, 2026-08-17. gpus-dist forward and reverse zones both say 10.1.96.40; the live VM is 10.1.96.46; inventory.yaml says .46.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
corroborates VLN-019stated in the source row. Second-source corroboration for duck.cloud.us.gl3 resolving to an address that is not duck's.
date_raised 2026-08-17
date_verified

Two independent sources now agree the zone data is wrong and inventory.yaml is right — a rare direction for this estate, and worth noting because most rows in this load run the other way.


PRG-025 — Enumeration surfaces missing from Stage 1

Field Value
id PRG-025
title Enumeration surfaces to add to any future reconciliation control
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 2, 2026-08-17. Pub/Sub topics, Cloud Scheduler jobs, host-resident services, and power devices were in no Stage 1 surface. Four declared Pub/Sub and Cloud Scheduler entities plus one Pub/Sub subscription (gpus_forms_clamav_worker_sub) could not be checked in either direction.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Feeds GOV-004 alongside GOV-016: that row specifies how a reconciliation control must join, this one specifies what it must look at. Power devices appear here and as GOV-015; the two are the same absence seen from the control side and the estate side.


PRG-026 — Seven entities declared with no identifying data, never observed live

Field Value
id PRG-026
title Seven entities declared with no identifying data and never observed live — unresolved in both directions
register Program work items
status open
priority / severity not assessed — none supplied
owner R. Chhetry
evidence Stage 2 A3, 2026-08-17. gl5_firewall (no serial, no MAC, no IP), synology_controller, visuals_storage_exp_1 through _3, vmware_storage. Unresolved in both directions.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

These are the entities GOV-016's join-key constraint cannot help with: with no serial, MAC or IP there is no key to join on at all. They can be neither confirmed nor refuted by any reconciliation control as currently specified.


Program work register — 2026-08-20 load (PRG-027)

PRG-027 — Enumerate what else on MAPLE runs under Python 3.6.8

Field Value
id PRG-027
title Enumerate every cron job, systemd unit and script on MAPLE that invokes the EOL Python 3.6.8 interpreter
register Program work
status open
priority T1 — derived from a HIGH security finding
owner R. Chhetry
evidence MAPLE, 2026-08-20: python3 --versionPython 3.6.8 (EOL 2021-12-23); python3.11 --versionPython 3.11.13 present but not the default. /usr/bin/python3 -m pip freeze carried reportlab==3.6.8, Pillow==8.4.0, requests==2.27.1, google-cloud-storage==2.0.0. Four production security reports had executed under this interpreter, as root from the root crontab, from ~2026-04 until 2026-08-20 — see VLN-037. Those four are remediated. Nothing else on the host has been examined. The root crontab visibly also carries /usr/local/bin/gpus-cloud-backup.sh (daily 02:00 UTC), whose interpreter is unknown.
acceptance test A written enumeration exists listing, for MAPLE: every root and user crontab entry, every enabled systemd unit, and every script under /opt, /usr/local/bin and /home/* that invokes python3, /usr/bin/python3, or a venv built from either — each classified as (a) confirmed not Python, (b) Python on 3.6.8 and needing migration, or (c) Python already on 3.11+. The enumeration is complete when every entry carries one of those three classifications and no entry is unclassified. Migration of what it finds is separate work and is NOT in this acceptance test — this row closes on knowing, not on fixing.
blocks
blocked_by
evidences VLN-037 — this row is that finding's residual, raised as work rather than left in a residual field.
date_raised 2026-08-20
date_verified

Why this is a row and not a residual line. VLN-037 states the residual accurately: the reports are off 3.6.8, the rest of the host is not, and what else runs under it is unknown. But a residual field has no owner and no acceptance test, so it reads as closed to anyone scanning the register — and this estate's recurring failure is precisely things that read as closed while nobody is looking at them (cf. VLN-010, VLN-011, and VLN-037 itself).

Suggested starting commands, offered as a starting point and not as the enumeration:

sudo crontab -l
for u in $(cut -d: -f1 /etc/passwd); do sudo crontab -l -u "$u" 2>/dev/null | sed "s|^|[$u] |"; done
ls -la /etc/cron.d/ /etc/cron.daily/ /etc/cron.hourly/
sudo systemctl list-units --type=service --state=running
sudo grep -rlE '(/usr/bin/)?python3([^.]|$)' /opt /usr/local/bin /etc/systemd/system 2>/dev/null
for v in $(sudo find /opt /home -name 'bin/python3' -path '*venv*' 2>/dev/null); do echo -n "$v: "; "$v" --version; done

The last line matters most and is the least obvious: a venv inherits the interpreter it was built from. /opt/gpus-reports/venv was a 3.6.8 venv until 2026-08-20 despite python3.11 being installed on the host, and looked identical from the outside to a 3.11 one. Any other venv on MAPLE may be in the same state, and ls will not tell you.


Program work register — 2026-09-01 load (PRG-029)

PRG-029 — Detect report silence: a freshness check over every latest.pdf

Field Value
id PRG-029
title Check the age of gs://gpus-infra-backups-wdc/reports/<type>/latest.pdf against each report's own cadence, for all eight types, and alert when one is older than it should be. Makes a report that did not run detectable, which today it is not by any means at all.
register Program work
status open
priority not assessed — none supplied
owner unassigned
evidence GOV-033, 2026-09-01: the September executive monthly aborted at 08:00:01 on a 30 s read timeout to /api/soc and produced no PDF, no mail and no error line. Every candidate signal was enumerated on the host and every one failed — the wrapper's own log "ERROR: PDF generation failed" is unreachable behind set -euo pipefail, nothing parses /var/log/gpus-reports.log, there is no Prometheus metric or textfile collector on MAPLE, and cron's MAILTO=root fired into an unread 439 KB file (GOV-034). Of the three fixes GOV-033 recommends, this is the only one that would have caught THIS failure rather than making the next one louder: the other two improve what happens when the report runs, and this one notices when it does not. It also covers the two reports with a single named recipient each, where the reader has no baseline for "should have arrived by now".
acceptance test A scheduled check reads the eight latest.pdf objects and alerts when any is older than its cadence allows (daily +1 day for approvals and volume; monthly +1 day for monthly, summary, prebooking; monthly-on-the-15th for newsletter; quarterly for quarterly; weekly for weekly). Demonstrated able to fail by pointing it at a deliberately stale object — or simply by running it against the current bucket, where reports/executive/latest.pdf should already be flagged from the 2026-09-01 miss. A check that runs and never alerts is not evidence it works; the demonstration is part of the test, per PRG-028.
blocks
blocked_by Needs a decision on where it runs and where it alerts — the same open question as PRG-028, and the two should be answered together rather than twice. Deliberately NOT blocked on GOV-034: report silence should become detectable without first fixing cron mail on seven hosts, and routing this through a channel that is already known broken would reproduce the failure it exists to catch.
evidences GOV-033 — this row is that finding's third recommendation, raised as work rather than left in a recommendation block. GOV-034 is the adjacent estate-wide gap and is deliberately a separate item.
date_raised 2026-09-01
date_verified

Why this is a row and not a line inside GOV-033. Same reasoning as PRG-027 and PRG-028, and it is the established pattern here rather than a new one: the GOV- row records what is missing and can be closed honestly once the finding is understood; the PRG- row carries the build, with an owner and an acceptance test. GOV-028PRG-028 was the first instance. Leaving this inside a recommendation block would give it no owner, no acceptance test and no place in a weekly review — and a reader scanning the gap register would see the failure documented and reasonably conclude something was being done about it.

One scoping note worth keeping. It is tempting to widen this to "alert on any failed cron job", which is GOV-034's territory and a much larger piece of work. Resist that: the eight latest.pdf objects already exist, are already written by the production path, and need no agent, no exporter and no mail. It is the cheapest thing in this family that would have worked, and its narrowness is the reason it can ship before the bigger question is answered.


Program work register — 2026-08-27 load (PRG-028)

PRG-028 — Alert on gpus-reports deploy drift, rather than making it merely visible

Field Value
id PRG-028
title Run GOV-028 acceptance command (b) on a schedule and alert on mismatch, so deployed-vs-repo drift in gpus-reports is detected rather than detectable
register Program work
status open
priority not assessed — none supplied
owner unassigned
evidence GOV-028 closed 2026-08-27: gpus-reports now carries a VERSION file and stamps the commit hash into every run's first stdout line, every degraded= line, the mailer's start line and every PDF footer. That makes drift visible. It does not prevent it, and nothing detects it unattended. The condition it fixes ran undetected for two days precisely because nobody was looking, and the guard does not change who is looking — it changes what they would see if they did. R5 has exactly one recipient.
acceptance test A scheduled job runs the currency comparison — git log --oneline "$deployed"..origin/main -- 'gpus-reports/*.py' 'gpus-reports/report_cron.sh' 'gpus-reports/requirements.txt' ':!gpus-reports/test_*.py' — and emits an alert when the output is non-empty. Demonstrated able to fail by pointing it at a stale marker (e.g. 10a48be, which names four undeployed commits) and observing the alert fire. A job that runs and never alerts is not evidence it works; the demonstration is part of the test.
blocks
blocked_by Needs a decision on where it runs and where it alerts. It cannot run on MAPLE alone — MAPLE has no clone of the repo, and a check that only reads the host cannot know what the host is behind.
evidences GOV-028 — this row is that finding's residual, raised as work rather than left in a residual field.
date_raised 2026-08-27
date_verified

Why this is a row and not a residual line. Same reasoning as PRG-027: a residual field has no owner and no acceptance test, so it reads as closed to anyone scanning the register. GOV-028 is closed and its closure is honest — the guard was built, deployed and verified — but a reader who sees only "drift guard: done" will reasonably conclude that drift is now handled. It is not. It is legible, which is a different property, and the gap between the two is exactly where the original two-day drift lived.

The awkward part, stated rather than left implicit. The check needs both sides — the deployed marker (on MAPLE) and the repo history (in a clone) — and no scheduled job today has both. That is the real work in this row; the comparison itself is one command.


Incomplete rows — 2026-08-17 load

ID Evidence Acceptance test
PRG-013 ✅ as supplied, dated missing
PRG-014 ✅ as supplied, dated missing
PRG-015 ✅ as supplied, dated missing
PRG-016 ✅ as supplied, dated missing
PRG-017 ✅ as supplied, dated missing
PRG-018 ✅ as supplied, dated missing
PRG-019 ✅ decision, dated missing
PRG-020 ✅ as supplied, dated missing
PRG-021 ✅ as supplied, dated missing
PRG-022 ⚠ partial — date, no source named missing
PRG-023 ✅ as supplied, dated missing
PRG-024 ✅ as supplied, dated missing
PRG-025 ✅ as supplied, dated missing
PRG-026 ✅ as supplied, dated missing

All fourteen lack an acceptance test; one has partial evidence. Evidence quality in this load is markedly better than the 2026-08-13 load — thirteen of fourteen rows carry a named source and a date, against six of sixteen last time. The acceptance-test gap is unchanged, and for the same reason: these are enumeration findings, and what "closed" means for each is a decision that has not been taken.

Completed

Initiative Completed Notes
Forms portal LIVE in production 2026-07-23 Go-live submission 28bd1ebf; routing worker f9a114b on MAPLE, override OFF; all 6 ingest addresses verified delivering; HappyFox tickets opening.
Broken-access-control remediation 2026-07-23 Phase-1 decrypt/list/submit removed; require_role fail-closed; resolve_user least-privilege; IDOR ownership checks on Phase-2. No evidence of exploitation (empty audit_log + empty users table, independently corroborated). Finding record → T1.8.
Forms 2.5(c) routing worker — Gates 3/4/5 2026-07-23 finalize_submission wired; MAPLE-resident Pub/Sub worker via localhost:25 Postfix; migrations 011 + 012 live. Residual propagation verification → Portal-propagation workstream.
Forms 2.5(d) — subject-template audit-persist + GAP-1 renderer fix + template reconciliation 2026-07-23 Audit-persist gate resolved; <%= Grant Funded => GAP-1 renderer fix landed; template reconciliation done.
Forms Portal Phase 2.5(b) — attachment upload wire-up + cleanup 2026-05-08 β closed with verification gap acknowledged (T1.7). Commits e17dacb (handler) + 28964c0 (cleanup). Migration 002 documents submission_deleted enum addition. See architecture/forms-phase2.5b-cleanup-closeout.md.
Okta Production cutover 2026-04-23 Production tenant live; group-based assignment; Preview kept as dev fallback
forms.greenpeace.us DNS + TLS 2026-04-21 CNAME → ghs.googlehosted.com; managed cert issued
Forms Portal Phase 1 (backend) 2026-04-20 Cloud SQL PG15, CMEK, IAM auth, AES-256-GCM envelope, RLS 4 roles
Meraki P1 inventory 2026-04 Org 395909, 5 networks, 32 devices
Portal backup cron on SKY 2026-03 gpus-portal-backup.sh nightly 02:30 → GCS. ⚠ Cron DIED ~2026-04-17 — see the T3 SKY portal-backup item; restore + backfill needed.
Okta Preview SSO across 4 portals 2026-03 OIDC PKCE, shared gpus-okta-auth.js

Change log

Version Date Author Change
v1.23 2026-09-01 R. Chhetry / Claude PRG-029 raised as the travel workstream closes — a freshness check over the eight reports/<type>/latest.pdf objects in GCS, so a report that did not run becomes detectable. It is GOV-033's third recommendation and, by that finding's own analysis, the only one of the three that would have caught the 2026-09-01 executive-monthly miss rather than making the next one louder. Raised as program work following the GOV-028PRG-028 precedent, and scoped deliberately not to depend on GOV-034 (cron MAILTO=root is inert on all seven hosts) — routing report alerting through a channel already known broken would reproduce the failure it exists to catch. Its where-does-it-run / where-does-it-alert question is the same one PRG-028 carries and the two should be answered together.
v1.21 2026-08-17 R. Chhetry / Claude Evidence-completion pass. Artifact paths appended to PRG-004, PRG-005, PRG-006, PRG-007, PRG-009, PRG-011, PRG-013, PRG-020, PRG-022, PRG-023 and HB-004 from the newly reachable enumeration corpus at ~/estate-enum-2026-08. Figures confirmed exactly: gpusa Cloud SQL 2; gpus-it buckets 18; unattached-disk ratios 7:3 and 15:9; 47 static IPs across three projects; 25 gpusa instances; 15 active DHCP leases. HB-004 gains evidencephoebe-capture/ holds vhost configs and logs for six candidate consumers, though the logs are unanalysed and its acceptance test stays INCOMPLETE. Two figures did not reproduce and were left unchanged: 126 host reservations against ~111 stated (the "workstation" filter is not recorded), and PRG-022's 154 remains unsourced. No acceptance test, status, severity or owner changed.
v1.20 2026-08-17 R. Chhetry / Claude 2026-08-17 enumeration load. New PRG-013PRG-026 section from priorities/backlog-2026-08-17.md v1.0 items B-25–B-38 (Stages 1, 1b, 1c, 2, 3, 3b). No pre-existing row edited; the 13 flagged Done rows from 2026-08-13 remain untouched. Owner R. Chhetry throughout except PRG-015, which names Rob MacMillan as co-decider. No severity or priority stated for any row — all read not assessed, none inferred. All 14 rows lack an acceptance test and are marked INCOMPLETE; PRG-022 is partial on evidence. Nothing is done, including PRG-019, which records a DECOMMISSION decision by R. Chhetry and not a verification. Headline context: 82 declared entries against 481 deduped live entities, gpusa zero against 174, 50.0% UNKNOWN at Stage 2, 24 of 25 gpusa VMs UNKNOWN→DERIVED at Stage 3 (gated by GOV-011 — DERIVED is not enforced). Companion updates: security/vuln/tracker.md v1.9 (VLN-023–VLN-036) and governance/gap-register.md v1.1 (GOV-010–GOV-019 plus a Method learnings section).
v1.19 2026-08-13 R. Chhetry / Claude New ## Program work register — schema v1 section added (PRG-001–PRG-012 program items, HB-001–HB-004 host-level blockers), mapped from provisional P-1–P-16. No pre-existing row in this document was edited, reworded, re-statused or removed; the new section carries its own status vocabulary and the rule that done means verified end-to-end via the real production path against live state. No row in the new section is done. 13 of 16 new rows lack an acceptance test and 6 lack evidence entirely — recorded as INCOMPLETE, not filled with guesses. Two corrections established first-hand: gpus-forms-clamav-worker is listed in inventory.yaml so P-3 is 3 unlisted services not 4; and inventory.yaml:149 carries the same water/ESXi-6.7 misattribution as VLN-004, which VLN-013's acceptance test does not cover. Host rows for gannet/emu/ostrich/catbird/phoebe do not exist in inventory.yaml, so the four HB- blockers name their host and are parked until PRG-002 creates them. Companion registers opened the same day: security/vuln/tracker.md v1.8 (VLN-013–VLN-022) and governance/gap-register.md v1.0 (GOV-001–GOV-009). Pre-existing Done rows flagged separately in priorities/flag-list-unverified-done-2026-08-13.md — list only, re-adjudication tracked as GOV-008.
v1.17 2026-07-24 R. Chhetry / Claude New T1 compliance item: credential-rotation control never closed out. The quarterly forms-portal-credential-rotation-quarterly control emits records (Q3 2026 evidence committed 61eefff) but is never back-filled — 3 ACTION-NEEDED carried since Q2 (KMS rotation timestamps, Cloud SQL backup verification, HappyFox rotation) + an MFA-enforcement flag; possibly two consecutive quarters with no live console verification. Same class as the contract-vs-code and DB↔repo drift findings — artifacts asserting a state nobody verified. Pre-announcement gate: confirm Okta MFA enforcement before the staff notice — forms authz now rests entirely on Okta identity and the announcement leads with "sign in with Okta".
v1.16 2026-07-23 R. Chhetry / Claude FORMS PORTAL LIVE IN PRODUCTION (go-live 2026-07-23, submission 28bd1ebf). Routing worker f9a114b on MAPLE, override OFF, all 6 ingest addresses verified delivering, HappyFox tickets opening. Phases 2.5(a)–(d) DONE; 2.5(c) Gates 3/4/5, GAP-1 renderer fix, and template reconciliation closed. Broken-access-control remediation COMPLETE — Phase-1 decrypt/list/submit removed, require_role fail-closed, resolve_user least-privilege, IDOR ownership checks on Phase-2; no evidence of exploitation (empty audit_log + empty users table, independently corroborated). Restructured T1 from "finish forms" to cutover + post-go-live hardening. Legacy forms.us.gl3 retires 2026-08-01 (a Saturday — flagged for reconsideration); parallel running until then; LDAP deprovisioning gap argues for cutting over sooner. New/organized T1 workstreams: Cutover (staff notice, decommission, in-flight handling, redirect-or-dark, duplicate-ticket overlap, LDAP gap); Detection pipeline (audit_log→SOC + Cloud Run→Wazuh sinks don't exist; rules 100026–100029 inert by starvation; routing-worker empty SyslogIdentifier) — blocks the drill program; Security hardening (retire auth.py authz; narrow backend-SA over-grants; Cloud Armor = LB+NEG arch change; /metrics + /health/deep exposure; CORS *; per-instance rate limiting; MAC validation); Content/data (<% = Note => in Finance Termination rendering literally; live-DB legacy-tag sweep; DB↔repo drift root cause — HIGH; Insurance checkbox unanswerable-as-No; HR Termination TODO; footer sweep; DSAR/erasure + retention-lock gap); Okta dedicated forms app (requested from Conan 2026-07-23); 2.5(e) HappyFox API (queue names from API, ticket-audit migration, per-queue double-ticket deconfliction, API preferred — DMARC spoofing case); Portal propagation (live/derived, verify pages render). New T1.8 documentation set (threat model, ASVS L2, PCI-DSS out-of-scope determination, OWASP→NIST→MITRE mapping, IR runbook + DRP RTO/RPO, BAC + governance finding records, 2.5(d) as-built, field-exposure regen) — depends on the assessment. T1.9 exercise program (SQLi tabletop 60m + blue-team 90m + red-team 90m) — BLOCKED on detection. Phase 3 HappyFox marked superseded by 2.5(e); T5-EXPANDED un-gated (Gate 5 landed). New adjacent T3 item: SKY portal-backup cron died ~2026-04-17 (portals/ snapshots stopped; server backups current).
v1.15 2026-06-10 R. Chhetry / Claude Forms 2.5(c) implementation underway. Design doc → v0.3 (commits f0b79a8, 9a5281b): §3a wire contract with purged=410 DECIDED, §3b async finalize, §6a actions-schema-as-built, §17 coverage triad, 2.5(e) relabel. Migration 011 APPLIED LIVE (audit_action enum 26→30, commit 4bdfdf9); migration 012 (submission_finalized, 30→31) approved 2026-06-10, applying. Gate 3 in progress per G3.0 decisions (C1 add GET, C2 routing_result finalize marker, C3 full predicate at finalize, C4 α publishes unconditionally, C5 = 012). Gate 4 = routing worker; Gate 5 = propagation (NEW, blocks "2.5(c) done"). 7-day Pub/Sub retention clock starts at Gate 3 push — Gate 4 within the window. New T5-EXPANDED filed: portal static-debt retirement (P0–P4) per 2026-06-10 read-only audit of both portals; P0 truth-fixes may interleave before Gate 4. Standing rule recorded: all portal content LIVE or DERIVED, never hardcoded; coverage standard to require render-from-source. Operational learnings logged: L2TP stale-tunnel 3rd occurrence, FortiClient utun6 confound, IAP 4003 host-level (break-glass promotion), legacy in_formfeed MySQL on public IP (decom precondition filed).
v1.14 2026-06-04 R. Chhetry / Claude Forms 2.5(c) routing pipeline design COMMITTED (v0.2, commit 40f90d5; live at architecture/forms-phase2.5c-design/). 2.5(c) moved from next-up to design-committed / implementation-pending. Locked: transport B-iii (MAPLE-resident Pub/Sub pull worker, localhost:25 Postfix, zero new secret, carries its own systemd + monitoring + IR/DR/drill); reuse forms_app role (RLS USING(TRUE) → GATE-4 state-coverage N/A); signed-URL attachment delivery (inline MIME deferred); interim minimal-plaintext body NOT gated on 2.5(d); lean audit — migration 011_routing_audit_actions.sql adds 4 enum values (submission_routed, submission_route_failed, email_sent, email_failed). Next code: 011 migration + finalize_submission wire-up (pure stub today); one code-time confirm — _status_to_wire at routes_phase2.py:57.
v1.13 2026-05-21 R. Chhetry / Claude α ClamAV close-out housekeeping. Sigrefresh build trigger now filters by gpus-forms-clamav-worker/** (A0) — fixes the double-build-on-every-push issue. Backfill drain complete (A1) — 3 fixture attachments (PNG/DOCX/XLSX, all rchhetry β-phase test files from 2026-05-08) re-fired via synthetic Pub/Sub publish, all clean, GCS bytes unchanged, 3 new audit_scanned_clean rows; attachments table now clean=5, pending=0. Design doc alpha-clamav-worker.md → v1.3 (§6 sweep reframed required, §10 DLQ subscription clarified, new §13a 7-point Cloud SQL access spec for new DB-using services). Commit 3 hardening row added to summary. Four new T3 candidates filed: VPN cold-start packet loss, IAP-to-MAPLE 4003 backend-fail, Cloud Scheduler missed-tick on cold-start, and "GCP terraform tree has no VCS" (the no-VCS finding supersedes the previously-filed "WDC VPN/route Terraform state drift" — drift framing was wrong; no remote ever existed to drift from). A2 deferred per the no-VCS surprise; tonight's mitigation is a local snapshot of the 11 .tf files.
v1.12 2026-05-20 R. Chhetry / Claude α ClamAV Commit 2 COMPLETE. Slack alerts + GCS quarantine tag + /dlq-alert + /sweep-stuck + Cloud Scheduler wiring shipped. EICAR test verified full infected pipeline end-to-end to #us-soc-alerts. Worker rev 00010-gr7. Migrations 009 + 010 (audit_action enum + scanner UPDATE/INSERT WITH CHECK widen). Terraform 6 resources applied (DLQ sub + scheduler SA + 2 scheduler jobs). 8 hardening/hygiene items filed for Commit 3.
v1.11 2026-05-19 R. Chhetry / Claude α ClamAV Commit 1 COMPLETE. Migrations 004-008 shipped (enum + IAM user + grants + RLS + sequence USAGE). Pipeline verified end-to-end on T1.7 fixture (07c3c9e0) and real upload (4a0f5bd5). Worker gpus-forms-clamav-worker live (rev 00002-dwx, 2Gi). α.1 follow-ups: stuck-scanning sweep, DLQ subscription + alert, backfill drain of 3 remaining pending rows.
v1.10 2026-05-14 R. Chhetry T1.7 forms portal frontend silent-attachment-drop bug closed.
v1.9 2026-05-08 R. Chhetry β phase: Phase 2.5(b) attachment upload + 2.5(b.cleanup) MIME/size truth consolidation COMPLETE. Three-layer verification PASS on 999cf0cc-…; 4-source-of-truth divergence collapsed to Config.ATTACHMENT_*; production env vars removed; schema migration 002 documents submission_deleted audit_action enum addition. Two orphan-intent submissions cleaned up (d20d2ac8, e93efbc3). New T1.7 filed: frontend silent-attachment-drop bug surfaced during β verification — SPA submit must be gated on attachment validation state. Verification gap on rev 00043-d9q acknowledged in closeout doc.
v1.8 2026-05-08 R. Chhetry New T1.6 workstream filed: forms portal SOC/observability integration. Gap surfaced during Phase 2.5(b) scoping — forms portal has been operationally invisible to SOC since cutover. 8 sub-items spanning logging/metrics/Wazuh/SOC tab/runbook/DRP/drill/ASVS. Sequenced after Phase 2.5 functional work, before T4 SOC Tickets.
v1.7 2026-05-07 R. Chhetry γ phase: design doc correction shipped (2de57c9) — forms-phase2.5a-design.md now matches shipped code. α phase: pulldown regression for yes/no booleans CLOSED (f31ec5f) — Flask route converter fix. T1.5 sub-section now fully closed (all 4 items a/b/c/d). New cross-cutting lesson on hypothesis falsification via empirical data added.
v1.6 2026-04-30 R. Chhetry Phase 2.5(a) COMPLETE — 3 commits (07cb75c, 10bcd0d, 83f85cb), submissions now persist end-to-end (verified API + DB + audit). γ phase shipped same day: font darkness (f5c79e0+e58dabc), NO_DATA_TYPES filter (6d9231a). T1.5 items b/c/d CLOSED; item a (pulldown regression) remains. Cross-cutting lessons added: design doc accuracy, MAPLE access pattern.
v1.5 2026-04-29 R. Chhetry Phase 2.5(a) Commit 1 SHIPPED (07cb75c, auth_v2 User.username via preferred_username). Commits 2+3 paused at design-complete state. New T1.5 sub-section "Forms Phase 2 UX gaps" added covering pulldown regression (yes/no booleans), COST CENTER duplicate, NOTESDIVIDER label leak, and font contrast.
v1.4 2026-04-28 R. Chhetry FieldRenderer pulldown bug FIXED (commit 5512d8e). Phase 2.5 "Phase 2 backend wire-up" promoted as T1 sub-section after discovery that Phase 2 submission stubs persist nothing. Cross-cutting lesson on Cloud SQL access added.
v1.3 2026-04-27 R. Chhetry Forms Phase 2 frontend cutover complete. Phase 2.1 cleanup queue (5 items) added under T1. FieldRenderer pulldown bug added. Cross-cutting Cloud Run env-var lesson documented.
v1.1 2026-04-24 R. Chhetry Re-sequenced: forms (T1) → Meraki (T2) → WDC foundation (T3). WDC items no longer compete with forms momentum.
v1.0 2026-04-24 R. Chhetry Initial draft — consolidated priorities from memory + recent session notes