GPUS-IT Priorities¶
Classification: CONFIDENTIAL — Internal Use Only Document:
priorities/gpus-it-priorities.md· v1.23 · 2026-09-01 · GPUS-IT Owner: Rajesh Chhetry · Review cadence: weekly (Fridays) + on every new initiative
Purpose¶
Living tracker of active and queued initiatives for the GPUS-IT infrastructure program. This document supersedes scattered priority notes across memory and Cowork; the canonical order of work lives here.
Rules of engagement:
- Every initiative lists its gating dependencies — don't start work that is gated on unfinished prerequisites.
- Every new asset added by an initiative must ship with its matching IR runbook (
rb-00N-*.md), DR procedure update indrp.md, and red/blue drill entry intabletop-playbooks.md+blue-team-drills.md. Infra docs without runbooks are not considered "done". - Status values:
Planned·In Progress·Blocked·Done·Deferred - When an item completes, move it to Completed with the completion date and delete it from the active queue.
Sequencing¶
Forms portal is LIVE in production as of 2026-07-23 (go-live submission 28bd1ebf; routing worker f9a114b on MAPLE; override OFF; all 6 ingest addresses verified delivering; HappyFox tickets opening). Phases 2.5(a)–(d) are DONE and the broken-access-control remediation is complete. The forms program has shifted from build to cutover + hardening.
The order is now: cutover forms, harden forms, then Meraki, then WDC foundation.
- Forms cutover (T1, the only hard external deadline). Legacy forms portal (
forms.us.gl3) retires 2026-08-01 — parallel running until then. The staff notice (email + Slack) and legacy decommission must land before that date. NOTE: Aug 1 is a Saturday — flag the cutover date for reconsideration (an LDAP deprovisioning gap argues for cutting over sooner, not later). - Forms hardening (T1). Detection pipeline (blocks the drill program), remaining security hardening, content/data fixes, dedicated Okta forms app, 2.5(e) HappyFox API path, and portal propagation. Documentation set and the security exercise program (SQLi tabletop + blue/red drills) are gated on the detection pipelines existing.
- Meraki (T2) — Sedita fix + Meraki SSO, then fold Meraki into status/SOC/MkDocs coverage.
- WDC foundation (T3) — ESXi + Synology NAS inventory, cleanup, monitoring, and documentation. Deferred intentionally: the hypervisor + NAS are stable enough today that finishing the forms cutover + hardening is higher-value.
Everything else (SOC Ticketing, Vendor Access, MySQL decom, status site automation) slots in after WDC foundation unless a specific gating dependency inverts the order.
Summary¶
| Tier | Initiative | Status | Target |
|---|---|---|---|
| T1 | Forms portal LIVE in production — go-live 2026-07-23 (sub 28bd1ebf); routing worker f9a114b on MAPLE, override OFF, all 6 ingest addresses verified delivering, HappyFox tickets opening |
Done | 2026-07-23 |
| T1 | Broken-access-control remediation — Phase-1 decrypt/list/submit endpoints removed; require_role fail-closed; resolve_user least-privilege; IDOR ownership checks on Phase-2 endpoints. NO evidence of exploitation (empty audit_log + empty users table, independently corroborated) |
Done | 2026-07-23 |
| T1 | Forms 2.5(a)–(d) backend + routing worker — submit/attachment/finalize wired, subject-template audit-persist gate resolved, GAP-1 renderer fix + template reconciliation landed | Done | 2026-07-23 |
| T1 | Forms cutover → forms.greenpeace.us — staff notice (email+Slack), legacy forms.us.gl3 decommission, in-flight handling, redirect-or-dark. Only hard external deadline. LDAP deprovisioning gap argues for cutting over sooner. STAFF-NOTICE CONTENT REQUIREMENT added 2026-08-25: the notice MUST state that portal mail arrives from alerts@greenpeace.us and is Gmail-tagged External, and that this is the permanent steady state — not a temporary condition. Per the 2026-08-25 decision that no send-as alias will be registered (see gpus-reports/README.md and architecture/forms-phase2.5c-design.md §7). A reader told this is temporary reports it as phishing three months later, and a phishing report against a real portal address costs more to unwind than the sentence costs to write. The four staff guides already carry the wording; the notice must match them, not paraphrase. |
In Progress | 2026-08-01 (Sat — reconsider) |
| T1 | Detection pipeline — audit_log→SOC + Cloud Run→Wazuh log sinks (neither exists); Wazuh rules 100026–100029 inert by starvation, not config. Blocks the exercise program |
Planned | 2026-08 |
| T1 | Security hardening — post-go-live — retire legacy auth.py authz primitive; narrow backend-SA project-level over-grants; Cloud Armor/WAF (arch change + cost case); /metrics + /health/deep exposure; CORS *; per-instance rate limiting; MAC field validation |
Planned | 2026-08 |
| T1 | Content / data fixes — ~~<% = Note => in Finance Termination template rendering literally in prod~~ DONE 2026-08-10 (template b1a51710e91b8b18 → {{ Note }}; it is the Facilities template, which Finance receives because it is bundled onto the Facilities destination line — audit §5.7); ~~live-DB legacy-tag sweep~~ DONE 2026-08-10: 1 of 48 live templates carried a legacy tag, and it was that one — estate now clean; DB↔repo drift root cause (HIGH); Insurance checkbox unanswerable-as-No; HR Termination authoring TODO; footer sweep; DSAR/erasure path; attachments retention lock |
Planned | 2026-08 |
| T1 | Credential-rotation control never closed out (compliance drift) — the quarterly forms-portal-credential-rotation-quarterly control emits records but is never back-filled: 3 ACTION-NEEDED carried since Q2 (KMS rotation timestamps, Cloud SQL backup verification, HappyFox credential rotation) + an MFA-enforcement flag; possibly two consecutive quarters with no live console verification. Same class as the contract-vs-code and DB↔repo drift findings — artifacts asserting a state nobody verified. PRE-ANNOUNCEMENT GATE: confirm Okta MFA enforcement before the staff notice goes out — forms authz now rests entirely on Okta identity and the announcement leads with "sign in with Okta". |
Planned | MFA: before cutover notice |
| T1 | Okta — dedicated forms app + group — requested from Conan 2026-07-23; then verify SPA redirect URI, aud validation, preview-tenant test, client-ID cutover |
In Progress | 2026-08 |
| T1 | 2.5(e) HappyFox API — queue names from API, ticket-audit migration, per-queue double-ticket deconfliction; API preferred over email-ingest (greenpeace.us DMARC p=none; help.greenpeace.org no DMARC → forged ticket plausible); ingest sender restrictions unverified |
Planned | 2026-08 |
| T1 | Portal propagation — inventory.yaml reflects live forms state; coverage + validate_portal_presence pass; status + SOC show forms LIVE (derived, not hardcoded); verify live pages render it |
Planned | 2026-08 |
| T1 | Okta cleanup — remove localhost redirect URIs (separate from the dedicated forms app above) | Planned | 2026-08 |
| T1 | Forms portal Phase 1.5 — legacy data migration | Planned | 2026-08 |
| T1 | Forms portal Phase 1.6 — ON CONFLICT refactor | Planned | 2026-08 |
| T1.8 | Forms documentation set (mkdocs) — threat model; OWASP ASVS L2; PCI-DSS applicability (out of scope, no CHD); OWASP→NIST 800-53→MITRE mapping; forms IR runbook + DRP entry; BAC security-finding record; governance finding; 2.5(d) as-built; field-exposure matrix regen; mitre-attack/threat-vectors/pentest-schedule/calendar updates. Depends on the assessment | Planned | 2026-08 |
| T1.9 | Forms security exercises — SQLi tabletop (60m), blue-team detection drill (90m), red-team simulation (90m). GATED on detection pipelines | Blocked | 2026-09 |
| T1 | α ClamAV scan worker — Commit 3 (hardening + hygiene, 8 items) | Filed | 2026-08 |
| T2 | Meraki cleanup — Sedita site + Meraki SSO | Planned | 2026-06 |
| T2 | Meraki integration — status/SOC/MkDocs coverage | Planned | 2026-06 |
| T3 | WDC foundation — ESXi inventory & cleanup | Planned | 2026-07 |
| T3 | WDC foundation — Synology NAS inventory & cleanup | Planned | 2026-07 |
| T3 | GCP terraform tree has no VCS — dedicated session: secrets review of tfvars/tfstate, .gitignore design, git init, push to private remote (CSR), import state (supersedes WDC VPN/route drift framing) |
Filed | 2026-06 |
| T3 | VPN cold-start packet loss — diagnosis + fix (Mac → GCP private subnet path drops on cold start; 2nd incident in 7 days) | Filed | 2026-06 |
| T3 | IAP-as-break-glass — IAP-to-MAPLE 4003 confirmed host-level 2026-06-10 (not IAP edge); fix VM-side so IAP works as break-glass path when VPN/SSH down | Filed | 2026-06 |
| T3 | Cloud Scheduler missed-tick on cold-start (sweep.clean missed 05:45 UTC 2026-05-21; investigate scheduler retry config vs worker min-instances tradeoff) | Filed | 2026-06 |
| T3 | SKY portal-backup cron died ~2026-04-17 — portals/ snapshots stop there (server backups current). Restore the cron; backfill the gap |
Filed | 2026-08 |
| T3 | 2.5(e) PRECONDITION — HappyFox ingest/API deconfliction. G4.4 proved email_template legs landing in HappyFox-watched inboxes (e.g. gpus-it-support@) auto-open tickets via email-to-ticket (#USITS00357259, 2026-06-12; worker dispatched no happyfox action — verified). Once 2.5(e) dispatches via API, forms with both action types against the same queue will DOUBLE-TICKET. Deconflict per-form email recipients vs API destinations before 2.5(e). The deconfliction points in OPPOSITE directions per form and there is no single rule: on it-support-request remove the EMAIL leg (the HF leg is the intended path); on employee-termination-notification remove the FIVE HF legs (the 3 email legs already cover all 5 owning teams — audit §5.10/§5.11). Miss the second and every termination opens 5 duplicate tickets across 5 queues. contract-extension-notification (3 HF + 1 email) is UNREVIEWED — do not assume either direction. ~~Context: org moved to API because HappyFox blocks greenpeace.us email~~ CORRECTED 2026-08-10: that claim is STALE or was never true. #USITS00365024 and #USITS00365073 both opened in queue 45 from mail the relay rewrote to alerts@greenpeace.us. HappyFox accepted it. The claim helped justify the move to API dispatch and was carried for months untested. The DMARC argument for API (row above) is unaffected and still stands. |
Filed | feeds 2.5(e) |
| T4 | SOC Ticketing tab | Planned | 2026-08 |
| T4 | Vendor Access Portal (replace SFTP) | Planned | 2026-08 |
| T5 | Status site automation — Phase A (live Cloud Run list) | Planned | 2026-09 |
| T5 | MySQL decommission | Planned | 2026-09 |
| T5 | Status site automation — Phase B (BigQuery billing export) | Planned | 2026-10 |
| T5 | Status site automation — Phase C (per-service cost attribution) | Planned | 2026-11 |
| T5 | T5-EXPANDED — Portal static-debt retirement (P0–P4, per 2026-06-10 audit; forms Gate 5 now landed → unqueued; P0 truth-fixes may interleave with forms hardening) | Filed | 2026-Q3 |
Recently completed forms/α items (see Completed table): forms 2.5(a)–(d), routing worker (Gate 4/5), BAC remediation, T1.7 SPA submit-gate, α ClamAV Commits 1 + 2.
T1 — Forms: cutover & post-go-live hardening¶
Forms portal is LIVE in production as of 2026-07-23. Go-live submission
28bd1ebf; routing workerf9a114bon MAPLE with override OFF; all 6 ingest addresses verified delivering; HappyFox tickets opening on real submissions. Phases 2.5(a)–(d) DONE (see the phase records further down this section, kept as history). The active T1 work is now cutover (the only hard external deadline) plus the hardening workstreams below. The phase-history sub-sections (Phase 2, 2.5(a)–(e), UX gaps, etc.) follow as completed records.
Cutover to forms.greenpeace.us — the only hard external deadline¶
- Status: In Progress. Legacy forms portal (
forms.us.gl3) retires 2026-08-01; new and legacy run in parallel until then. - ⚠ Date flag: 2026-08-01 is a Saturday. Flag for reconsideration — a weekend decommission means no staffed coverage if a rollback or in-flight-submission issue surfaces. Weigh against the LDAP argument below, which pushes the other way (sooner).
- Items:
- Staff notice (email + Slack). Operational, independent of every other workstream — has no technical dependency and must go out before Aug 1. Do not let it wait on hardening.
- Legacy decommission of
forms.us.gl3— plan the teardown; decide redirect-or-dark for the old hostname. - In-flight submission handling — drain/route anything submitted to the legacy portal during the overlap window.
- Duplicate-ticket ambiguity during overlap — both portals feed the same ingest addresses, so a form submitted on both (or mid-migration) can double-open tickets. Deconflict before/during the overlap (ties into 2.5(e) per-queue deconfliction below).
- LDAP deprovisioning gap — argues for cutting over SOONER. The Okta→LDAP sync keeps breaking, so a disabled account may still authenticate to the legacy portal. Every day of parallel running is a day a deprovisioned user retains legacy access. This is the security case for an earlier cutover, in tension with the Saturday-date concern above.
Detection pipeline — blocks the drill program¶
- Status: Planned. Gating dependency for the T1.9 exercise program — there is no point running blue-team detection drills against detections that cannot fire.
- Items:
audit_log→ SOC pipeline — does not exist. Forms audit events never reach SOC. (Overlaps the long-standing T1.6 observability workstream; this is the concrete blocker for drills.)- Cloud Run → Wazuh log sink — does not exist. Wazuh rules 100026–100029 are inert by starvation, not by misconfiguration — the rules exist but no forms/Cloud Run events are being fed to them.
- Verify Wazuh rule bodies +
ossec.confon MAPLE — confirm the rule definitions and agent config are actually what we think before wiring the feed. - Routing-worker
systemdunit has an emptySyslogIdentifier— fix so the worker's logs are attributable in the journal/sink (prerequisite for the Cloud Run→Wazuh feed to be useful).
Security hardening — post-go-live (remaining)¶
The broken-access-control remediation is DONE (see the dedicated record below). These are the remaining hardening items surfaced during/after go-live:
- Retire legacy
auth.pyauthz primitive (Step 4). Migrate/api/admin/reloadtoauth_v2; theuserstable becomes login-audit only. (Completes the auth-module consolidation begun in Phase 2.1 item 5.) - Backend service-account over-grants. The backend SA holds project-level
roles/cloudkms.cryptoKeyEncrypterDecrypter+ project-wideroles/storage.objectAdmin. Narrow to the specific key and bucket scope (least privilege). - Cloud Armor / WAF — NOTE this is an architecture change, not a policy toggle.
forms.greenpeace.usis a direct Cloud Run domain mapping today; adding Cloud Armor requires standing up a new external HTTP(S) load balancer + serverless NEG in front. Treat as an architecture change + cost case, not a quick enablement. /metricsis unauthenticated — leaks per-form submission volumes. Gate it./health/deepexposes DB + KMS status publicly — reduce to an authenticated/internal probe.- CORS
origins="*"— tighten to the known SPA origin(s). - Rate limiting is
memory://per-instance — the effective limit multiplies by instance count (so the real ceiling is N× the configured value). Move to a shared backend or account for instance count. - MAC field accepts malformed input — add validation.
Broken-access-control remediation — DONE (2026-07-23)¶
- Status: COMPLETE. Recorded here as a closed security finding; the full finding record (finding + remediation + no-exploitation conclusion) is a T1.8 documentation deliverable.
- What was fixed:
- Phase-1
decrypt/list/submitendpoints removed. require_rolemade fail-closed.resolve_usermoved to least-privilege.- IDOR ownership checks added on the Phase-2 endpoints.
- Phase-1
- No evidence of exploitation — corroborated independently by two facts: the
audit_logis empty of any decrypt/view events, and theuserstable is empty. Two independent signals both consistent with "never exploited."
Content / data fixes¶
<% = Note =>legacy placeholder surviving in the Finance Termination template — renders literally into production tickets. Highest-visibility content bug post-go-live.- Live-DB sweep for other legacy tags — a repo sweep provably misses drift (see DB↔repo root-cause item), so the sweep must run against the live database, not the repo.
- DB↔repo drift root cause — HIGH IMPORTANCE. If the repo is not the authoritative source for templates/forms, then the gap-2 sweep, the field-exposure matrix, and the CI leak-gate are all validating the wrong artifact. Root-cause which store is authoritative before trusting any repo-based check.
- Insurance checkbox: required with only "Yes" — unanswerable as "No" (a required field the user cannot legitimately decline). Fix, and sweep for the same pattern across other forms.
- HR Termination template contains an unfinished authoring TODO line — remove/complete before it renders to a recipient.
forms.us.gl3→forms.greenpeace.usfooter sweep — replace stale legacy-hostname footers.- No DSAR / subject-erasure path — there is no implemented data-subject-access-request or erasure flow; the retention purge executor is unlocated; the attachments bucket retention (2555d) is still UNLOCKED. Privacy/retention gap to close.
Okta — dedicated forms app + group¶
- Status: In Progress — dedicated forms app + group requested from Conan 2026-07-23.
- Then: verify the SPA redirect URI,
aud(audience) validation, run a preview-tenant test, and perform the client-ID cutover to the new app. - Related but separate: the existing "remove localhost redirect URIs" cleanup (Phase 2.1 / Okta cleanup section) is a different task from standing up the dedicated app.
2.5(e) — HappyFox API integration¶
- Status: Planned. Native HappyFox API ticket creation, distinct from today's live email-ingest path.
- Items:
- Queue names from the API (stop hardcoding queue identifiers).
- Ticket-audit migration — persist HappyFox ticket IDs / outcomes into the audit trail.
- PER-QUEUE double-ticket deconfliction — a single form routing to N queues risks doubling a ticket across all N once both the email-ingest and API paths are active. This is the same hazard as the cutover overlap duplicate-ticket item; solve once, per queue. (See the T3 "2.5(e) PRECONDITION" row for the G4.4 evidence.)
- API vs email-ingest decision — security prefers API. Rationale:
greenpeace.usis DMARCp=none(vs.org'sp=reject), andhelp.greenpeace.orghas no DMARC record at all, so a forged portal-looking email ticket is plausible on the ingest path. The authenticated API path removes that spoofing surface. - HappyFox ingest sender restrictions unverified — confirm whether the ingest inboxes actually restrict senders (if not, the spoofing risk above is live today).
Portal propagation¶
- Status: Planned (residual of forms Gate 5). The standing rule applies: portal content must be LIVE or DERIVED, never hardcoded, and must render from sources, not merely have endpoints exist.
- Items:
inventory.yamlreflects the live forms state.- Coverage +
validate_portal_presencepass. status+socportals show forms LIVE — derived, not hardcoded.- Verify the live pages actually render it (the audit standard: endpoints existing is not enough; the frontend must display the derived value).
Phase 2 React + Okta (forms frontend SPA)¶
- Status: Phase 2 frontend SPA: COMPLETE 2026-04-27. Live on
forms.greenpeace.us. Backend revgpus-forms-backend-00033-bfx, frontend revgpus-forms-frontend-00003-74f. 28 forms render with editorial typography. Auth pipeline (Okta OIDC PKCE → JWKS →/api/forms) fully verified end-to-end. - Goal: Ship
gpus-forms-frontendas React SPA with Okta OIDC PKCE auth, callinggpus-forms-backendwith a validated JWT bearer. - Gating: None —
gpus-forms-backendis live on Cloud Run,forms.greenpeace.usresolves, Okta Production cutover complete (2026-04-23). - Deliverables:
forms-frontend/ingpus-infra-portalsrepo- JWT validation middleware on
gpus-forms-backend(and on status / security / soc backends — Phase 2 is org-wide, not forms-only) - Cloud Run deploy + Cloud Build trigger wired
- Asset docs required: Update
iar.mdwith forms-frontend service; no new IR runbook needed (covered by existing portal runbooks); blue-team drill entry for "forged/expired JWT rejected" inblue-team-drills.md. - Phase 2 follow-up: FieldRenderer pulldown lookup — COMPLETE 2026-04-28 (commit
5512d8e). Phase 1's/api/forms/<id>serializer now returnsid/label/pulldown_id(matchingcontract.ts). New endpointGET /api/pulldowns/<name>added withPulldownResponseshape. Backend revgpus-forms-backend-00034-8tn.
Phase 2.5 — Phase 2 backend wire-up¶
- Discovered 2026-04-28: the Phase 2 SPA has been live but Phase 2's three submission endpoints (
POST /api/submissions,POST .../attachments,POST .../submit) are unwired stubs inroutes_phase2.py. They return fake UUIDs and a hardcoded'TBD@greenpeace.us'literal. Submitting a form via the SPA appears to work to the user but persists nothing — no DB write, no GCS upload, no email/HappyFox call. The well-formed routing model exists (actionstable populated by the MySQL migrator with per-formaction_type+destination+template_id) but no submit-time code reads it. There is also no mailer module in the backend at all. - Implementation scope:
- Wire
create_submission(routes_phase2.py:102) to write to thesubmissionstable with KMS envelope encryption for non-searchable fields, returning real UUIDs. Pattern reference:routes/submissions.py:37(Phase 1's working submit handler).- STATUS: COMPLETE 2026-04-30
- Commit 1 SHIPPED 2026-04-29:
07cb75c(auth_v2User.usernameviapreferred_usernameclaim, revgpus-forms-backend-00036-89b). - Commit 2 SHIPPED 2026-04-30:
10bcd0d(design doc tomkdocs-portal/docs/architecture/forms-phase2.5a-design.md). - Commit 3 SHIPPED 2026-04-30:
83f85cb(routes_phase2.pycreate_submissionwire-up, revgpus-forms-backend-00038-5kg). - Verified end-to-end 2026-04-30: API contract response shape, DB persistence (encrypted fields + audit log), MAPLE-side inspection of test submission
0bac9326-4968-477b-8d10-d7d6f457e2a8. - Phase 2 SPA submissions now ACTUALLY persist (was theater since Phase 2 cutover).
- Commit 1 SHIPPED 2026-04-29:
- STATUS: COMPLETE 2026-04-30
- Wire
upload_attachments(routes_phase2.py:128) to GCS bucketgpus-forms-attachmentswith ClamAV scan and attachment row inserts.- STATUS: COMPLETE 2026-05-08 (β closeout — see
architecture/forms-phase2.5b-cleanup-closeout.md)- 2.5(b) handler shipped 2026-05-08:
e17dacb(revgpus-forms-backend-00041-mcf), three-layer verification PASS on test submission999cf0cc-…(GCS object + attachments row + audit_log row, all consistent to the microsecond). - 2.5(b.cleanup) shipped 2026-05-08:
28964c0(revgpus-forms-backend-00043-d9q), 4-source MIME/size truth (prod env vars,Config, module-level shadows,_schema.yamlhint) collapsed toConfig.ATTACHMENT_MAX_BYTES+Config.ATTACHMENT_ALLOWED_MIMEas single source. Production env varsMAX_UPLOAD_BYTES+ALLOWED_MIME_TYPESremoved. - Schema migration
forms-backend/schema/002_add_submission_deleted_action.sqldocumentssubmission_deletedaudit_action enum value (added in production during β step 8 to enable orphan-intent cleanup). - ClamAV scan handled in Phase 2.5(b.2) — see dedicated section below, COMPLETED 2026-05-19. At β time all attachments landed with
clamav_status='pending'; partial indexidx_attachments_clamavwas already in place for the scanner. - Verification gap accepted: end-to-end positive case on rev
00043-d9qblocked by T1.7 SPA bug (silent attachment drop). Pre-cleanup positive cases (5c57b2e6 docx, 62d2fc73 xlsx) plus mechanical-substitution diff correctness accepted as evidence of backend behavior. Rationale in closeout doc.
- 2.5(b) handler shipped 2026-05-08:
- STATUS: COMPLETE 2026-05-08 (β closeout — see
- Wire
finalize_submission(routes_phase2.py:151) to:- STATUS: COMPLETE — SHIPPED TO PRODUCTION 2026-07-23 (go-live). The routing worker is live: rev
f9a114bon MAPLE, override OFF, all 6 ingest addresses verified delivering, HappyFox tickets opening on real submissions (go-live submission28bd1ebf). Gates 3, 4, and 5 all landed. Locked decisions (from v0.2) as shipped:- Transport/shape: B-iii — MAPLE-resident routing worker, Pub/Sub pull subscriber, sends via
localhost:25Postfix reusing thereport_mailer.pypattern. Zero new secret. First non-serverless service in the rebuild → carries its own systemd unit + monitoring + IR/DR/drill obligations (see the routing-workerSyslogIdentifiernote under Detection). - DB role: reuse
forms_app(no new role; RLSUSING(TRUE)makes the GATE-4 state-coverage work N/A). - Attachments: signed-URL delivery in body (inline MIME deferred).
- Interim body: minimal plaintext, NOT gated behind 2.5(d).
- Audit: lean — migration
011_routing_audit_actions.sql(4 enum values) +012(submission_finalized), both live.
- Transport/shape: B-iii — MAPLE-resident routing worker, Pub/Sub pull subscriber, sends via
- Gate close-out (2026-06-10 → 2026-07-23): Gate 3 (C1 GET, C2
routing_resultfinalize marker, C3 full predicate at finalize, C4 α publishes unconditionally, C5 = migration 012), Gate 4 (routing worker), and Gate 5 (propagation: inventory + coverage-standard portal-propagation section + status Services section + SOC alerting + ops triad + α backfill) all DONE. Residual Gate-5 propagation verification is carried forward as the Portal propagation workstream below (confirm live pages actually render forms LIVE, derived not hardcoded). SELECT actions FROM actions WHERE form_id = <...> ORDER BY action_order- For each action, branch by
action_type:happyfox_template→ render viatemplatestable, POST to HappyFox API (secrets already in Secret Manager:gpus-forms-happyfox-api-key,gpus-forms-happyfox-auth-code)email_template/email_raw→ render template, SMTP send via Postfix on MAPLE (currently no SMTP client in backend)
- Aggregate
routing_resultintosubmissions.routing_resultJSONB.
- STATUS: COMPLETE — SHIPPED TO PRODUCTION 2026-07-23 (go-live). The routing worker is live: rev
- Build template render layer (Jinja2 against
templates(id, body)). - Build HappyFox API client (configures the integration that SOC dashboard memory #20 noted as "not configured").
- Decide policy on
email_template/email_rawactions: keep, drop, or redirect through HappyFox? (Per Rajesh: HappyFox API is the operational current path because Google's stricter email policies broke the email-to-HappyFox flow.)
- Wire
- Estimated effort: ~~multi-session work~~ DONE. Sequence ran (a)→(b)→(c)→(d) with the routing/HappyFox path shipped at go-live 2026-07-23.
- Gating: completed the "forms portal actually works" milestone. The 2.5(e) HappyFox API client (native ticket creation, distinct from today's email-ingest path) is now its own workstream below. The SQLi tabletop drill is unblocked functionally but is now GATED on the detection pipeline existing (no point drilling detection that can't fire) — see T1.9.
Phase 2.5(b.2) — α ClamAV scan worker¶
- Status: Filed 2026-05-19. COMPLETED 2026-05-19 (5 migrations 004-008 shipped — enum + SQL user + table grants + tight scanner RLS + USING widen + sequence USAGE; pipeline verified end-to-end: 07c3c9e0 fixture + 4a0f5bd5 real upload both clean with audit rows; worker rev
gpus-forms-clamav-worker-00002-dwxon Cloud Run, 2Gi, min-instances=0). - What shipped: new Cloud Run service
gpus-forms-clamav-worker(Pub/Sub-push ongpus-forms-attachmentsOBJECT_FINALIZE→ claim →clamscan→ verdict toattachments+audit_log). Scoped Terraform (SA + IAM + topic + DLQ + OIDC push sub; VPN drift untouched, see T3 row). CSR Cloud Build triggers (push + weekly sigrefresh). Design docalpha-clamav-worker.mdv1.2. - α.1 / Commit-2 follow-ups: shipped 2026-05-20 in Commit 2 — see
### Phase 2.5(b.2) — Commit 2immediately below.
Phase 2.5(b.2) — Commit 2 (Slack + GCS tag + DLQ + sweep)¶
- Status: COMPLETED 2026-05-20 (worker rev
gpus-forms-clamav-worker-00010-gr7on:c9c2016; migrations 009 + 010 shipped; Terraform 6 resources applied inclamav-worker.tf; EICAR test verified the full infected pipeline end-to-end at 16:45 UTC). - What shipped:
- Slack-post for infected verdicts — placeholder-tolerant fetch from Secret Manager (
gpus-forms-clamav-slack-webhook), lazy module-global cache, cold-start dance documented; verifiedclass=wiredon first call during EICAR. - GCS quarantine metadata tag (
quarantined=true,quarantine_reason=<sig>,quarantined_at=<iso8601>) — best-effort with try/except per C4 ordering (tag first, Slack second; audit row is the system of record). /dlq-alertroute + DLQ push subscription ongpus-forms-attachment-uploaded-dlq— closes the design §10 gap surfaced at Commit 1 GATE 4. Always returns 200 (no DLQ-of-DLQ)./sweep-stuckroute + Cloud Scheduler tick every 15 min — closes α.1 stuck-scanningrecovery. AtomicUPDATE … RETURNING, one audit row per flipped id withrevision_at_sweepfromK_REVISIONfor deploy correlation, Slack only on flips > 0.- Cloud Scheduler API enabled on
gpus-infra;clamav-scheduler@gpus-infra.iamSA wired withrun.invokeron the worker (resource-scoped) +cloudbuild.builds.editoron the project. - Weekly sigrefresh scheduler job (Sun 02:00 UTC) fires the existing
gpus-forms-clamav-worker-sigrefreshbuild trigger — closes the README's "weekly auto-refresh — NOT yet wired" follow-up from Commit 1.
- Slack-post for infected verdicts — placeholder-tolerant fetch from Secret Manager (
- Verification: EICAR test at 16:45 UTC. Full pipeline green:
event.received→claim.ok→clamscan→INFECTED signature=Eicar-Test-Signature→quarantine.tag.ok→slack.webhook.resolved class=wired(first fetch — lazy init confirmed) → Slack post landed in#us-soc-alerts→ audit rowattachment_scanned_infectedwritten.- All 5 verification checks (worker log, GCS metadata, attachments row, audit_log row, Slack channel visual) passed.
- Bonus signal: the
*/15sweep tick fired naturally at 16:45:03 UTC during the test window, loggingsweep.clean count=0— gate 2d wiring proven live end-to-end without manual stimulus. - Test fixture (1 submission + 1 attachment + 1 audit row + 1 GCS object) cleaned up post-validation; attachments now back at pre-test state (
clean=2, pending=3).
- α.1 / Commit 3 follow-ups (hardening / hygiene — NOT Commit-2 blockers):
- Tighten
cloudbuild.builds.editorto per-trigger IAM on justgpus-forms-clamav-worker-sigrefresh(currently project-wide — over-broad blast radius). - Rename
SLACK_PLACEHOLDER_PREFIX_OK→SLACK_VALID_PREFIX(constant name misleads; logic is correct). - Normalize
_sweep_stuckaudit INSERT to use the_AUDIT_ACTIONdict (currently hardcodes the enum value as a string literal; PG casts implicitly so it works, but inconsistent with_record). - Sigrefresh build trigger needs
--included-files=gpus-forms-clamav-worker/**filter — currently fires on every push tomainregardless of path, racing with the main trigger and burning a build slot. forms-backend/schema/002 ordinal collision (002_rls.sql+002_add_submission_deleted_action.sql) — rename the second to a higher ordinal.- Backfill drain: 3 remaining
pendingattachments (real user uploads pre-dating Commit-1 worker deployment) still need scanning — carried over from Commit 1; safe to drain via the established byte-correct re-fire pattern, OR let/sweep-stuckpick them up if theiruploaded_atis in the stuck window. - Commit the
~/terraform/gpus-infra/terraform/clamav-worker.tfCommit-2 additions to that repo (applied to GCP but the .tf change is local-only).
- Tighten
- T3 candidates filed during this arc (separate from Commit 3):
- VPN cold-start packet loss — 2nd incident in 7 days; routing to private-IP Cloud SQL (10.34.0.0/24, Private Services Access) failed mid-session.
- IAP-to-MAPLE 4003 backend-fail — only matters as a backup access path when direct SSH is also down, but should still resolve.
Phase 2.5(a) design doc correction¶
-
Status: CLOSED 2026-05-07 via
2de57c9. Four corrections applied tomkdocs-portal/docs/architecture/forms-phase2.5a-design.mdto match shipped code (commit83f85cb):_audit_v2helper signature (kwargs + body)- allow-list deny audit action (
"auth_failure"not"submission_denied") SubmissionFieldattachment (session.flush+submission_id, not relationship)- second
_audit_v2call site (consistency fix)
Plus a process-note appendix appended to the doc capturing the read-first-discipline lesson. - Discovered: 2026-04-30 during Phase 2.5(a) Commit 3 implementation. - Context: The design doc at
mkdocs-portal/docs/architecture/forms-phase2.5a-design.mdwas committed in10bcd0dand contained 3 bugs in the helper specs that Code caught before Commit 3 shipped: 1._audit_v2helper signature — design saidactor=actorkwarg; correct isactor_username+ 4 other Phase 1 fields (target_id,target_type,details,request_id,success). 2. Allow-list deny audit action — design said"submission_denied"; that value isn't in theaudit_actionENUM. Correct is"auth_failure"withdetails={"reason": "not_in_allow_list"}andsuccess=False— matches Phase 1's existing pattern. 3.SubmissionFieldattachment pattern — design usedrow.submission = submission(assumes ORM relationship that doesn't exist). Correct is Phase 1'ssession.flush()+submission_id=submission.idpattern.
Forms Phase 2 UX gaps¶
- Discovered: 2026-04-29 PM during Phase 2.5(a) browser regression check.
- Status: T1.5 sub-section is now FULLY CLOSED — items a/b/c/d all closed. Section can stay in the doc as a closed-finding record or be moved to a "Recently completed" archive at section author's discretion. Item (a) CLOSED 2026-05-07 via
f31ec5f. Items (b), (c), (d) CLOSED 2026-04-30 via 2026-04-30 γ phase commits. - Priority order within sub-section: a (pulldown — user-visible blocker for form submission accuracy) → d (font — visible polish) → b, c (cosmetic).
-
Items:
a. Pulldown regression — yes/no booleans — CLOSED 2026-05-07 via
f31ec5f.Root cause: Flask `<string>` route converter (the default) does not match forward slash. Pulldown names containing `/` (e.g. "Grant Funded ? (Yes/No)", "Bargain Type (In/Out)", "Shipping Label/Box Required?") failed to route — Cloud Run URL processing decoded `%2F` back to `/` before route matching, splitting the path into multiple segments that no route consumed. Result: 404 → SPA renders "Could not load options". Fix: change the `@bp.get` decorator from `"/pulldowns/<name>"` to `"/pulldowns/<path:name>"` — single character change. Diagnosis chain used DB inspection (via MAPLE access pattern established 2026-04-30): hypothesis 1 (missing pulldown rows) was falsified by data showing all referenced `pulldown_name`s exist in the `pulldowns` table. Pattern noticed in failing data: all 7 failing names contained `/`.b. COST CENTER duplicate field — CLOSED 2026-04-30 via
6d9231a(resolved as a side effect of theNO_DATA_TYPESfilter — duplicate was a divider/instructions row with that label).c. NOTESDIVIDER label leak — CLOSED 2026-04-30 via
6d9231a(forms-backendroutes/forms.pyfiltersNO_DATA_TYPESfrom form-detail serializer).d. Font darkness / contrast — CLOSED 2026-04-30 via
f5c79e0+e58dabc(forms-frontend.field-labelcolor#4a4a44— WCAG AA at ~9.6:1 contrast over--surface-warm).
Phase 2.1 — Phase 1 Production cutover completion¶
- Goal: Finish the Okta Preview → Production cutover that only partially landed in Phase 1's auth/config layer. Phase 1 was originally built against the Okta Preview tenant; the Production cutover on 2026-04-23 left several stale references and shim layers in place. The 5 items below were surfaced 2026-04-27 during the forms-frontend smoke test and consolidate the remaining cleanup. None are user-visible bugs today; they are maintenance traps.
- Gating: None — Phase 2 SPA is live; cleanup can land at any time.
- Items (priority order):
forms-backend/auth.py— replaceaudience=config.OKTA_AUDIENCEwithaudience=config.OKTA_CLIENT_ID. Remove theOKTA_AUDIENCECloud Run env var workaround. Estimated 1 commit, ~5 min.forms-backend/config.py— changeOKTA_ISSUER = f"https://{os.environ.get('OKTA_DOMAIN', 'greenpeaceeu.oktapreview.com')}"to readOKTA_ISSUERdirectly with defaulthttps://greenpeaceeu.okta.com, matchingauth_v2.py's pattern. RemoveOKTA_DOMAINenv var afterward. Estimated 1 commit, ~10 min.forms-backendsecurity headers — remove CSP/HSTS/X-Frame-Options/etc. from Flask middleware. nginx (forms-frontend/nginx.conf) is now the sole source for response-time security headers. The duplicate headers don't break anything but create maintenance traps. Estimated 1 commit, ~15 min.forms-backendCSP — any CSP that must remain in Flask should remove stalegreenpeaceeu.oktapreview.comreferences. Replace withgreenpeaceeu.okta.com. Stale since the 2026-04-23 Preview→Prod cutover. Estimated 1 commit, ~10 min.- Consolidate
forms-backend/auth.pyandauth_v2.pyinto a single module once both stable. The_key_for_kidretry pattern shipped 2026-04-27 inauth.py(commitd0a7867) carries forward. The v2 module's signin/role pattern carries forward. Defer until after a few weeks of stable production traffic. Estimated 1 PR-sized change.
T1.9 — Forms security exercise program (SQLi tabletop + blue/red drills)¶
- Status: BLOCKED — gated on the Detection pipeline (
audit_log→SOC and Cloud Run→Wazuh log sinks must exist first; Wazuh rules 100026–100029 are inert by starvation today). No point running a blue-team detection drill against detections that cannot fire. - Goal: Exercise detection and response for SQL injection against the live forms backend.
- Components: SQLi tabletop (60m) → blue-team detection drill (90m) against the live endpoint → red-team simulation (90m) adversarial test.
- Deliverables: Updates to
tabletop-playbooks.mdandblue-team-drills.md; Wazuh (+ any WAF) rule tuning if gaps surface; executive summary. - Owner: Rajesh.
Phase 3 HappyFox integration (forms backend) — SUPERSEDED by 2.5(e)¶
- Status: Superseded. The auto-ticket routing shipped at go-live via the routing worker's
happyfox_template/ email-ingest legs. Remaining native-API work is tracked as 2.5(e) — HappyFox API integration above (queue names from API, ticket-audit migration, per-queue double-ticket deconfliction, API-vs-ingest decision). Kept here for lineage. - Historical goal: Wire form submissions to auto-create HappyFox tickets per form's routing config; retry + DLQ on failed ticket creation; ticket ID written back to submission record.
- Asset docs required (carried into T1.8):
iar.md(HappyFox as integrated system); forms IR runbook "HappyFox API outage" + "ticket creation failure → DLQ drain"; blue-team drill entry for "spoofed HappyFox webhook".
Forms portal Phase 1.5 — legacy data migration¶
- Goal: Migrate existing legacy-form submissions (where applicable) into the new schema.
- Gating: Phase 3 HappyFox integration live (so migrated records route correctly on any post-migration edits).
- Scope: One-time batch import; validate row counts, checksum fields, encrypted columns round-trip correctly; retain legacy source read-only for 90 days post-migration.
Forms portal Phase 1.6 — ON CONFLICT refactor¶
- Goal: Refactor upsert logic in
gpus-forms-backendto use properON CONFLICTclauses rather than check-then-insert race conditions. - Gating: Phase 1.5 complete (don't refactor insert paths mid-migration).
- Scope: Submission insert, audit_log append, pulldown cache refresh.
Okta cleanup — remove localhost redirect URIs¶
- Goal: Remove dev-convenience localhost redirect URIs from the Okta Production app (
0oavvg1y33wTWFsmP417) once Cloud Run deploy of forms-frontend verifies. - Gating: Phase 2 Cloud Run deploy verified — met (forms live in production). Unblocked; can land at any time. Coordinate with the dedicated-forms-app cutover above so the URI cleanup targets the right app.
- Done-when: Production app redirect URI list contains only
https://*.greenpeace.us/*entries. Preview tenant (greenpeaceeu.oktapreview.com, client0oadhpjktd5UfCMDm0x7) retained as dev fallback.
T1.6 — Forms portal SOC / observability integration¶
STATUS: NOT STARTED — gap surfaced 2026-05-08 during Phase 2.5(b) scoping discussion
Priority: P2 — should ship after Phase 2.5 functional work completes (2.5b/c/d/e), before SOC Tickets workstream (T4) starts. T4 will assume forms portal events flow into SOC; this workstream makes that true.
Context:
The forms portal (forms.greenpeace.us, gpus-forms-backend, gpus-forms-frontend, Cloud SQL gpus-forms-db) has been operationally invisible to SOC since Phase 1 cutover. While the rest of GPUS infrastructure (WDC servers SKY/RAIN/SUN/WIND, GCP VMs OAK/MAPLE/CEDAR) emits structured events into Wazuh + ELK + Prometheus and is visible across the 15 tabs of soc.greenpeace.us, forms portal emits zero events into that pipeline.
Existing forms-portal capabilities:
- KMS envelope encryption for sensitive form fields (Phase 2.5a)
- Audit log table in Cloud SQL with actor_username, actor_ip, target_id, request_id, success
- IAM-protected Cloud SQL access via service accounts
- Allow-list authorization on form submission
Observability gaps:
- No SOC dashboard tab — the 15 tabs at soc.greenpeace.us don't include "Forms"
- No Wazuh rule coverage for forms-portal events
- No Cloud Logging → CEDAR/MAPLE pipeline for forms events (audit log lives only in Postgres, not exported)
- No Prometheus metrics emitted from forms-backend
- No alert routing for forms-portal anomalies (failed auth bursts, validation errors, malicious upload attempts)
- No IR runbook (rb-00N-forms-*.md) for forms-portal incidents
- No DR procedure for forms-portal in drp.md
- No red/blue drill in tabletop-playbooks.md
Compliance framing:
- PCI DSS: not applicable (no cardholder data flows through forms portal)
- NIST 800-53 / 800-171: applicable as best-practice — control families AC (Access Control), AU (Audit & Accountability), SC (System & Communications Protection), SI (System & Information Integrity). Audit log + KMS encryption partially satisfy AU/SC; AC + SI need observability work.
- MITRE ATT&CK: not a compliance framework, but rest of GPUS infrastructure maps detected events to MITRE techniques. Forms portal events should reach the same taxonomy.
- OWASP ASVS: directly relevant to the SPA + API surface; needs explicit verification pass.
- CIS: applies to host hardening (not Cloud Run services directly); covered for forms portal's underlying platform.
- GPUS internal IRP/DRP framework: forms portal needs runbook + DRP + drill coverage matching what the rest of infrastructure has.
Sub-items (provisional scope — refine when work starts):
a. Audit log → Cloud Logging structured events. Forms-backend audit_log table inserts should also emit structured Cloud Logging entries with proper severity. Pipeline: forms-backend → Cloud Logging → log sink → CEDAR (Elastic) for indexing.
b. Wazuh rule additions for forms-portal events. New rule IDs in the 100020+ range (memory entry on Wazuh ruleset). Cover auth_failure (HIGH severity), attachment_rejected (MEDIUM — potential abuse), submission_created (LOW — informational), attachment_uploaded (LOW). Emit MITRE technique tags where applicable.
c. Prometheus metrics from forms-backend. Standard four (request count, latency, error rate, in-flight) plus domain-specific (submissions_per_minute, auth_failure_rate, upload_size_p99, etc.). Scrape via MAPLE.
d. New Forms tab on soc.greenpeace.us. Tab structure: submission volume, auth-failure rate, attachment activity, slowest queries, recent rejections, threat hunting view (filter audit_log by anomaly patterns).
e. IR runbook rb-006-forms-portal-incident.md. Cover scenarios: compromised submitter account, mass-upload abuse, infected attachment in GCS (forward-look at ClamAV scenario), Cloud SQL unavailability, KMS key rotation, DEK compromise.
f. DR procedure for forms-portal in drp.md. Cover Cloud SQL point-in-time recovery, GCS bucket recovery, Cloud Run rollback pattern, Okta tenant outage fallback.
g. Red/blue drill in tabletop-playbooks.md + blue-team-drills.md. Scenario: malicious attachment uploaded by compromised submitter. Validate detection pipeline end-to-end.
h. OWASP ASVS verification pass. Walk the standard against the forms portal surface; document gaps.
Estimated scope: 3-5 sessions if done as a focused workstream. Could be done incrementally if items a-d are prioritized first (functional observability) and e-h follow (process + verification).
Dependencies:
- Phase 2.5(b)/(c)/(d)/(e) ideally complete first (gives stable surface to instrument)
- Wazuh rule slot range coordination (memory entry ranges)
- Cooperation with T4 SOC Tickets workstream (forms events should auto-ticket)
T1.7 — Forms portal frontend: client-side validation must block submit, not just attachment¶
STATUS: FILED 2026-05-08. COMPLETED 2026-05-14 (commit f175810 shipped 2026-05-12; browser-verified end-to-end 2026-05-14: oversized blocks submit, wrong-MIME blocks submit, clear re-enables, positive case sub 9ac940c8).
Severity: Medium (silent data loss, recipient-visible).
Scope: Frontend only (forms-frontend SPA).
Bug: When client-side validation rejects an attachment (oversized, wrong MIME, or other triggers), the SPA hides the attachment but does not disable the submit button. Submission proceeds to the backend without the attachment. Recipient team sees a complete-looking submission with no attachment and assumes it was intentional.
Known instances (both cleaned up):
d20d2ac8-c19d-48df-a0a4-4f833a750e9b— 2026-05-08 17:15 UTC, csv test, cleaned in β step 8 (audit_log id=10).e93efbc3-4c6e-4b7a-adc9-169d4aedec70— 2026-05-08 18:09 UTC, docx test post-env-var-removal, cleaned in β closeout (audit_log id=12).
Probable triggers (uncharacterized):
- Server-side rejection signal (oversized, wrong MIME)
- Stale
/api/configcache showing the old MIME allowlist (frontend may have cached the pre-cleanup 4-MIME list withtext/csvand withoutdocx/xlsx) - Possibly other (see closeout doc
architecture/forms-phase2.5b-cleanup-closeout.md)
Fix shape (TBD by frontend session):
- Submit button should be disabled while any attachment field has a validation error
- OR submit handler should check for unresolved attachment errors before POSTing to
/submit - OR both
Relationship to other T-priorities:
- Filed below T1.6 (forms portal SOC/observability integration)
- Independent of α (ClamAV) and 2.5(c)/(d)/(e) tracks
- Should be addressed before any meaningful production user testing — silent attachment drop produces forensically-confusing audit trails (submission_created with no attachment_uploaded) and recipient-team confusion
Estimated: 1 session.
T1.8 — Forms documentation set (mkdocs)¶
STATUS: Planned — depends on the security assessment. The assessment (threat model + ASVS) is the upstream artifact several of these records cite; sequence the assessment first, then the derived docs. All new/updated docs live in the mkdocs portal.
Assessment (do first):
- Threat model for the forms portal surface (SPA + API + routing worker + Cloud SQL + GCS + HappyFox path).
- OWASP ASVS L2 assessment — walk the standard against the live surface; document gaps.
- PCI-DSS applicability assessment — expected outcome out of scope (no cardholder data flows through forms); record the determination so it's on file, not assumed.
- Control mapping: OWASP → NIST 800-53 → MITRE ATT&CK, matching the taxonomy the rest of the estate already uses.
Records / runbooks (derive from the assessment):
- Forms IR runbook (
rb-006-forms-portal-incident.mdper T1.6) + DRP entry indrp.mdwith Cloud Run + Cloud SQL RTO/RPO. - Security finding record — the broken-access-control finding, its remediation, and the no-exploitation conclusion (empty
audit_log+ emptyuserstable, independently corroborated). - Governance finding — the API contract documented the Phase-1
decrypt/list/submitendpoints as removed while the code kept serving them; nothing checks that the contract's claims match the code. Record the finding and propose a contract-vs-code conformance check. - 2.5(d) as-built — document the subject-template audit-persist behavior as shipped.
- Regenerate the field-exposure matrix (against the authoritative store — see the DB↔repo drift item; regenerating against the repo is invalid if the repo isn't authoritative).
- Update
mitre-attack.md,threat-vectors.md,pentest-schedule.md,calendar.md.
T2 — Meraki¶
Meraki cleanup — Sedita site + Meraki SSO¶
- Goal: Close out the Meraki P2 follow-ups identified after the P1 inventory (org 395909, 5 networks, 32 devices — completed 2026-04).
- Gating: T1 forms tier complete.
- Scope:
- Sedita site misconfiguration — resolve (specifics to be confirmed at start-of-work)
- Meraki SSO — currently broken, wire to Okta Production
- Deliverables: Fixes verified end-to-end (Okta login → Meraki dashboard for an admin test user); Sedita site returns to expected operational state.
Meraki integration — status/SOC/MkDocs coverage¶
- Goal: Fold Meraki into the same documentation and monitoring posture as the rest of the estate — Meraki currently lacks matching coverage.
- Gating: Meraki cleanup above complete.
- Deliverables:
- Status site: Meraki org card (device count, online/offline, firmware currency)
- SOC site: Meraki alerts surfaced (security events, config changes, WAN uplink loss)
- MkDocs: new
architecture/meraki-network.mddescribing org, networks, devices, admin model, SSO posture iar.mdentries for Meraki org + each site- Syslog from Meraki → WIND (so Wazuh indexes Meraki events into CEDAR)
- IR runbook:
rb-006-meraki-compromise.md(admin account takeover, rogue config push, AP impersonation) - DR procedure: Meraki config backup/restore procedure added to
drp.md(Meraki backs up config in-cloud, but document how to roll back + how to replace a bricked device) - Red/blue drill: tabletop "Meraki admin credentials leaked" in
tabletop-playbooks.md; blue-team detection drill for "unexpected config change outside change window" inblue-team-drills.md
T3 — WDC foundation¶
SKY portal-backup cron died (~2026-04-17) — ADJACENT (non-forms)¶
- Status: Filed. Not a forms item, but an open backup-coverage gap worth surfacing alongside the WDC work.
- What: The
gpus-portal-backup.shcron on SKY (nightly 02:30 → GCS, shipped 2026-03; see Completed) stopped running ~2026-04-17.portals/snapshots stop at that date. Server backups are current — only the portal-content snapshot stream is affected. - Do: Root-cause why the cron stopped, restore it, and backfill the snapshot gap (≈Apr 17 → now). Note this compounds the T5-EXPANDED P0 finding that
soc-site+forms-frontendwere already unchecked by portal-backup coverage — restoring the cron should also extend coverage to those.
ESXi inventory & cleanup¶
- Goal: Bring the ESXi hypervisor — currently undocumented — under the same documentation, monitoring, and IR posture as everything else.
- Gating: T1 + T2 complete.
- Deliverables:
- Inventory: ESXi version, licensing, VMs hosted, networking, hardware health, management plane exposure
- Confirm or sever ESXi ↔ NAS coupling (decision: is the NAS a datastore, a backup target, or both?)
- Cleanup: disable unused accounts, rotate admin credentials, enable syslog → WIND
- Reconfigure: NTP, DNS, timezone, email alerts →
gpus-it-security@greenpeace.org - Monitoring: Prometheus scrape via
vmware_exporter, Grafana dashboard, Wazuh agent on guest VMs where feasible - Status site: ESXi card (version, uptime, VM count, datastore usage)
- SOC site: ESXi in asset coverage map
- MkDocs:
architecture/wdc-hypervisor.md wdc-hostregistry.csventry;iar.mdentry- IR runbook:
rb-007-esxi-compromise.md(hypervisor takeover, guest escape, management plane breach) - DR procedure: ESXi host failure recovery in
drp.md - Red/blue drill: tabletop "ESXi vCenter creds leaked" in
tabletop-playbooks.md; blue-team detection drill "unexpected VM clone / snapshot export" inblue-team-drills.md
- Known risk: ESXi 6.7 is already flagged as EOL in
tracker.md(VLN-004). Inventory may surface the need to accelerate hypervisor replacement — if so, that becomes its own T-tier item.
Synology NAS inventory & cleanup¶
- Goal: Same as ESXi above, for the Synology NAS.
- Gating: ESXi inventory complete (likely coupled — NAS may be serving as an ESXi datastore, which affects cleanup sequencing).
- Deliverables:
- Inventory: model, firmware, volumes, shares, users, backup targets, relationship to ESXi
- Cleanup: disable unused accounts, rotate admin credentials, enable SNMP + syslog → WIND
- Reconfigure: NTP, DNS, timezone, email alerts
- Monitoring: Prometheus scrape via SNMP, Grafana dashboard
- Status site: Synology card (volume health, SMART, firmware)
- SOC site: in asset coverage map
- MkDocs:
architecture/wdc-nas.md wdc-hostregistry.csventry;iar.mdentry (classification, owner, retention, criticality)- IR runbook:
rb-008-nas-compromise.md(ransomware on shares, credential theft, firmware tampering) - DR procedure: NAS failure + volume rebuild in
drp.md - Red/blue drill: tabletop "Synology admin portal exposed" in
tabletop-playbooks.md; blue-team detection drill "mass file encryption on shares" inblue-team-drills.md
GCP terraform tree has no VCS (supersedes "WDC VPN/route TF state drift")¶
- Goal: Initialize version control on
~/terraform/gpus-infra/terraform/(currently untracked on rchhetry's Mac, single point of failure) so terraform changes can be reviewed, rolled back, and "what's the source of truth" has an answer. - Discovered: 2026-05-21, during α ClamAV Commit 2 close-out (A2).
git -C ~/terraform/gpus-infra/terraform statusreturnedfatal: not a git repository, and exhaustive.gitsearch across/Users/rchhetryconfirmed no repo contains these .tf files. The previously-filed "WDC VPN/route Terraform state drift" item assumed a remote repo existed to drift from — the framing was wrong; the no-VCS problem is the parent. - Acute risk mitigation (in effect 2026-05-21):
~/Downloads/terraform-snapshots/2026-05-21/holds md5-verified copies of all 11 .tf files (no tfvars/tfstate/tfplan copied — those need secrets review first). Local-disk redundancy only; not version control. - Inherited drift (still real, blocked until VCS exists):
terraform planagainst live state showsgoogle_compute_vpn_tunnel.wdc_tunnellocal_traffic_selectorforcing replacement;google_compute_route.onprem_mgmt+google_compute_route.onprem_prodmust be replaced as dependents;google_compute_instance.{cedar,maple,openvas}in-place updates. Replacing tunnel + routes would tear down the WDC↔GCP site-to-site VPN (SKY/RAIN DNS-DHCP, SUN Prometheus, WIND ELK). Interim mitigation: all clamav-worker Terraform applied-target-scoped so the drift is never actioned. - Deliverables for the dedicated session:
- Secrets review of
terraform.tfvars+terraform.tfstate*+tfplan: enumerate values, identify what must NOT enter VCS. - Design
.gitignore: minimumterraform.tfvars,*.tfstate*,tfplan,.terraform/, plus anything from secrets review. git initin~/terraform/gpus-infra/terraform/; first commit of the 11 .tf files (and any safe-to-commit ancillary files).- Decide remote: Cloud Source Repositories (matches existing convention for
gpus-infra-portals) vs private GitHub. - Push initial commit; document the remote in the worker README.
- With VCS in place: drift triage — root-cause the VPN tunnel
local_traffic_selectormismatch, decide reconcile direction (update Terraform to match live, OR plan a maintenance-window apply that recreates tunnel/routes), if recreation: scheduled change window with WDC-connectivity-loss comms. - Remove the
-targetworkaround note from the clamav-worker README once unscoped apply is safe.
- Secrets review of
- Gating: independent of WDC inventory work; should run before any further unscoped
terraform apply. Best-suited to a dedicated session — touching secrets + remote setup + drift reconcile + maintenance window planning needs full attention.
T4 — Queued¶
SOC Ticketing tab¶
- Goal: Replace ad-hoc alert triage with tracked tickets on
soc.greenpeace.us. - Gating: T3 complete (stable asset inventory before we wire ticketing to it).
- Sources: Wazuh (level ≥ 10), Prometheus alertmanager, AIDE change alerts, Fail2ban bans, OpenVAS critical/high.
- Dedup: 5 min window on (source, rule_id, host).
- SLA: Critical = 15 min ack / 4 hr resolve · High = 1 hr ack / 24 hr resolve · breach → Slack
#soc-alerts+ email. - Asset docs required:
iar.mdupdate; IR runbookrb-009-soc-ticketing-outage.md; blue-team drill for "silent alert drop" (ingestion pipeline broken but tickets still showing green).
Vendor Access Portal (replace SFTP)¶
- Goal: Zero-trust replacement for the current SFTP vendor drop.
- Gating: SOC Ticketing in place (so vendor-portal anomalies ticket correctly from day one).
- Controls: Vendor IP whitelist via Cloud Armor, signed expiring URLs (max 72h), full audit trail, per-vendor bucket prefixes, ClamAV scan before internal consumption.
- Asset docs required:
iar.md; IR runbookrb-010-vendor-portal-abuse.md(stolen signed URL, vendor account compromise); blue-team drill for "vendor credential used from unexpected geography".
T5 — Backlog¶
Status site automation — Phase A¶
- Goal: Eliminate hardcoded values in
status-site/index.html. - Scope: Cloud Run service list via
gcloud run services listat render time; server count viaservers.py; cost block labelled "last updated YYYY-MM-DD" (still manual this phase, but honest about staleness). - Estimated effort: 2h.
MySQL decommission¶
- Goal: Retire the legacy MySQL instance. Remaining dependencies to be confirmed during inventory.
- Gating: Legacy migration path confirmed (see Forms Phase 1.5 outcome — likely overlap).
- Asset docs required:
iar.mdremoval;drp.mdupdate to drop MySQL recovery procedure; final backup captured and sealed in Coldline GCS with 7yr retention lock before shutdown. - 2026-06-10: legacy
in_formfeedMySQL found on PUBLIC IP34.171.123.238— ownership unconfirmed. Decommission precondition: drift check viafrom_mysql.py --dry-run; authorized-networks check needed.
Status site automation — Phase B¶
- BigQuery billing export, live current-month spend, trend, forecast.
- Requires GPI budget approval for BigQuery storage + query cost (est. <$5/mo).
Status site automation — Phase C¶
- Per-service cost attribution, budget alerts.
- Depends on Phase B (BigQuery export) landing.
T5-EXPANDED — Portal static-debt retirement¶
- Filed: 2026-06-10, per read-only audit of both portals (status + SOC).
- Sequencing: forms 2.5(c) Gate 5 has now landed (go-live 2026-07-23), so this is no longer gated behind it. P0 truth-fixes may interleave with the forms cutover/hardening workstreams; the remainder queues behind them.
- Phases:
- P0 — truth-fixes:
- SOC posture fail-open bug: unreachable host renders green "Compliant"; hardcoded auditd/SELinux/firewall columns.
- Stale-wrong Exec risk register — DRP/IRP marked "not documented" but both exist.
- "All 8 services" undercount.
- Reports
last_generatedalwaysNone. - Portal-backup coverage gap:
soc-site+forms-frontendunchecked.
- P1 — wire already-served data: server cards from
/api/status; discarded/api/carbon;/api/reportsfields; Governance link-out. - P2 — author missing canonical sources:
defense-in-depth.md,threat-model.md,risk-register.yaml(status & SOC registers currently DISAGREE), compliance-scores YAML — then build-time render runbooks/redblue/compliance (fixes SOC missing rb-006/007). - P3 — new collectors: VPN, DNS serial, Prometheus range charts, posture score, FLEET→inventory, Cloud Run table from inventory (overlaps forms Gate 5).
- P4: BigQuery billing (= old Status-site Phase B, GPI-gated), control matrix, git-log audit trail.
- P0 — truth-fixes:
- Standing rule (recorded 2026-06-10): all portal content must be LIVE or DERIVED — never hardcoded. The coverage standard is to be extended to require that elements RENDER FROM sources, not merely that endpoints exist — the audit found backends over-serving and frontends discarding (e.g.
/api/carbonfetched then thrown away).
Cross-cutting / lessons learned¶
Cloud Run env vars + Cloud Build deploys¶
Cloud Run env vars set out-of-band via gcloud run services update --update-env-vars are wiped on every Cloud Build deploy IF the cloudbuild.yaml uses --set-env-vars (destructive replace) instead of --update-env-vars (merge). Discovered 2026-04-27 in forms-backend/cloudbuild.yaml — fixed in commit 1ca4461. Other three backends (status, soc, security) don't have this bug because their cloudbuild.yaml files don't pass any env-var flag.
Lesson: any new backend cloudbuild.yaml should either omit env-var flags entirely (preserve) or use --update-env-vars (merge). Never --set-env-vars unless the deploy is intentionally the source of truth for ALL env vars.
Cloud SQL access for developer-side diagnostics¶
gpus-forms-db is private-IP only (10.34.0.3). Reaching it requires presence inside the gpus-infra VPC. Tested 2026-04-28:
- Laptop direct: blocked (no VPN to forms-db's service-peering range)
- Cloud Shell +
cloud-sql-python-connector: blocked (timeout to private IP from Cloud Shell's managed network) - Cloud Shell +
cloud-sql-proxy --private-ip: blocked (proxy bound locally fine but the dial to10.34.0.3:3307timed out)
Workable paths for developer-side diagnostic queries:
- SSH into MAPLE/OAK/CEDAR (all in
gpus-infraVPC), run script there. Caveat: that VM's service account principal needs Postgres-side grants (CLOUD_IAM_USER+SELECT). - Add VPC peering between Cloud Shell's project network and
gpus-infra(administrative work, deferred).
Lesson: Phase 2.5 implementation work needs MAPLE-based or peered DB access established as a prerequisite. inspect_actions.py (forms-backend/migrate/, committed 9eecc35) is ready to run from any in-VPC environment.
Memory entries describing "current bugs" age fast¶
Diagnostic memos written during one session describe state at that moment, not current state. The 2026-04-22 count-drift memo described a real bug in auth.py — fixed one day later in commit dd810e0 — but the memo persisted in memory and led to a Phase 2.5(a) design draft that proposed re-fixing the already-fixed bug. Code surfaced the staleness by reading current auth.py before any edits.
Lesson: Read-first discipline includes reading current code, not just current memory. When a memory entry describes a bug, verify the bug still exists by reading the affected module's current state (and grep recent commits for likely fix language).
Design doc accuracy under read-first discipline (2026-04-30)¶
Three design-doc bugs were caught by Code's "STOP and tell me if anything doesn't fit" gate before Commit 3 shipped:
_audithelper signature mismatch with actualAuditLogmodel- Invented
audit_actionenum value not in schema - Assumed SQLAlchemy relationship that wasn't declared
All three would have crashed the handler at runtime (or worse, the third would have silently produced submissions with NO field rows).
Lesson: design docs that "mirror Phase 1 patterns" must be drafted from current reads of those patterns, not from memory of earlier reads. The three failures had a common root: the doc described what was remembered of Phase 1's behavior, not what Phase 1 actually does at the line number the design doc claims to mirror. Verifying the source pattern at draft time would have caught all three.
Read-first discipline doesn't end at "read the file once during investigation." It applies again at every implementation moment that references that file's content.
MAPLE access for ad-hoc DB queries (2026-04-30)¶
Established and verified working pattern for ad-hoc Cloud SQL inspection from MAPLE (Phase 2.5 implementation work depends on this for diagnostic queries):
- SSH user:
cloudadmin(NOTmonitadmin, which is SUN/WIND only). - MAPLE has
cloud-sql-proxyv2 +psqlpre-installed. - MAPLE does NOT have
python3.11orgit— Python connector path requiressudo dnf install. - Postgres user
maple-agent@gpus-infra.iamexists with SELECT grants onsubmissions,submission_fields,audit_log(no GRANT needed).
Standard one-liner:
ssh cloudadmin@maple "cloud-sql-proxy --auto-iam-authn --private-ip \
gpus-infra:us-central1:gpus-forms-db &" && sleep 5 && \
psql "host=127.0.0.1 port=5432 dbname=gpus_forms user=maple-agent@gpus-infra.iam sslmode=disable" \
-c "<query>"
inspect_actions.py at forms-backend/migrate/ can run from MAPLE only if python3.11 + git are installed first. Defer Python install until there's a real need beyond what proxy + psql handles.
Hypothesis falsification via empirical data (2026-05-07)¶
α phase pulldown regression demonstrated a clean hypothesis-disconfirmation chain:
- Symptom: yes/no pulldowns failing "Could not load options"
- Initial hypothesis: missing pulldown rows in DB
- Hypothesis falsified: DB query showed all 7 failing
pulldown_names exist inpulldownstable with proper["Yes", "No"]values - Pattern in failing data: all 7 names contained
/ - New hypothesis: Flask string converter doesn't match
/; URL decoding splits the path - Fix: change
<name>to<path:name>in route decorator - Verified: dropdown opens with Yes/No options
Lesson: when a fix idea seems obvious ("add the missing pulldown rows"), check the data first. The obvious fix shipped without the falsification step would have been a no-op (rows already exist) and the bug would persist with diagnostic time wasted plus user trust diminished.
This pattern is reusable: when an outage looks like "data is missing," verify by query before shipping the inverse ("add the data"). The reverse pattern — when data exists but isn't being read — points at a different layer (routing, auth, encoding, serialization).
Operational learnings — 2026-06-10¶
- L2TP stale tunnel, 3rd occurrence: tunnel reports "up" while the path is dead. Health checks must probe internal IPs, not tunnel state.
- FortiClient
utun6route confound: FortiClient's interface can shadow routes and confuse VPN path diagnosis — rule it out before blaming the site-to-site tunnel. - IAP 4003 = host-level: the IAP-to-MAPLE 4003 backend-fail is host-level, not IAP edge. Promotes the T3 item to IAP-as-break-glass framing (see Summary).
- Legacy
in_formfeedMySQL on PUBLIC IP34.171.123.238: ownership unconfirmed. Decommission item filed under T5 MySQL decom with drift-check precondition (from_mysql.py --dry-run); needs authorized-networks check.
Program work register — schema v1 (opened 2026-08-13)¶
This section uses the schema below. Nothing above it does, and nothing above it was changed. The tiered tables, the T1–T5 sections, the Completed table and the change log keep their original
Planned / In Progress / Blocked / Done / Deferred / Filedvocabulary and their original wording. No pre-existing row in this document was edited, reworded, re-statused or removed — including the rows markedDonethat predate the status rule. Those are flagged, not changed, in Flag List — Unverified Done Rows; re-adjudication is a later pass tracked asGOV-008.Schema:
id · title · register · status · priority · owner · evidence · acceptance test · blocks · blocked_by · date_raised · date_verified.Status vocabulary for this section only:
open·in progress·blocked·done.donemeans verified end-to-end via the real production path against live state — authored, present, loaded or validated is not done. No row in this section isdone.
evidencerecords what was read, where, and on what date.acceptance teststates in advance the observable condition that closes the row. A row missing either is marked⚠ INCOMPLETEon that field and is not filled with a plausible guess. Most rows in this section are incomplete on the acceptance test; see Incomplete rows.
Two ID series¶
| Series | Meaning |
|---|---|
PRG- |
Program work items — belong to the program queue |
HB- |
Host-level blockers — belong to a host, not to the program |
HB- rows are the four items that block a specific host's disposition. They are
deliberately not filed as program items: filing them there would make the
whole disposition pass appear blocked on four unrelated questions, when in fact
each blocks exactly one host.
⚠ The host rows these four attach to do not exist yet.
gannet,emu,ostrich,catbirdandphoebeappear nowhere ininventory.yaml(checked 2026-08-13, zero matches each). Onlycatbirdandphoebeappear anywhere underdocs/at all, and only as/32allowlist entries atcompliance/iar.md:381-382. There is therefore no host row to attach anHB-blocker to. EachHB-row below names its host inblocksand is parked untilPRG-001/PRG-002create the host rows. This is the one place the instruction "attach them to the host rows they block" could not be carried out as written, and the reason is a missing anchor, not a judgement call.
Index — program work items¶
| ID | Title | Status | Priority | Blocked by |
|---|---|---|---|---|
| PRG-001 | Stage 1c — enumerate Meraki org 395909 and all three Synology units | open | not assessed | — |
| PRG-002 | Stage 2 — diff enumerated.yaml vs inventory.yaml, then disposition pass |
blocked | not assessed | PRG-001 |
| PRG-003 | Document the previously unlisted Cloud Run services | open | not assessed | — |
| PRG-004 | Identify 2 undocumented Cloud SQL instances in gpusa |
open | not assessed | — |
| PRG-005 | Scope 18 buckets in gpus-it against the "empties and retires entirely" commitment |
open | not assessed | — |
| PRG-006 | Unattached disks and 47 static IPs — billing impact | open | not assessed | — |
| PRG-007 | 35 orphan DNS names + 18 orphan Puppet node definitions | open | not assessed | — |
| PRG-008 | phoenix into inventory.yaml and under a liveness check |
open | not assessed | GOV-004 (derived) |
| PRG-009 | desert, river, star — three powered-off VMs on flower, disposition unknown |
open | not assessed | — |
| PRG-010 | water hosts zero VMs and is mis-documented as hosting ocean |
open | not assessed | — |
| PRG-011 | ~111 workstation DHCP reservations against 15 active leases | open | low — "not urgent" (as supplied) | — |
| PRG-012 | Billing export — highest-value cost action, not yet pulled | open | highest-value cost action (as supplied) | — |
Index — host-level blockers¶
| ID | Title | Blocks | Status | Priority |
|---|---|---|---|---|
| HB-001 | Tamr support on Rocky 9 — vendor question | gannet |
open | not assessed |
| HB-002 | GPI consumers of us.gl3? |
emu / ostrich rebuild |
open | first on the critical path (as supplied) |
| HB-003 | sssd/nslcd check across gpusa |
catbird delete |
open | not assessed |
| HB-004 | phoebe's real consumers |
phoebe disposition |
open | not assessed |
PRG-001 — Stage 1c: enumerate Meraki org 395909 and all three Synology units¶
| Field | Value |
|---|---|
| id | PRG-001 |
| title | Stage 1c — enumerate Meraki org 395909 and all three Synology units |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ Partial. Supplied: estate total 455, which "currently excludes network edge and storage layer entirely". The figure is specific but its source artifact and read date were not given. To complete: the Stage 1 output the 455 comes from, with its date. |
| acceptance test | ⚠ INCOMPLETE — none supplied. "Enumerate X" states the work, not the observable condition that closes it. Not guessed. To complete: the enumeration artifact containing a stated count of Meraki devices for org 395909 and all three Synology units. |
| blocks | PRG-002 |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Related: VLN-021 covers the Meraki org's admin access levels; this row covers
its device inventory. Same org (395909), different question.
PRG-002 — Stage 2: diff enumerated.yaml against inventory.yaml, then disposition pass¶
| Field | Value |
|---|---|
| id | PRG-002 |
| title | Stage 2 — diff enumerated.yaml against inventory.yaml, then run the disposition pass |
| register | Program work items |
| status | blocked |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ INCOMPLETE — none supplied. No artifact, path or date was given for this row. Not guessed. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | HB-001, HB-002, HB-003, HB-004 — this pass is what creates the host rows those four attach to |
| blocked_by | PRG-001 (stated) |
| date_raised | 2026-08-13 |
| date_verified | — |
PRG-003 — Document the previously unlisted Cloud Run services¶
| Field | Value |
|---|---|
| id | PRG-003 |
| title | Document the previously unlisted Cloud Run services: gpus-security-backend, gpus-soc-site, gpus-status-backend — and reconcile the status of gpus-forms-clamav-worker |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | Confirmed first-hand 2026-08-13 against inventory.yaml (repo root): gpus-security-backend 0 matches, gpus-soc-site 0 matches, gpus-status-backend 0 matches — all three genuinely absent. gpus-forms-clamav-worker is present, as cloud_services.gpus_forms_clamav_worker (underscore form). The cloud_services block holds nine entries, all forms/ClamAV-related. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. Note a bare "present in inventory.yaml" would be a weak test here: GOV-003 establishes that the cloud-services-render sentinel satisfies portal presence for cloud_services entities unconditionally, so an entry can be added and pass coverage without any portal actually rendering it. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Correction to the item as supplied. It was filed as "the 4 previously unlisted Cloud Run services" including
gpus-forms-clamav-worker. That service is listed. The count is 3 unlisted, not 4. The item is recorded with the corrected count rather than the supplied one, because the supplied count is checkable and wrong; nothing was merged or reworded beyond that.
PRG-004 — Identify 2 undocumented Cloud SQL instances in gpusa¶
| Field | Value |
|---|---|
| id | PRG-004 |
| title | Identify 2 undocumented Cloud SQL instances in gpusa; retention and backup obligations unknown |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ INCOMPLETE — none supplied. The count (2) and the project (gpusa) are given, but no listing, command output or read date. Not guessed. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/gpusa-it-infrastructure-306400__sql__rchhetry-at-greenpeace.org.json. Confirmed: 2 Cloud SQL instances in gpusa. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
"Retention and backup obligations unknown" is the material part: an
undocumented database is also an unclassified one. Note VLN-017 — gcloud
list verbs exit 0 on permission denial — means any enumeration of gpusa that
did not check stderr may have under-reported.
PRG-005 — Scope 18 buckets in gpus-it against the "empties and retires entirely" commitment¶
| Field | Value |
|---|---|
| id | PRG-005 |
| title | Scope the 18 buckets in gpus-it against the "empties and retires entirely" commitment |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ Partial. Count (18) and project (gpus-it) supplied; no listing or read date. The "empties and retires entirely" commitment is referenced but its source document was not cited. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/gpus-it-infrastructure__buckets__rchhetry-at-greenpeace.org.json. Confirmed: 18 buckets in gpus-it. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
PRG-006 — Unattached disks and 47 static IPs; billing impact¶
| Field | Value |
|---|---|
| id | PRG-006 |
| title | Unattached disks (gpus-infra 7 disks : 3 instances; gpus-it 15 : 9) and 47 static IPs — billing impact |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied. (Billing impact is stated as a consequence, not as a priority.) |
| owner | unassigned |
| evidence | ⚠ Partial. Ratios and counts supplied and specific — gpus-infra 7 disks against 3 instances, gpus-it 15 against 9, and 47 static IPs — but no listing, command output or read date. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): All four figures confirmed. raw/gpus-infra__disks__*.json = 7 disks against raw/gpus-infra__instances__*.json = 3 instances; raw/gpus-it-infrastructure__disks__*.json = 15 against 9 instances; and the three *__addresses__*.json captures total 47 static IPs across gpus-infra, gpus-it-infrastructure and gpusa-it-infrastructure-306400. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
The billing impact here is quantifiable but not yet quantified — PRG-012
(billing export) is what would size it. No dependency is asserted between them
because none was stated.
PRG-007 — 35 orphan DNS names and 18 orphan Puppet node definitions¶
| Field | Value |
|---|---|
| id | PRG-007 |
| title | 35 orphan DNS names in cloud.us.gl3 plus 18 orphan Puppet node definitions — likely one uncleaned decommission wave |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ Partial. Counts supplied (35 DNS names, 18 Puppet node definitions) with the zone named; no dump, command output or read date. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task1-arecord-crossref-20260813.txt (captured 2026-08-13) for the DNS side, and raw/puppet__nodes-parsed__phoenix-ssh-rchhetry.json for the node definitions. See GOV-018 — the 35 figure this row inherits was superseded twice and now stands at 50. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
"Likely one uncleaned decommission wave" is recorded as supplied — it is a
hypothesis, not a finding, and the row does not treat it as established.
Distinct from VLN-020, which covers five node definitions that can never
match a certname; that is a naming defect, this is orphaned scope. The two may
overlap and neither row assumes it.
PRG-008 — phoenix into inventory.yaml and under a liveness check¶
| Field | Value |
|---|---|
| id | PRG-008 |
| title | Bring phoenix into inventory.yaml and under a liveness check |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ Partial — supplemented first-hand. No source or date was supplied. Confirmed 2026-08-13: phoenix has zero matches in inventory.yaml, consistent with the item as filed. |
| acceptance test | phoenix present in inventory.yaml and covered by a liveness check. (Derived from the item's own stated target state, not invented — but see the blocker below, which means the second half is not currently satisfiable.) |
| blocks | — |
| blocked_by | GOV-004 — derived, not supplied. GOV-004 records that no GL5 liveness monitoring exists; the second half of this acceptance test cannot be met until it does. Flagged as a derived dependency so it can be rejected on review. |
| date_raised | 2026-08-13 |
| date_verified | — |
PRG-009 — desert, river, star: three powered-off VMs on flower, disposition unknown¶
| Field | Value |
|---|---|
| id | PRG-009 |
| title | desert, river, star — three VMs discovered on flower, powered off, no DNS, no DHCP presence; disposition unknown |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ Partial. Four specific observations supplied — discovered on flower, powered off, no DNS, no DHCP presence — but no source artifact or read date. Confirmed first-hand 2026-08-13 that none of desert, river or star appears in inventory.yaml (zero matches each). Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-esxi-inventory-20260813.txt, the same authenticated vSphere capture that carries the VLN-013 build numbers. |
| acceptance test | ⚠ INCOMPLETE — none supplied. "Disposition unknown" states the open question, not the condition that closes it. Not guessed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Note the host these sit on: flower is the hypervisor VLN-013 identifies as
running the 2018 GA 6.7 build that has never been patched.
PRG-010 — water hosts zero VMs and is mis-documented as hosting ocean¶
| Field | Value |
|---|---|
| id | PRG-010 |
| title | water hosts zero VMs and is mis-documented as hosting ocean; capture CPU, RAM and datastore capacity to test consolidating fire and flower onto water |
| register | Program work items |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | Confirmed first-hand 2026-08-13 in inventory.yaml:146-152. The water entry reads desc: First on-prem ESXi host — currently hosts Ocean (KACE SMA) (:149) and vms: [ocean] (:152) — the mis-documentation is in the inventory itself, not only in prose. ocean is separately defined at inventory.yaml:988 with fqdn: ocean.wdc.us.gl3. The same entry also carries hardware_model: "" # TODO: walk-around fill-in (:151), i.e. the CPU/RAM/datastore capture this row calls for has an existing empty slot waiting for it. |
| acceptance test | CPU, RAM and datastore capacity captured for water, sufficient to test consolidation of fire and flower onto it. (Stated in the item; recorded as given.) |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Second occurrence of the VLN-013 misattribution, found while evidencing this row.
inventory.yaml:149also recordshypervisor: VMware ESXi 6.7forwater. PerVLN-013's authenticated vSphere API evidence, water is on 8.0.3 (build24022510).VLN-013's acceptance test names onlyVLN-004and will not catch this line. Raised here rather than silently widening VLN-013's acceptance test; see the open question in the session log entry for 2026-08-13.
PRG-011 — ~111 workstation DHCP reservations against 15 active leases¶
| Field | Value |
|---|---|
| id | PRG-011 |
| title | ~111 registered workstation DHCP reservations against 15 active leases — ghost-asset exposure |
| register | Program work items |
| status | open |
| priority | low — recorded from the supplied qualifier "not urgent". This is the only priority signal given for any PRG- row and is carried as stated, not converted to a tier. |
| owner | unassigned |
| evidence | ⚠ Partial. Counts supplied (~111 reservations, 15 active leases); the ~ is carried through rather than resolved. No source artifact or read date. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task3-dhcp-parsed-20260813.txt and raw/dhcp-static-reservations-sky-20260813.txt. The lease side confirms exactly: the sky census reads active 15 against backup 162, free 88, TOTAL 265 distinct IPs. ⚠ Discrepancy — figure did not reproduce. The static-reservation capture contains 126 host declarations on sky, not the ~111 stated. 126 includes infrastructure hosts — the first stanza is host sky itself — so ~111 is presumably workstations after filtering, but the filter is not recorded. The stated figure is left unchanged pending that definition. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
PRG-012 — Billing export: highest-value cost action, not yet pulled¶
| Field | Value |
|---|---|
| id | PRG-012 |
| title | Billing export — flagged as the highest-value cost action, not yet pulled |
| register | Program work items |
| status | open |
| priority | highest-value cost action — recorded as supplied. Carried verbatim rather than mapped to a tier, since no tier was given. |
| owner | unassigned |
| evidence | ⚠ INCOMPLETE — none supplied. No artifact, path or date. "Not yet pulled" is a state, not evidence of who established it or when. Not guessed. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Note the existing backlog rows Status site automation — Phase B (BigQuery billing export) and Phase C (per-service cost attribution) sit at T5 with 2026-10 / 2026-11 targets. This row is filed as the highest-value cost action. That tension is recorded, not resolved — reconciling the two would mean re-statusing a pre-existing row, which this pass does not do.
Host-level blockers¶
Each row below blocks exactly one host. None is a program blocker. The host
rows they attach to do not exist yet — see the warning at the top of this
section — so blocks names the host and its pending disposition rather than a
row id.
HB-001 — Tamr support on Rocky 9¶
| Field | Value |
|---|---|
| id | HB-001 |
| title | Tamr support on Rocky 9 — vendor question |
| register | Host-level blocker |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ INCOMPLETE — none supplied. Not guessed. |
| acceptance test | ⚠ INCOMPLETE — none supplied. A vendor answer is the implied output, but the condition that makes it sufficient was not stated. Not guessed. |
| blocks | gannet — host row does not exist; created by PRG-002 |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
HB-002 — GPI consumers of us.gl3?¶
| Field | Value |
|---|---|
| id | HB-002 |
| title | GPI consumers of us.gl3? |
| register | Host-level blocker |
| status | open |
| priority | first on the critical path — recorded as supplied |
| owner | unassigned |
| evidence | ⚠ INCOMPLETE — none supplied. Not guessed. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | emu / ostrich rebuild — host rows do not exist; created by PRG-002 |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
This is the highest-priority signal supplied anywhere in the program list, and
it is attached to a host row that does not yet exist. Both emu and ostrich
are the DNS masters carrying the VLN-015 NOPASSWD sudo finding.
HB-003 — sssd/nslcd check across gpusa¶
| Field | Value |
|---|---|
| id | HB-003 |
| title | sssd/nslcd check across gpusa |
| register | Host-level blocker |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ INCOMPLETE — none supplied. Not guessed. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | catbird delete — host row does not exist; created by PRG-002. catbird currently appears only as a /32 allowlist entry at compliance/iar.md:382, marked "Remove after Phase 4 cutover" |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
HB-004 — phoebe's real consumers¶
| Field | Value |
|---|---|
| id | HB-004 |
| title | phoebe's real consumers |
| register | Host-level blocker |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | ⚠ INCOMPLETE — none supplied. Not guessed. Evidence now available — phoebe-capture/ under ~/estate-enum-2026-08, captured 2026-08-13. This row was INCOMPLETE on evidence; the directory holds Apache vhost configs and access/error logs for six candidate consumers: budget.us.gl3, forms.us.gl3, lam.cloud.us.gl3, phoebe.cloud.us.gl3, redirects.cloud.us.gl3 and sharperlight.us.gl3, plus messages-tail.txt. forms.us.gl3 carries the largest logs — four weekly rotations through 2026-07 totalling ~5 MB — and lam.cloud.us.gl3 has rotations running back to 2025-12. This is raw material, not an answer: the logs have not been analysed, so which of the six are live consumers versus redirect stubs is still open, and the acceptance test remains INCOMPLETE. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not guessed. |
| blocks | phoebe disposition — host row does not exist; created by PRG-002. phoebe currently appears only as a /32 allowlist entry at compliance/iar.md:381, marked "Remove after Phase 4 cutover" |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
GOV-004 names this host as the origin of the missing-reconciliation-control
class ("the phoebe mechanism"). No blocks relation is asserted between them
because none was stated.
Incomplete rows — program register¶
Recorded as incomplete rather than completed with a plausible guess.
| ID | Evidence | Acceptance test |
|---|---|---|
| PRG-001 | ⚠ partial — 455 figure, no source or date | ⚠ missing |
| PRG-002 | ⚠ missing | ⚠ missing |
| PRG-003 | ✅ confirmed first-hand | ⚠ missing |
| PRG-004 | ⚠ missing | ⚠ missing |
| PRG-005 | ⚠ partial — counts, no source or date | ⚠ missing |
| PRG-006 | ⚠ partial — counts, no source or date | ⚠ missing |
| PRG-007 | ⚠ partial — counts, no source or date | ⚠ missing |
| PRG-008 | ⚠ partial — supplemented first-hand | ✅ derived from the item's stated target |
| PRG-009 | ⚠ partial — supplemented first-hand | ⚠ missing |
| PRG-010 | ✅ confirmed first-hand | ✅ as supplied |
| PRG-011 | ⚠ partial — counts, no source or date | ⚠ missing |
| PRG-012 | ⚠ missing | ⚠ missing |
| HB-001 | ⚠ missing | ⚠ missing |
| HB-002 | ⚠ missing | ⚠ missing |
| HB-003 | ⚠ missing | ⚠ missing |
| HB-004 | ⚠ missing | ⚠ missing |
13 of 16 rows lack an acceptance test. 6 of 16 lack evidence entirely. The pattern is consistent and worth naming: the program list was supplied as a list of observations and questions, which is what it is good at. Acceptance tests are decisions about what "finished" means, and those had not been taken for most of these items. Every one is recoverable — the fastest route is one pass stating the close condition per row.
Program work register — 2026-08-17 load (PRG-013 – PRG-026)¶
Continues the schema v1 section above. Same schema, same status vocabulary, same rule:
donemeans verified end-to-end via the real production path against live state. No row in this load isdone, includingPRG-019, which records a disposition decision rather than a verification.Source:
priorities/backlog-2026-08-17.mdv1.0, items B-25 – B-38. Owner is R. Chhetry throughout exceptPRG-015, where the source text names Rob MacMillan as co-decider. No severity or priority was stated for any row in this load, so every row readsnot assessed; none is inferred. No acceptance test was stated for any of the fourteen, so all fourteen are marked INCOMPLETE on that field.
Index¶
| ID | Title | Status | Priority | Notes |
|---|---|---|---|---|
| PRG-013 | Consolidate fire and flower onto water | open | not assessed | RAM headroom thin |
| PRG-014 | WS2025 golden image family is a dependency, not a future item | open | not assessed | second image lane required |
| PRG-015 | magpie contradiction — keep vs. project committed to retire entirely | open | not assessed | decision: R. Chhetry + Rob MacMillan |
| PRG-016 | Seven unmanaged gpus-it hosts unexplained | open | not assessed | — |
| PRG-017 | Two undocumented Meraki networks plus three unassigned APs | open | not assessed | — |
| PRG-018 | Meraki WDC edge is an MX95 HA pair; the asset registry says MX100 | open | not assessed | registry is wrong |
| PRG-019 | Vendor transfer stack — DECOMMISSION | open | not assessed | decision, not a verification |
| PRG-020 | Stage 3 result — 24 of 25 gpusa VMs moved UNKNOWN → DERIVED | open | not assessed | gated by GOV-011 |
| PRG-021 | The estate spans two GCP projects for Puppet purposes | open | not assessed | — |
| PRG-022 | gpusa UNKNOWN count is 154, not 146 | open | not assessed | — |
| PRG-023 | gpus-dist staleness profile — 95% over three years old |
open | not assessed | — |
| PRG-024 | duck IP contradiction confirmed from a second source | open | not assessed | corroborates VLN-019 |
| PRG-025 | Enumeration surfaces missing from Stage 1 | open | not assessed | — |
| PRG-026 | Seven entities declared with no identifying data, never observed live | open | not assessed | — |
PRG-013 — Consolidate fire and flower onto water¶
| Field | Value |
|---|---|
| id | PRG-013 |
| title | Consolidate fire and flower onto water |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Authenticated vSphere API, 2026-08-13. water: Xeon Gold 5416S, PowerEdge R660xs, 63.5 GB RAM at 3.8%, 1.66 TB datastore at 0.1%, zero VMs. fire: Xeon E5640 (2010 silicon), R610, 48 GB at 70%, hosts all four core servers. flower: Xeon E5-2697 v3, R630, 128 GB at 11.6%, local datastore 80.6% full. Combined RAM in use across fire and flower is ~48 GB against water's 63.5 GB. All three already mount the same vmstorage NFS datastore, so compute moves without storage moving. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): capacity figures at raw/task3-esxi-capacity-20260813.txt, inventory at raw/task2-esxi-inventory-20260813.txt. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Retires two EOL 6.7 hosts onto supported hardware already owned. Recorded as supplied: RAM headroom is thin — per-VM allocations are needed before committing. ~48 GB into 63.5 GB leaves little margin, and the figure is current usage rather than allocation.
Depends in practice on VLN-023: the shared vmstorage NFS datastore that makes
this move cheap is the same array currently running degraded with no spare.
Consolidating three hosts onto one storage path does not change that array's
state, but it does concentrate what depends on it.
PRG-014 — WS2025 golden image family is a dependency, not a future item¶
| Field | Value |
|---|---|
| id | PRG-014 |
| title | WS2025 golden image family is a dependency, not a future item — a second image lane is required |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stage 2 Part B, 2026-08-17. inventory.yaml declares duck, grebe and nfs-gw as REPLACE targeting Windows Server 2025, while the golden-image program currently plans a single Rocky 9 family. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
The framing is the point and is carried as supplied: this is a dependency of the current program, not a later phase. Three declared REPLACE targets have no image lane to be replaced onto.
PRG-015 — magpie contradiction¶
| Field | Value |
|---|---|
| id | PRG-015 |
| title | magpie declared disposition: keep inside a project committed to empty and retire entirely |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry and Rob MacMillan — the source row names this as a decision for both |
| evidence | Stage 2 Part B, 2026-08-17. inventory.yaml declares disposition: keep, role_status: confirmed-data-team, os: rhel-8, os_eol: 2029-05-31, live RUNNING — in gpus-it-infrastructure, which is committed to empty and retire entirely. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Recorded as supplied: either the commitment has an unwritten exception or
magpie needs a destination. Both branches are decisions, not investigations —
which is why this row names two owners and no acceptance test rather than a
research task. Note PRG-021: magpie is also one of the two hosts that make the
estate span two GCP projects for Puppet purposes.
PRG-016 — Seven unmanaged gpus-it hosts unexplained¶
| Field | Value |
|---|---|
| id | PRG-016 |
| title | Seven unmanaged gpus-it hosts unexplained — running, never Puppet-managed, unaccounted for in either repo |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stages 3 and 3b, 2026-08-17. duck, falcon, grebe, gull, kestrel, nfs-gw, woodpecker — running, never Puppet-managed, and nothing in either repo accounts for them. falcon, gull and kestrel appear only as DNS records; grebe, nfs-gw and woodpecker appear nowhere. falcon's public record (35.223.86.92, 2026-07-27) is the newest content in gpus-dist. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Three of the seven — duck, grebe, nfs-gw — are the same hosts PRG-014 records
as REPLACE targets needing a WS2025 lane, and duck carries the IP contradiction
in VLN-019 / PRG-024. That falcon's DNS record is the newest content in an
otherwise 2017-era repo (PRG-023) is recorded as supplied, without inference
about what it means.
PRG-017 — Two undocumented Meraki networks plus three unassigned APs¶
| Field | Value |
|---|---|
| id | PRG-017 |
| title | Two undocumented Meraki networks — DC Apartments and Major Gifts — plus three unassigned APs |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Meraki API, 2026-08-13. DC Apartments (4 APs — MR33×3, MR52×1) and Major Gifts (systems-manager only), plus three unassigned APs (MR32×2, MR33×1). The working brief says three sites; there are five networks and 32 devices, 29 assigned. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
This closes out the Meraki half of PRG-001 (Stage 1c) as an enumeration
result, though PRG-001's own acceptance test remains unstated and that row is
not re-statused here.
PRG-018 — Meraki WDC edge is an MX95 HA pair; the registry says MX100¶
| Field | Value |
|---|---|
| id | PRG-018 |
| title | Meraki WDC edge is an MX95 HA pair with warm spare enabled — the asset registry says MX100 and is wrong |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Meraki API, 2026-08-13. MX95 HA pair, warm spare enabled, primarySerial Q2XN-V4XE-UQKX, firmware wired-26-1-5. The asset registry says MX100. WAN1 virtual IP 38.140.146.68 matches the documented VPN endpoint; the physical MXs hold .66 and .67. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Another instance of the GOV-017 pattern: the entity is present in the
registry and its declared attributes are wrong. The virtual-vs-physical IP
detail matters for any firewall or peer configuration derived from the registry.
PRG-019 — Vendor transfer stack: DECOMMISSION¶
| Field | Value |
|---|---|
| id | PRG-019 |
| title | Vendor transfer stack — DECOMMISSION (relay module, transfer crons, storehouse hosts) |
| register | Program work items |
| status | open — a disposition decision is recorded; the decommission itself is not started and certainly not verified |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | R. Chhetry decision, 2026-08-17: stale code, not required, removed with the move to Rocky Linux. Covers the relay module, the transfer crons and the storehouse hosts. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| evidenced_by | VLN-035 — stated in the source row. dataflow.us.gl3 resolving to a host that does not exist supports the decommission. |
| date_raised | 2026-08-17 |
| date_verified | — |
This is a decision, not a verification — do not record as done
Carried verbatim from the source. A disposition decision by the owner is an
input to work, not evidence that work happened. Under the status rule,
done would require the stack removed and verified against live state.
Status is open.
Premise corrections recorded with the decision. These correct the picture the decision was taken against and are kept because they change what the decommission actually has to remove:
- There are 27 cron resources but only 13 potentially-active transfers —
13 are declared twice (
ensure => presentunderif($live == true),ensure => absentin the else branch) andpushPSIis unconditionally absent. - Three scripts do not do what their names say.
pushFacterperforms no transfer at all and only downcases filenames locally.pushFPRexits 0 at line 7 ("Turning this script off for now") and has fired every 30 minutes doing nothing.pushPSIis explicitly deactivated. - Scripts live in the Puppet control repo at
modules/relay/files/sbin/with-roosterand-whistlervariants;sourceselect => firstpicks the hostname variant and whistler is the RUNNING one. Endpoint routing comes frompushtab.
pushFPR is the second instance of LOADED ≠ REACHABLE in this row alone — see
M-01. Note that this disposition does not close VLN-024: the plaintext
credentials in pushtab survive deletion of the code that reads them.
PRG-020 — Stage 3 result: 24 of 25 gpusa VMs moved UNKNOWN → DERIVED¶
| Field | Value |
|---|---|
| id | PRG-020 |
| title | Stage 3 classification result — 24 of 25 gpusa VMs moved UNKNOWN → DERIVED |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stage 3, 2026-08-17. 24 of 25 gpusa VMs moved UNKNOWN → DERIVED. Only quail remains UNKNOWN (TERMINATED, no node definition). Across all 44 node definitions: 38 DERIVED, 5 ASSERTED-STALE, 1 ASSERTED. kingfisher resolved to DERIVED via the testing modulepath, where postgresql is defined — puppetEnvironment => testing. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/gpusa-it-infrastructure-306400__instances__rchhetry-at-greenpeace.org.json — confirmed: 25 instances in gpusa, the denominator of the 24-of-25 claim. Classification output at raw/stage3-noderows.json and raw/stage3-class-resolution.json, with the testing-modulepath resolution that reclassified kingfisher at raw/stage3-class-resolution-testing.json. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | GOV-011 — derived. DERIVED is not enforced; the source row states this row is subject to that gap. |
| date_raised | 2026-08-17 |
| date_verified | — |
DERIVED is a hypothesis, not evidence
This row looks like progress and is recorded as a classification result,
not a verification one. Per GOV-011, thirteen RUNNING VMs with node
definitions have never produced a retained catalog — including phoenix
itself — so a DERIVED statement on such a host is a hypothesis. kingfisher
resolving only via the testing modulepath is worth watching for the same
reason.
PRG-021 — The estate spans two GCP projects for Puppet purposes¶
| Field | Value |
|---|---|
| id | PRG-021 |
| title | The estate spans two GCP projects for Puppet purposes — comparing node definitions against gpusa alone over-reports orphans by two |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stage 3, 2026-08-17. magpie and kingfisher have node definitions and live in gpus-it-infrastructure, not gpusa. Comparing node definitions against gpusa VMs alone over-reports orphans by two. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
This is the fact that corrected the middle orphan figure in GOV-018 (52 → 50).
Both hosts are already carried elsewhere: magpie in PRG-015, kingfisher in
PRG-020.
PRG-022 — gpusa UNKNOWN count is 154, not 146¶
| Field | Value |
|---|---|
| id | PRG-022 |
| title | gpusa UNKNOWN count is 154, not 146 |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | ⚠ Partial. Supplied: the count is 154, not 146, as read 2026-08-17. The date is given; no source artifact or query was named. To complete: the enumeration output the 154 was read from. ⚠ Still unsourced after searching the enumeration corpus 2026-08-17. No artifact stating the 154 figure was located under ~/estate-enum-2026-08. The row remains partial: the date is recorded, the source is not. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
One of the four counts M-08 records as overturned during these sessions.
PRG-023 — gpus-dist staleness profile¶
| Field | Value |
|---|---|
| id | PRG-023 |
| title | gpus-dist staleness profile — 900 of 946 files are more than three years old |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stage 3b per-file provenance, 2026-08-17. 900 of 946 files (95%) are more than three years old; 800 were last touched in 2017. The sole current content is DNS zones — zones/greenpeaceusa.org*.zone at 2026-07-27 is the newest file in the repo. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage3b-dist-provenance.txt and raw/stage3b-present-dates.json, with the file index at raw/stage3b-dist-index.json. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Recorded as supplied: DNS is the consistent exception to staleness across both
repos. That pattern matters for GOV-014's mirror-before-teardown constraint —
the one live thing in an otherwise dead repo is the thing most likely to be
missed if the repo is written off as stale.
PRG-024 — duck IP contradiction confirmed from a second independent source¶
| Field | Value |
|---|---|
| id | PRG-024 |
| title | duck IP contradiction confirmed from a second independent source |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stage 3b, 2026-08-17. gpus-dist forward and reverse zones both say 10.1.96.40; the live VM is 10.1.96.46; inventory.yaml says .46. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| corroborates | VLN-019 — stated in the source row. Second-source corroboration for duck.cloud.us.gl3 resolving to an address that is not duck's. |
| date_raised | 2026-08-17 |
| date_verified | — |
Two independent sources now agree the zone data is wrong and inventory.yaml is
right — a rare direction for this estate, and worth noting because most rows in
this load run the other way.
PRG-025 — Enumeration surfaces missing from Stage 1¶
| Field | Value |
|---|---|
| id | PRG-025 |
| title | Enumeration surfaces to add to any future reconciliation control |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stage 2, 2026-08-17. Pub/Sub topics, Cloud Scheduler jobs, host-resident services, and power devices were in no Stage 1 surface. Four declared Pub/Sub and Cloud Scheduler entities plus one Pub/Sub subscription (gpus_forms_clamav_worker_sub) could not be checked in either direction. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Feeds GOV-004 alongside GOV-016: that row specifies how a reconciliation
control must join, this one specifies what it must look at. Power devices
appear here and as GOV-015; the two are the same absence seen from the control
side and the estate side.
PRG-026 — Seven entities declared with no identifying data, never observed live¶
| Field | Value |
|---|---|
| id | PRG-026 |
| title | Seven entities declared with no identifying data and never observed live — unresolved in both directions |
| register | Program work items |
| status | open |
| priority / severity | not assessed — none supplied |
| owner | R. Chhetry |
| evidence | Stage 2 A3, 2026-08-17. gl5_firewall (no serial, no MAC, no IP), synology_controller, visuals_storage_exp_1 through _3, vmware_storage. Unresolved in both directions. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
These are the entities GOV-016's join-key constraint cannot help with: with no
serial, MAC or IP there is no key to join on at all. They can be neither
confirmed nor refuted by any reconciliation control as currently specified.
Program work register — 2026-08-20 load (PRG-027)¶
PRG-027 — Enumerate what else on MAPLE runs under Python 3.6.8¶
| Field | Value |
|---|---|
| id | PRG-027 |
| title | Enumerate every cron job, systemd unit and script on MAPLE that invokes the EOL Python 3.6.8 interpreter |
| register | Program work |
| status | open |
| priority | T1 — derived from a HIGH security finding |
| owner | R. Chhetry |
| evidence | MAPLE, 2026-08-20: python3 --version → Python 3.6.8 (EOL 2021-12-23); python3.11 --version → Python 3.11.13 present but not the default. /usr/bin/python3 -m pip freeze carried reportlab==3.6.8, Pillow==8.4.0, requests==2.27.1, google-cloud-storage==2.0.0. Four production security reports had executed under this interpreter, as root from the root crontab, from ~2026-04 until 2026-08-20 — see VLN-037. Those four are remediated. Nothing else on the host has been examined. The root crontab visibly also carries /usr/local/bin/gpus-cloud-backup.sh (daily 02:00 UTC), whose interpreter is unknown. |
| acceptance test | A written enumeration exists listing, for MAPLE: every root and user crontab entry, every enabled systemd unit, and every script under /opt, /usr/local/bin and /home/* that invokes python3, /usr/bin/python3, or a venv built from either — each classified as (a) confirmed not Python, (b) Python on 3.6.8 and needing migration, or (c) Python already on 3.11+. The enumeration is complete when every entry carries one of those three classifications and no entry is unclassified. Migration of what it finds is separate work and is NOT in this acceptance test — this row closes on knowing, not on fixing. |
| blocks | — |
| blocked_by | — |
| evidences | VLN-037 — this row is that finding's residual, raised as work rather than left in a residual field. |
| date_raised | 2026-08-20 |
| date_verified | — |
Why this is a row and not a residual line. VLN-037 states the residual accurately: the reports are off 3.6.8, the rest of the host is not, and what else runs under it is unknown. But a residual field has no owner and no acceptance test, so it reads as closed to anyone scanning the register — and this estate's recurring failure is precisely things that read as closed while nobody is looking at them (cf. VLN-010, VLN-011, and VLN-037 itself).
Suggested starting commands, offered as a starting point and not as the enumeration:
sudo crontab -l
for u in $(cut -d: -f1 /etc/passwd); do sudo crontab -l -u "$u" 2>/dev/null | sed "s|^|[$u] |"; done
ls -la /etc/cron.d/ /etc/cron.daily/ /etc/cron.hourly/
sudo systemctl list-units --type=service --state=running
sudo grep -rlE '(/usr/bin/)?python3([^.]|$)' /opt /usr/local/bin /etc/systemd/system 2>/dev/null
for v in $(sudo find /opt /home -name 'bin/python3' -path '*venv*' 2>/dev/null); do echo -n "$v: "; "$v" --version; done
The last line matters most and is the least obvious: a venv inherits the
interpreter it was built from. /opt/gpus-reports/venv was a 3.6.8 venv
until 2026-08-20 despite python3.11 being installed on the host, and looked
identical from the outside to a 3.11 one. Any other venv on MAPLE may be in the
same state, and ls will not tell you.
Program work register — 2026-09-01 load (PRG-029)¶
PRG-029 — Detect report silence: a freshness check over every latest.pdf¶
| Field | Value |
|---|---|
| id | PRG-029 |
| title | Check the age of gs://gpus-infra-backups-wdc/reports/<type>/latest.pdf against each report's own cadence, for all eight types, and alert when one is older than it should be. Makes a report that did not run detectable, which today it is not by any means at all. |
| register | Program work |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | GOV-033, 2026-09-01: the September executive monthly aborted at 08:00:01 on a 30 s read timeout to /api/soc and produced no PDF, no mail and no error line. Every candidate signal was enumerated on the host and every one failed — the wrapper's own log "ERROR: PDF generation failed" is unreachable behind set -euo pipefail, nothing parses /var/log/gpus-reports.log, there is no Prometheus metric or textfile collector on MAPLE, and cron's MAILTO=root fired into an unread 439 KB file (GOV-034). Of the three fixes GOV-033 recommends, this is the only one that would have caught THIS failure rather than making the next one louder: the other two improve what happens when the report runs, and this one notices when it does not. It also covers the two reports with a single named recipient each, where the reader has no baseline for "should have arrived by now". |
| acceptance test | A scheduled check reads the eight latest.pdf objects and alerts when any is older than its cadence allows (daily +1 day for approvals and volume; monthly +1 day for monthly, summary, prebooking; monthly-on-the-15th for newsletter; quarterly for quarterly; weekly for weekly). Demonstrated able to fail by pointing it at a deliberately stale object — or simply by running it against the current bucket, where reports/executive/latest.pdf should already be flagged from the 2026-09-01 miss. A check that runs and never alerts is not evidence it works; the demonstration is part of the test, per PRG-028. |
| blocks | — |
| blocked_by | Needs a decision on where it runs and where it alerts — the same open question as PRG-028, and the two should be answered together rather than twice. Deliberately NOT blocked on GOV-034: report silence should become detectable without first fixing cron mail on seven hosts, and routing this through a channel that is already known broken would reproduce the failure it exists to catch. |
| evidences | GOV-033 — this row is that finding's third recommendation, raised as work rather than left in a recommendation block. GOV-034 is the adjacent estate-wide gap and is deliberately a separate item. |
| date_raised | 2026-09-01 |
| date_verified | — |
Why this is a row and not a line inside GOV-033. Same reasoning as
PRG-027 and PRG-028, and it is the established pattern here rather than a
new one: the GOV- row records what is missing and can be closed honestly
once the finding is understood; the PRG- row carries the build, with an
owner and an acceptance test. GOV-028 → PRG-028 was the first instance.
Leaving this inside a recommendation block would give it no owner, no
acceptance test and no place in a weekly review — and a reader scanning the
gap register would see the failure documented and reasonably conclude
something was being done about it.
One scoping note worth keeping. It is tempting to widen this to "alert on
any failed cron job", which is GOV-034's territory and a much larger piece
of work. Resist that: the eight latest.pdf objects already exist, are
already written by the production path, and need no agent, no exporter and no
mail. It is the cheapest thing in this family that would have worked, and its
narrowness is the reason it can ship before the bigger question is answered.
Program work register — 2026-08-27 load (PRG-028)¶
PRG-028 — Alert on gpus-reports deploy drift, rather than making it merely visible¶
| Field | Value |
|---|---|
| id | PRG-028 |
| title | Run GOV-028 acceptance command (b) on a schedule and alert on mismatch, so deployed-vs-repo drift in gpus-reports is detected rather than detectable |
| register | Program work |
| status | open |
| priority | not assessed — none supplied |
| owner | unassigned |
| evidence | GOV-028 closed 2026-08-27: gpus-reports now carries a VERSION file and stamps the commit hash into every run's first stdout line, every degraded= line, the mailer's start line and every PDF footer. That makes drift visible. It does not prevent it, and nothing detects it unattended. The condition it fixes ran undetected for two days precisely because nobody was looking, and the guard does not change who is looking — it changes what they would see if they did. R5 has exactly one recipient. |
| acceptance test | A scheduled job runs the currency comparison — git log --oneline "$deployed"..origin/main -- 'gpus-reports/*.py' 'gpus-reports/report_cron.sh' 'gpus-reports/requirements.txt' ':!gpus-reports/test_*.py' — and emits an alert when the output is non-empty. Demonstrated able to fail by pointing it at a stale marker (e.g. 10a48be, which names four undeployed commits) and observing the alert fire. A job that runs and never alerts is not evidence it works; the demonstration is part of the test. |
| blocks | — |
| blocked_by | Needs a decision on where it runs and where it alerts. It cannot run on MAPLE alone — MAPLE has no clone of the repo, and a check that only reads the host cannot know what the host is behind. |
| evidences | GOV-028 — this row is that finding's residual, raised as work rather than left in a residual field. |
| date_raised | 2026-08-27 |
| date_verified | — |
Why this is a row and not a residual line. Same reasoning as PRG-027:
a residual field has no owner and no acceptance test, so it reads as closed to
anyone scanning the register. GOV-028 is closed and its closure is honest —
the guard was built, deployed and verified — but a reader who sees only
"drift guard: done" will reasonably conclude that drift is now handled. It is
not. It is legible, which is a different property, and the gap between the
two is exactly where the original two-day drift lived.
The awkward part, stated rather than left implicit. The check needs both sides — the deployed marker (on MAPLE) and the repo history (in a clone) — and no scheduled job today has both. That is the real work in this row; the comparison itself is one command.
Incomplete rows — 2026-08-17 load¶
| ID | Evidence | Acceptance test |
|---|---|---|
| PRG-013 | ✅ as supplied, dated | ⚠ missing |
| PRG-014 | ✅ as supplied, dated | ⚠ missing |
| PRG-015 | ✅ as supplied, dated | ⚠ missing |
| PRG-016 | ✅ as supplied, dated | ⚠ missing |
| PRG-017 | ✅ as supplied, dated | ⚠ missing |
| PRG-018 | ✅ as supplied, dated | ⚠ missing |
| PRG-019 | ✅ decision, dated | ⚠ missing |
| PRG-020 | ✅ as supplied, dated | ⚠ missing |
| PRG-021 | ✅ as supplied, dated | ⚠ missing |
| PRG-022 | ⚠ partial — date, no source named | ⚠ missing |
| PRG-023 | ✅ as supplied, dated | ⚠ missing |
| PRG-024 | ✅ as supplied, dated | ⚠ missing |
| PRG-025 | ✅ as supplied, dated | ⚠ missing |
| PRG-026 | ✅ as supplied, dated | ⚠ missing |
All fourteen lack an acceptance test; one has partial evidence. Evidence quality in this load is markedly better than the 2026-08-13 load — thirteen of fourteen rows carry a named source and a date, against six of sixteen last time. The acceptance-test gap is unchanged, and for the same reason: these are enumeration findings, and what "closed" means for each is a decision that has not been taken.
Completed¶
| Initiative | Completed | Notes |
|---|---|---|
| Forms portal LIVE in production | 2026-07-23 | Go-live submission 28bd1ebf; routing worker f9a114b on MAPLE, override OFF; all 6 ingest addresses verified delivering; HappyFox tickets opening. |
| Broken-access-control remediation | 2026-07-23 | Phase-1 decrypt/list/submit removed; require_role fail-closed; resolve_user least-privilege; IDOR ownership checks on Phase-2. No evidence of exploitation (empty audit_log + empty users table, independently corroborated). Finding record → T1.8. |
| Forms 2.5(c) routing worker — Gates 3/4/5 | 2026-07-23 | finalize_submission wired; MAPLE-resident Pub/Sub worker via localhost:25 Postfix; migrations 011 + 012 live. Residual propagation verification → Portal-propagation workstream. |
| Forms 2.5(d) — subject-template audit-persist + GAP-1 renderer fix + template reconciliation | 2026-07-23 | Audit-persist gate resolved; <%= Grant Funded => GAP-1 renderer fix landed; template reconciliation done. |
| Forms Portal Phase 2.5(b) — attachment upload wire-up + cleanup | 2026-05-08 | β closed with verification gap acknowledged (T1.7). Commits e17dacb (handler) + 28964c0 (cleanup). Migration 002 documents submission_deleted enum addition. See architecture/forms-phase2.5b-cleanup-closeout.md. |
| Okta Production cutover | 2026-04-23 | Production tenant live; group-based assignment; Preview kept as dev fallback |
| forms.greenpeace.us DNS + TLS | 2026-04-21 | CNAME → ghs.googlehosted.com; managed cert issued |
| Forms Portal Phase 1 (backend) | 2026-04-20 | Cloud SQL PG15, CMEK, IAM auth, AES-256-GCM envelope, RLS 4 roles |
| Meraki P1 inventory | 2026-04 | Org 395909, 5 networks, 32 devices |
| Portal backup cron on SKY | 2026-03 | gpus-portal-backup.sh nightly 02:30 → GCS. ⚠ Cron DIED ~2026-04-17 — see the T3 SKY portal-backup item; restore + backfill needed. |
| Okta Preview SSO across 4 portals | 2026-03 | OIDC PKCE, shared gpus-okta-auth.js |
Change log¶
| Version | Date | Author | Change |
|---|---|---|---|
| v1.23 | 2026-09-01 | R. Chhetry / Claude | PRG-029 raised as the travel workstream closes — a freshness check over the eight reports/<type>/latest.pdf objects in GCS, so a report that did not run becomes detectable. It is GOV-033's third recommendation and, by that finding's own analysis, the only one of the three that would have caught the 2026-09-01 executive-monthly miss rather than making the next one louder. Raised as program work following the GOV-028 → PRG-028 precedent, and scoped deliberately not to depend on GOV-034 (cron MAILTO=root is inert on all seven hosts) — routing report alerting through a channel already known broken would reproduce the failure it exists to catch. Its where-does-it-run / where-does-it-alert question is the same one PRG-028 carries and the two should be answered together. |
| v1.21 | 2026-08-17 | R. Chhetry / Claude | Evidence-completion pass. Artifact paths appended to PRG-004, PRG-005, PRG-006, PRG-007, PRG-009, PRG-011, PRG-013, PRG-020, PRG-022, PRG-023 and HB-004 from the newly reachable enumeration corpus at ~/estate-enum-2026-08. Figures confirmed exactly: gpusa Cloud SQL 2; gpus-it buckets 18; unattached-disk ratios 7:3 and 15:9; 47 static IPs across three projects; 25 gpusa instances; 15 active DHCP leases. HB-004 gains evidence — phoebe-capture/ holds vhost configs and logs for six candidate consumers, though the logs are unanalysed and its acceptance test stays INCOMPLETE. Two figures did not reproduce and were left unchanged: 126 host reservations against ~111 stated (the "workstation" filter is not recorded), and PRG-022's 154 remains unsourced. No acceptance test, status, severity or owner changed. |
| v1.20 | 2026-08-17 | R. Chhetry / Claude | 2026-08-17 enumeration load. New PRG-013–PRG-026 section from priorities/backlog-2026-08-17.md v1.0 items B-25–B-38 (Stages 1, 1b, 1c, 2, 3, 3b). No pre-existing row edited; the 13 flagged Done rows from 2026-08-13 remain untouched. Owner R. Chhetry throughout except PRG-015, which names Rob MacMillan as co-decider. No severity or priority stated for any row — all read not assessed, none inferred. All 14 rows lack an acceptance test and are marked INCOMPLETE; PRG-022 is partial on evidence. Nothing is done, including PRG-019, which records a DECOMMISSION decision by R. Chhetry and not a verification. Headline context: 82 declared entries against 481 deduped live entities, gpusa zero against 174, 50.0% UNKNOWN at Stage 2, 24 of 25 gpusa VMs UNKNOWN→DERIVED at Stage 3 (gated by GOV-011 — DERIVED is not enforced). Companion updates: security/vuln/tracker.md v1.9 (VLN-023–VLN-036) and governance/gap-register.md v1.1 (GOV-010–GOV-019 plus a Method learnings section). |
| v1.19 | 2026-08-13 | R. Chhetry / Claude | New ## Program work register — schema v1 section added (PRG-001–PRG-012 program items, HB-001–HB-004 host-level blockers), mapped from provisional P-1–P-16. No pre-existing row in this document was edited, reworded, re-statused or removed; the new section carries its own status vocabulary and the rule that done means verified end-to-end via the real production path against live state. No row in the new section is done. 13 of 16 new rows lack an acceptance test and 6 lack evidence entirely — recorded as INCOMPLETE, not filled with guesses. Two corrections established first-hand: gpus-forms-clamav-worker is listed in inventory.yaml so P-3 is 3 unlisted services not 4; and inventory.yaml:149 carries the same water/ESXi-6.7 misattribution as VLN-004, which VLN-013's acceptance test does not cover. Host rows for gannet/emu/ostrich/catbird/phoebe do not exist in inventory.yaml, so the four HB- blockers name their host and are parked until PRG-002 creates them. Companion registers opened the same day: security/vuln/tracker.md v1.8 (VLN-013–VLN-022) and governance/gap-register.md v1.0 (GOV-001–GOV-009). Pre-existing Done rows flagged separately in priorities/flag-list-unverified-done-2026-08-13.md — list only, re-adjudication tracked as GOV-008. |
| v1.17 | 2026-07-24 | R. Chhetry / Claude | New T1 compliance item: credential-rotation control never closed out. The quarterly forms-portal-credential-rotation-quarterly control emits records (Q3 2026 evidence committed 61eefff) but is never back-filled — 3 ACTION-NEEDED carried since Q2 (KMS rotation timestamps, Cloud SQL backup verification, HappyFox rotation) + an MFA-enforcement flag; possibly two consecutive quarters with no live console verification. Same class as the contract-vs-code and DB↔repo drift findings — artifacts asserting a state nobody verified. Pre-announcement gate: confirm Okta MFA enforcement before the staff notice — forms authz now rests entirely on Okta identity and the announcement leads with "sign in with Okta". |
| v1.16 | 2026-07-23 | R. Chhetry / Claude | FORMS PORTAL LIVE IN PRODUCTION (go-live 2026-07-23, submission 28bd1ebf). Routing worker f9a114b on MAPLE, override OFF, all 6 ingest addresses verified delivering, HappyFox tickets opening. Phases 2.5(a)–(d) DONE; 2.5(c) Gates 3/4/5, GAP-1 renderer fix, and template reconciliation closed. Broken-access-control remediation COMPLETE — Phase-1 decrypt/list/submit removed, require_role fail-closed, resolve_user least-privilege, IDOR ownership checks on Phase-2; no evidence of exploitation (empty audit_log + empty users table, independently corroborated). Restructured T1 from "finish forms" to cutover + post-go-live hardening. Legacy forms.us.gl3 retires 2026-08-01 (a Saturday — flagged for reconsideration); parallel running until then; LDAP deprovisioning gap argues for cutting over sooner. New/organized T1 workstreams: Cutover (staff notice, decommission, in-flight handling, redirect-or-dark, duplicate-ticket overlap, LDAP gap); Detection pipeline (audit_log→SOC + Cloud Run→Wazuh sinks don't exist; rules 100026–100029 inert by starvation; routing-worker empty SyslogIdentifier) — blocks the drill program; Security hardening (retire auth.py authz; narrow backend-SA over-grants; Cloud Armor = LB+NEG arch change; /metrics + /health/deep exposure; CORS *; per-instance rate limiting; MAC validation); Content/data (<% = Note => in Finance Termination rendering literally; live-DB legacy-tag sweep; DB↔repo drift root cause — HIGH; Insurance checkbox unanswerable-as-No; HR Termination TODO; footer sweep; DSAR/erasure + retention-lock gap); Okta dedicated forms app (requested from Conan 2026-07-23); 2.5(e) HappyFox API (queue names from API, ticket-audit migration, per-queue double-ticket deconfliction, API preferred — DMARC spoofing case); Portal propagation (live/derived, verify pages render). New T1.8 documentation set (threat model, ASVS L2, PCI-DSS out-of-scope determination, OWASP→NIST→MITRE mapping, IR runbook + DRP RTO/RPO, BAC + governance finding records, 2.5(d) as-built, field-exposure regen) — depends on the assessment. T1.9 exercise program (SQLi tabletop 60m + blue-team 90m + red-team 90m) — BLOCKED on detection. Phase 3 HappyFox marked superseded by 2.5(e); T5-EXPANDED un-gated (Gate 5 landed). New adjacent T3 item: SKY portal-backup cron died ~2026-04-17 (portals/ snapshots stopped; server backups current). |
| v1.15 | 2026-06-10 | R. Chhetry / Claude | Forms 2.5(c) implementation underway. Design doc → v0.3 (commits f0b79a8, 9a5281b): §3a wire contract with purged=410 DECIDED, §3b async finalize, §6a actions-schema-as-built, §17 coverage triad, 2.5(e) relabel. Migration 011 APPLIED LIVE (audit_action enum 26→30, commit 4bdfdf9); migration 012 (submission_finalized, 30→31) approved 2026-06-10, applying. Gate 3 in progress per G3.0 decisions (C1 add GET, C2 routing_result finalize marker, C3 full predicate at finalize, C4 α publishes unconditionally, C5 = 012). Gate 4 = routing worker; Gate 5 = propagation (NEW, blocks "2.5(c) done"). 7-day Pub/Sub retention clock starts at Gate 3 push — Gate 4 within the window. New T5-EXPANDED filed: portal static-debt retirement (P0–P4) per 2026-06-10 read-only audit of both portals; P0 truth-fixes may interleave before Gate 4. Standing rule recorded: all portal content LIVE or DERIVED, never hardcoded; coverage standard to require render-from-source. Operational learnings logged: L2TP stale-tunnel 3rd occurrence, FortiClient utun6 confound, IAP 4003 host-level (break-glass promotion), legacy in_formfeed MySQL on public IP (decom precondition filed). |
| v1.14 | 2026-06-04 | R. Chhetry / Claude | Forms 2.5(c) routing pipeline design COMMITTED (v0.2, commit 40f90d5; live at architecture/forms-phase2.5c-design/). 2.5(c) moved from next-up to design-committed / implementation-pending. Locked: transport B-iii (MAPLE-resident Pub/Sub pull worker, localhost:25 Postfix, zero new secret, carries its own systemd + monitoring + IR/DR/drill); reuse forms_app role (RLS USING(TRUE) → GATE-4 state-coverage N/A); signed-URL attachment delivery (inline MIME deferred); interim minimal-plaintext body NOT gated on 2.5(d); lean audit — migration 011_routing_audit_actions.sql adds 4 enum values (submission_routed, submission_route_failed, email_sent, email_failed). Next code: 011 migration + finalize_submission wire-up (pure stub today); one code-time confirm — _status_to_wire at routes_phase2.py:57. |
| v1.13 | 2026-05-21 | R. Chhetry / Claude | α ClamAV close-out housekeeping. Sigrefresh build trigger now filters by gpus-forms-clamav-worker/** (A0) — fixes the double-build-on-every-push issue. Backfill drain complete (A1) — 3 fixture attachments (PNG/DOCX/XLSX, all rchhetry β-phase test files from 2026-05-08) re-fired via synthetic Pub/Sub publish, all clean, GCS bytes unchanged, 3 new audit_scanned_clean rows; attachments table now clean=5, pending=0. Design doc alpha-clamav-worker.md → v1.3 (§6 sweep reframed required, §10 DLQ subscription clarified, new §13a 7-point Cloud SQL access spec for new DB-using services). Commit 3 hardening row added to summary. Four new T3 candidates filed: VPN cold-start packet loss, IAP-to-MAPLE 4003 backend-fail, Cloud Scheduler missed-tick on cold-start, and "GCP terraform tree has no VCS" (the no-VCS finding supersedes the previously-filed "WDC VPN/route Terraform state drift" — drift framing was wrong; no remote ever existed to drift from). A2 deferred per the no-VCS surprise; tonight's mitigation is a local snapshot of the 11 .tf files. |
| v1.12 | 2026-05-20 | R. Chhetry / Claude | α ClamAV Commit 2 COMPLETE. Slack alerts + GCS quarantine tag + /dlq-alert + /sweep-stuck + Cloud Scheduler wiring shipped. EICAR test verified full infected pipeline end-to-end to #us-soc-alerts. Worker rev 00010-gr7. Migrations 009 + 010 (audit_action enum + scanner UPDATE/INSERT WITH CHECK widen). Terraform 6 resources applied (DLQ sub + scheduler SA + 2 scheduler jobs). 8 hardening/hygiene items filed for Commit 3. |
| v1.11 | 2026-05-19 | R. Chhetry / Claude | α ClamAV Commit 1 COMPLETE. Migrations 004-008 shipped (enum + IAM user + grants + RLS + sequence USAGE). Pipeline verified end-to-end on T1.7 fixture (07c3c9e0) and real upload (4a0f5bd5). Worker gpus-forms-clamav-worker live (rev 00002-dwx, 2Gi). α.1 follow-ups: stuck-scanning sweep, DLQ subscription + alert, backfill drain of 3 remaining pending rows. |
| v1.10 | 2026-05-14 | R. Chhetry | T1.7 forms portal frontend silent-attachment-drop bug closed. |
| v1.9 | 2026-05-08 | R. Chhetry | β phase: Phase 2.5(b) attachment upload + 2.5(b.cleanup) MIME/size truth consolidation COMPLETE. Three-layer verification PASS on 999cf0cc-…; 4-source-of-truth divergence collapsed to Config.ATTACHMENT_*; production env vars removed; schema migration 002 documents submission_deleted audit_action enum addition. Two orphan-intent submissions cleaned up (d20d2ac8, e93efbc3). New T1.7 filed: frontend silent-attachment-drop bug surfaced during β verification — SPA submit must be gated on attachment validation state. Verification gap on rev 00043-d9q acknowledged in closeout doc. |
| v1.8 | 2026-05-08 | R. Chhetry | New T1.6 workstream filed: forms portal SOC/observability integration. Gap surfaced during Phase 2.5(b) scoping — forms portal has been operationally invisible to SOC since cutover. 8 sub-items spanning logging/metrics/Wazuh/SOC tab/runbook/DRP/drill/ASVS. Sequenced after Phase 2.5 functional work, before T4 SOC Tickets. |
| v1.7 | 2026-05-07 | R. Chhetry | γ phase: design doc correction shipped (2de57c9) — forms-phase2.5a-design.md now matches shipped code. α phase: pulldown regression for yes/no booleans CLOSED (f31ec5f) — Flask route converter fix. T1.5 sub-section now fully closed (all 4 items a/b/c/d). New cross-cutting lesson on hypothesis falsification via empirical data added. |
| v1.6 | 2026-04-30 | R. Chhetry | Phase 2.5(a) COMPLETE — 3 commits (07cb75c, 10bcd0d, 83f85cb), submissions now persist end-to-end (verified API + DB + audit). γ phase shipped same day: font darkness (f5c79e0+e58dabc), NO_DATA_TYPES filter (6d9231a). T1.5 items b/c/d CLOSED; item a (pulldown regression) remains. Cross-cutting lessons added: design doc accuracy, MAPLE access pattern. |
| v1.5 | 2026-04-29 | R. Chhetry | Phase 2.5(a) Commit 1 SHIPPED (07cb75c, auth_v2 User.username via preferred_username). Commits 2+3 paused at design-complete state. New T1.5 sub-section "Forms Phase 2 UX gaps" added covering pulldown regression (yes/no booleans), COST CENTER duplicate, NOTESDIVIDER label leak, and font contrast. |
| v1.4 | 2026-04-28 | R. Chhetry | FieldRenderer pulldown bug FIXED (commit 5512d8e). Phase 2.5 "Phase 2 backend wire-up" promoted as T1 sub-section after discovery that Phase 2 submission stubs persist nothing. Cross-cutting lesson on Cloud SQL access added. |
| v1.3 | 2026-04-27 | R. Chhetry | Forms Phase 2 frontend cutover complete. Phase 2.1 cleanup queue (5 items) added under T1. FieldRenderer pulldown bug added. Cross-cutting Cloud Run env-var lesson documented. |
| v1.1 | 2026-04-24 | R. Chhetry | Re-sequenced: forms (T1) → Meraki (T2) → WDC foundation (T3). WDC items no longer compete with forms momentum. |
| v1.0 | 2026-04-24 | R. Chhetry | Initial draft — consolidated priorities from memory + recent session notes |