Skip to content

Vulnerability Tracker

Classification: CONFIDENTIAL — Internal Use Only Document: security/vuln/tracker.md · v1.34 · 2026-09-01 · GPUS-IT


Summary

Severity Open In Progress Remediated
Critical (CVSS ≥ 9.0) 1 0 0
High (CVSS 7.0–8.9) 2 1 1
Medium (CVSS 4.0–6.9) 2 2 0
Low (CVSS < 4.0) 2 0 1

What this table counts — and a recount, 2026-09-08

Scope: VLN-001VLN-012 only — the two tables immediately below. It does not cover the Security findings register (VLN-013 onward), which holds the majority of the estate's open findings. A reader taking these twelve as the total will understate the backlog by a wide margin.

Recounted 2026-09-08 against the rows, and it had drifted. The previous version totalled 14 across tables holding 12 rows: High read 2 / 2 / 1 and Medium 3 / 2 / 0, neither of which matched. The counts above are derived from the rows rather than maintained by hand.

VLN-010 and VLN-011 carry no CVSS by design (footnotes 1 and 2) and are counted under their argued severity, High. A CVSS-banded table cannot place a row that deliberately has no score; that is a limitation of the banding, not a missing score.

Two register-hygiene defects — raised 2026-08-10, DECIDED and applied 2026-08-13

Read this first if you followed a reference to VLN-009 and did not find what you expected. Both defects below were recorded on 2026-08-10 and left open for an owner decision. That decision was taken on 2026-08-13 and is applied in this version (v1.7). This note is retained, not deleted, so that references written between 2026-03-16 and 2026-08-13 still resolve.

Defect 1 — VLN-009 was used twice. It was carried by both the open Critical passwordless-sudo finding and the remediated chronyd/NTP entry.

DECIDED 2026-08-13: VLN-009 stays with the passwordless-sudo finding (VLN-009) — its row, page, nav entry and all cross-references are unchanged. The chronyd/NTP entry is reassigned to VLN-012 and now appears under that id in Remediated vulnerabilities below.

If you arrived here from a pre-2026-08-13 reference that meant "VLN-009 = chronyd / NTP / timestamp skew", the row you want is now VLN-012. Nothing about that finding changed — same CVSS 2.6, same remediation date 2026-03-10, same conclusion. Only the identifier moved.

Defect 2 — VLN-011 had a published finding page but no row here. It was registered in the nav and under security/ but never entered in this table; adding it was out of scope for the VLN-010 publication pass.

DECIDED 2026-08-13: the missing row is added to Open vulnerabilities below, pointing at the existing finding page. The page and nav entry were already correct and are unchanged.

No id other than the chronyd entry was renumbered. VLN-010 and VLN-011 are untouched, and no new id was minted for the passwordless-sudo finding.


Open vulnerabilities

ID Component Description CVSS Severity Affected Target Status
VLN-009 Sudo / SSH Passwordless root (NOPASSWD: ALL) for the backend SSH account on all four WDC hosts; key is mounted into three internet-facing Cloud Run services 9.9 Critical SKY/RAIN/SUN/WIND Q3 2026 Open
VLN-011 Forms / Routing worker Forms HappyFox Dispatch Never Ran — routing_worker.py:680-694 matched every happyfox_template action, wrote status='deferred', recorded success=true and continued; no HappyFox client existed anywhere in the repo (HAPPYFOX_URL declared at config.py:78, never read). 0 of 37 submissions had ever received a HappyFox ticket id; 19 sat at happyfox_status='deferred'. Two forms carried no email action, so they delivered nothing while the UI told submitters "The relevant team has been notified." Confirmed 2026-08-06 read-only against live Cloud SQL + repo. it-support-request CLOSED end to end — fix deployed and a real ticket confirmed opened (finding §11). new-employee-notification-contractor-intern still HELD; contract-extension-notification UNREVIEWED. Full detail: VLN-011 n/a² High Forms portal (Cloud Run + MAPLE routing worker) Q3 2026 In Progress — 1 of 2 affected forms verified closed; second form held
VLN-001 SSH Config No MFA — password auth disabled but no TOTP/U2F second factor 7.5 High ALL Q2 2026 Open
VLN-002 Network No VLAN segmentation between server roles 7.2 High ALL Q3 2026 Open
VLN-003 GCP IAM Service account key stored on-disk — no Workload Identity Federation 6.5 Medium GCP Q2 2026 In progress
VLN-004 ESXi 6.7 EOL hypervisor — no vendor security patches since Oct 2023. ⚠ SCOPE DISPUTED 2026-08-13 — see VLN-013: authenticated vSphere API shows the 6.7 builds are on fire (17700523, 6.7 U3) and flower (8169922, 6.7 GA 2018, never patched); water is on 8.0.3 (24022510) and is not EOL. The Affected cell below is believed wrong. It is left unchanged pending VLN-013's acceptance test rather than corrected here 6.3 Medium WATER hypervisor (disputed — see VLN-013) Q4 2026 Open
VLN-005 Identity No SSO — separate credentials per service, no central identity 5.4 Medium ALL Q3 2026 Open
VLN-006 DNS Recursive resolver accepts queries from all internal hosts 4.9 Medium SKY/RAIN Q2 2026 In progress
VLN-007 Logging No centralized SIEM alerting — manual log review only 3.8 Low WIND Q3 2026 Open
VLN-008 Backup No automated backup integrity verification or restore testing 3.1 Low ALL Q2 2026 Open

¹ VLN-010 carries no CVSS deliberately. CVSS scores a weakness that an attacker exploits against a system; this is a detection control that does not run. There is no attack vector, no privilege requirement and no impact metric that describes it honestly, and any score produced would be false precision in a register people make prioritisation decisions from. Severity is assessed as High on operational grounds: the control is the primary alert-fatigue remediation for the estate, it is documented as delivered, and 99.0% of level-12 volume is the noise it was written to remove. Full detail: VLN-010.

² VLN-011 carries no CVSS. The source finding (security/finding-2026-08-06-forms-happyfox-never-dispatched.md, v1.1) assesses severity as High — silent loss of user-submitted requests and states no CVSS vector. None is invented here. As with VLN-010, this is a delivery control that did not run rather than a weakness an attacker exploits, and a score produced after the fact would be false precision in a register used for prioritisation. Severity High is carried through from the finding page as written.


Remediated vulnerabilities

ID Component Description CVSS Remediated Notes
VLN-012 NTP chronyd not verified — potential log timestamp skew 2.6 2026-03-10 chronyd verified and syncing on all 4 servers. Formerly VLN-009; reassigned 2026-08-13 to resolve a duplicate ID. The original VLN-009 designation is retained by the passwordless-sudo finding.
VLN-010 Wazuh / Detection Baseline rules 100016/100017/100018 were loaded, valid and counted but never evaluated — Wazuh sorts siblings by level and stops at first match, so the level-3 baselines sat behind the level-12 catch-all 100015 and could not be reached. The tuning written to suppress ~184 daemon alerts/day never took effect, and every observable was green: the rules loaded, the count was correct, validation passed. Six hypotheses were eliminated, all of which assumed rejection at load; none was n/a¹ 2026-09-08 Fixed by re-parenting the three baselines as children of 100015 via <if_sid>, verified before deploy on 4.14.4-rc2 and deployed 2026-08-11 18:29:50 UTC. All of D1–D6 measured and passed. D6(a) mail down ~95% at the inbox, counted on the real recipient to=<rajesh.chhetry@greenpeace.us>: 351 across the fix boundary, then 19 / 18 / 25 / 7. Two earlier needles (ossec, wazuh) both returned 0 and both were meaningless — postfix logs addresses, not the sending application. Report as a RATE: email_maxperhour=12 capped Wazuh at 288/day against ~184/day pre-fix, so 351 is a throttled sample, not an alert count — "351 became 19" is wrong. D6(b) indexed volume flat at 182–184/day vs 183–184/day pre-fix — flat IS the pass, since log_alert_level=3 keeps every suppressed event indexed; the fix removes mail, not documents. D6(c) 100016/100017/100018 non-zero and exclusively level 3. Compound failure (silence mistaken for suppression) excluded by measurement, not assumption. BT-003 proven in production: 100015 fired 15 times since 2026-08-12, every one at level 12 — the daemon noise went without touching the detection. Reproduction scripts retained on MAPLE. See the finding


Security findings register — schema v1 (opened 2026-08-13)

This section uses the schema below; the tables above do not. The legacy tables (VLN-001VLN-012) keep their original columns and are unchanged. New findings from 2026-08-13 onward are recorded here against the full schema: id · title · register · status · severity · owner · evidence · acceptance test · blocks · blocked_by · date_raised · date_verified.

Status vocabulary: open · in progress · blocked · done. done means verified end-to-end via the real production path against live state. Authored, present, loaded or validated is not done.

One row in this section is done: VLN-041, closed 2026-08-25. It is the first, and it is closed against a live observation on the deployed revision — a real Okta token returning 200 on the endpoint that runs the changed code, correlated to that revision's digest in the Cloud Run request log. Not "authored", not "merged", not "the build was green".

VLN-014 is why that bar is written the way it is. It was recorded done on 2026-08-21 against a verification performed 2026-08-17; a re-query on 2026-08-24 found the condition no longer held, and the row is back to in progress. See the regression note on that row — it is a worked example of why this rule says verified end-to-end against live state and not verified once. VLN-041 carries the same exposure: it is closed against today's revision, and a future revision that reintroduces the defect would reopen it. What makes that unlikely rather than merely hoped-for is condition 2 — the guard is executed by a gate that has demonstrably failed a build over exactly this suite.

evidence records what was read, where, and on what date. acceptance test states in advance the observable condition that closes the row. A row missing either is marked INCOMPLETE and is not filled with a plausible guess. Fields recorded as not assessed were not supplied and were not inferred.

Index

ID Title Status Severity Blocked by
VLN-013 ESXi 6.7 EOL is on fire and flower — VLN-004 misattributes it to water open not assessed
VLN-014 sky and rain serve different DNSSEC signings — converged 2026-08-17, diverged again 2026-08-19 in progress High (as supplied)
VLN-015 NOPASSWD sudo on emu and ostrich — VLN-009 scope extends beyond WDC open not assessed
VLN-016 One person holds three identity strings with inverted access across systems open not assessed
VLN-017 gcloud list verbs exit 0 on permission denial open not assessed
VLN-018 All 10 Cloud Run services in gpus-infra are ingress: all open not assessed
VLN-019 duck.cloud.us.gl3 resolves to an address that is not duck's open not assessed
VLN-020 Five Puppet node definitions can never match a certname open not assessed
VLN-021 3 of 4 Meraki org admins hold orgAccess: full on the VPN-terminating org open not assessed
VLN-022 Published forms go-live record asserts a delivery claim VLN-011 disproves open not assessed GOV-009 (correction cannot be verified live)
VLN-023 vmstorage running a degraded RAID 6 with no spare — NFS datastore for every WDC VM in progress CRITICAL
VLN-024 Plaintext vendor credentials in the Puppet control repo (pushtab) open CRITICAL
VLN-025 puppetCrypt exists and pushtab bypasses it open HIGH
VLN-026 vmstorage has no backup or replication task open HIGH GOV-019 (derived)
VLN-027 Meraki IPsec pre-shared keys returned in plaintext open HIGH
VLN-028 SMBv1 enabled with signing off on all three Synology units open HIGH
VLN-029 Key material in a control repo scheduled for deletion open HIGH
VLN-030 Undocumented second VPN peer (Target, IKEv1) open MEDIUM
VLN-031 Zero 2FA across all three Synology units open MEDIUM
VLN-032 Meraki org 395909 has no SSO and no IdP open MEDIUM
VLN-033 synstorage volumes at 93.1% and 82.3%, status attention open MEDIUM
VLN-034 rooster cannot compile in production — archive::backup body commented out open MEDIUM
VLN-035 dataflow.us.gl3 resolves to a host that does not exist open MEDIUM
VLN-036 wdc-wap-5 status alerting open LOW
VLN-037 Four production security reports ran on EOL Python 3.6.8 for ~4 months remediated HIGH
VLN-038 No Cloud Build in the estate ran unit tests — controls implemented as tests were never enforced in progress High (argued)
VLN-039 Pre-build check steps couple every deploy to PyPI availability — a gate can block an incident fix in progress Medium (argued)
VLN-040 The dependency scan runs, finds 24 advisories, and its exit code is discarded — green on every build since April in progress High
VLN-041 auth_v2 took the JWT algorithm allowlist from the unverified header of the token being validated done 2026-08-25 High (argued)
VLN-042 CORS(app, origins="*") on an unauthenticated, ingress: all backend — any site a staff member visits can read /metrics and /health/deep from their browser done 2026-08-25 MEDIUM (argued)
VLN-043 The forms SPA reads none of the three variables its own env template defines, and the template's Okta clientId is the pre-cutover one open LOW (argued)
VLN-044 The Submitted page promised every submitter an approval step to watch and an outcome email — true for 1 form of 29. Gate written against a field that never existed in progress — fixed, not yet deployed MEDIUM (argued)
VLN-045 /status/<id> is headed "Travel request." for every form — a Facilities Support Request's own status page names a workflow it is not in in progress — fixed 2026-08-26, not yet verified in a browser LOW (argued)
VLN-046 The confirmation page asserted "Ticket — NOT CREATED" while a ticket existed — CONFIRMED by #USFIN00367708. Copy fixed 2026-08-26 in progress — fixed, not yet verified in a browser LOW (argued)
VLN-047 Migration 022's revise path is deployed and no UI constructs its URL — and /status/<id> asserts the request "cannot be edited" while offering a button that produces an unlinked resubmission done 2026-08-26 — acceptance test met; first revision 35750caf links to its parent MEDIUM (argued)
VLN-048 The portal-coverage gate passes every cloud_services entity automatically — present = (sentinel in blob) or (eid.lower() in blob) and both portals carry the sentinel in a comment. Two live Cloud Run services sat unrendered for three weeks, gate green open MEDIUM (argued)
VLN-049 The travel form's served instructions told every submitter a returned request "cannot be edited" — the VLN-047 claim, live at version: 2, on the same column GOV-021 was raised about. A green test pinned it done 2026-08-27 — 4 surfaces fixed in the v3 bump; served column verified against gpus_forms MEDIUM (argued)
VLN-050 The HappyFox ticket body (travelrequest001) is a third, uncounted render surface that enumerates its fields — the v3 bump added three and never touched it. One real ticket (c84bfb45, Madison Carter) was rendered by the broken template 6m06s before the fix landed CLOSED 2026-09-01 — verified on a real ticket, #USITS00368946 MEDIUM (argued, raised from LOW)
VLN-051 approval_decision_recorded / approval_notification_sent are in AUDIT_ACTIONS but in neither ship_actions nor excluded_actions — 31 of 33 declared. The non-repudiation audit row reaches no consumer open MEDIUM (argued)
VLN-052 approval_access's denial reasons are absent from authz_reasons.json, so a denied approval attempt falls to rule 100030 at level 3 — below the ≥ 10 auto-ticket threshold. A forged-approval attempt pages nobody open MEDIUM (argued)
VLN-053 CEDAR's log indexer serves the entire security-alert corpus with no authentication — Elasticsearch 8.19.13 on 172.16.0.13:9200, xpack.security.enabled: false, plain HTTP, no TLS. Anonymous _cat/_count/_search over 155 daily wazuh-alerts-* indices back to 2026-04-07 (339 indices total). Reachability CORRECTED 2026-09-08 (v1.1) — v1.0 said "anything that can reach the subnet" and counted 113 WDC workstations, reasoning from the GCP rule cedar-ingress alone. CEDAR's host firewall is narrower and is the effective control: 9200 is permitted from 172.16.0.12 (MAPLE) and 10.8.0.0/28 only; the WDC LAN gets 5140/tcp and ICMP, not 9200. The 10.8.0.0/28 end is the Serverless VPC connector, enumerated rather than assumed: 5 of 10 Cloud Run services attach to it (gpus-forms-backend, gpus-forms-clamav-worker, gpus-security-backend, gpus-soc-backend, gpus-status-backend), all ingress: all — so the distance from the internet to the corpus is one SSRF/RCE in any of five services, with no credential to steal. The cloud layer is still open to the 113 workstations and only firewalld withholds them, which is a latent trap if the host firewall is ever stopped. No audit logging is possible with security off, so the access is unattributable and the exposure cannot be scoped retrospectively. Write/delete are almost certainly equally unauthenticated — reasoned from the absent auth layer, deliberately not tested open High (argued)
VLN-054 Webmin on CEDAR and MAPLE is reachable only from the Cloud Run VPC connectorminiserv.pl runs as root, active and enabled at boot on both hosts, bound 0.0.0.0:10000 + [::]:10000. The host firewall on each permits 10000/tcp from 10.8.0.0/28 and nothing else — the Serverless VPC connector, where no administrator sits and 5 of 10 ingress: all Cloud Run services egress. The management network 192.168.124.0/24 gets 22/tcp (and Grafana on MAPLE) but not 10000, so no operator can reach Webmin over the network at all; real use goes via ssh -L to loopback, which needs no rule. The exposure therefore buys nothing and removing it costs nothing. Auth is real but single-layered: TLS enforced (ssl_enforce=2, HSTS, no SSLv2/3), session-based, passdelay=1, blockhost_failures=5 — but blockhost_time=60 is sixty seconds, there is no 2FA, and no allow= IP allow-list, so firewalld is again the only control. Also contradicts defense-in-depth.md, which claims all admin interfaces are management-network-only: Prometheus, Kibana and Webmin are not reachable from that network, and all four are reachable from a connector the doc never mentions. CEDAR and MAPLE are absent from access-control/roles.md entirely. Exactly one account exists, it is root, and its password field is x — Unix/PAM auth against /etc/shadow, confirmed. So the Webmin login is the host's root password: no lower-privilege account to compromise first, no second factor, and blockhost_failures=5/blockhost_time=60 permits ~5 guesses/minute indefinitely, offered to the origin five ingress: all Cloud Run services egress from. Established by printing field lengths and then a single character — no credential material left the host. …and on CEDAR that login cannot succeed. passwd -S root returns LK — the root password is locked in /etc/shadow, so pam_unix refuses every attempt. The sole account cannot authenticate: nobody can log into Webmin on CEDAR. Severity DOWNGRADED High → Medium 2026-09-08 on that check, after three revisions had published High. The chain x → "the login is the root password" was correct at every step and still wrong, because it stopped one layer short of the credential store — the same read-one-layer-and-stop shape recorded against this author in VLN-053 §4. What remains is a root-owned miniserv.pl, bound 0.0.0.0 + [::], enabled at boot, needlessly reachable from the connector, whose only live attack surface is pre-authentication — and a locked account is no defence there at all. Pre-auth question RESOLVED 2026-09-08 against the vendor advisory list: none published for 2.621. Every later fix (2.640, 2.641, 2.652, 2.653 — incl. CVE-2026-42210/56022 2FA bypass, CVE-2026-22678 XSS, an SSRF and a privesc) requires an authenticated session, so none is reachable while root cannot log in. The Critical trigger did not fire. What keeps it at Medium rather than Low: the mitigation is a state of /etc/shadow, not a control in Webmin. Setting a root password for any unrelated reason — recovery, vendor instruction, console rescue — silently re-opens a root login on a connector-facing port with no config change, no firewall change and no alert, and makes that entire authenticated-CVE backlog live in the same moment. "Safe because root happens to be locked" is not "safe because Webmin cannot authenticate that way". 2.621 is ≥4 security releases behind — unreachable here, live on any host with a usable login. MAPLE measured, not assumed, and it matches: passwd -S rootLK and cut -d: -f1,2 miniserv.usersroot:x. All three facts hold on both hosts, so nobody can log into Webmin on either. Measured rather than inferred because a locked root only yields an unusable login if root is the sole account and delegates to Unix auth — a second account, or root with a real stored hash, would authenticate inside Webmin and never consult /etc/shadow. Parity on one input is not parity on the conjunction. Both hosts rate Medium. If §7.2 ever comes true and a root password is set, MAPLE's consequence is the worse of the two: it is the Wazuh manager, so a root compromise there also compromises the pipeline meant to notice a compromise. Remediation: disable the service (systemctl disable --now webmin), on three independent lines of evidence — nobody can log in (locked root, sole account), nothing in the repo consumes port 10000 while the other connector rules do have consumers, and miniserv.log records no logins at all (its only three entries are the author's own 401 probes, which also resolved a miniserv.conf mtime that had been flagged as possible third-party activity). Disabling also converts §7.2's state into a real control: with the service gone, setting a root password later cannot silently re-open anything. Check MAPLE's own log first rather than reusing CEDAR's open Medium (both hosts; was High)

VLN-013 — ESXi 6.7 EOL is on fire and flower; VLN-004 misattributes it to water

Field Value
id VLN-013
title ESXi 6.7 EOL on fire and flower — VLN-004 misattributes this to water
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence Authenticated vSphere API, 2026-08-13: fire build 17700523 (6.7 U3); flower build 8169922 (6.7 GA, 2018, never patched); water build 24022510 (8.0.3). fire hosts sky, rain, sun and wind. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-esxi-inventory-20260813.txt. All three build numbers confirmed present in the capture: 17700523, 8169922, 24022510.
acceptance test WIDENED 2026-08-17 per GOV-017 (B-22). No artifact asserts that water is ESXi 6.7 or that water hosts VMs — covering (a) VLN-004's row above, (b) inventory.yaml:146-152, and (c) any downstream generated output derived from either. (Original test, 2026-08-13: "VLN-004 corrected to name fire and flower; water removed from its scope." Superseded — it closed on one artifact and left the other two asserting the same wrong thing.)
blocks
blocked_by
date_raised 2026-08-13
date_verified

Note: flower is on the 2018 GA build and has never been patched. VLN-004 above still reads WATER hypervisor; a scope-dispute marker has been added to that row pointing here. VLN-004's text is otherwise unchanged pending this row's closure.

RESOLVED 2026-08-17 — acceptance test widened

Raised 2026-08-13: the same misattribution was found in a second place while evidencing PRG-010inventory.yaml:149 records hypervisor: VMware ESXi 6.7 for water, and :152 records vms: [ocean] on a host that hosts zero VMs. The acceptance test then named only VLN-004, so correcting VLN-004 would have closed this row while inventory.yaml — the artifact three governance documents call the single source of truth — kept asserting the same wrong thing. It was flagged rather than widened unilaterally, and left as open item 2 in the 2026-08-13 session log.

Decided 2026-08-17: widened. The test now reads no artifact asserts water is ESXi 6.7 or hosts VMs, covering VLN-004, inventory.yaml:146-152 and downstream generated output. This closes open item 2 from the 2026-08-13 session log.

Justification is GOV-017, which found the disagreement is broader than water: flower is declared as ESXi 6.8, a release that does not exist, and fire is declared as hosting nothing while carrying all four core WDC servers. Those are separate rows and are not folded into this one.

VLN-014 — sky and rain serve different DNSSEC signings — REGRESSED 2026-08-19, REOPENED 2026-08-24

Field Value
id VLN-014
title rain served a 2026-07-29 DNSSEC signing while sky served 2026-08-12; converged 2026-08-17, then diverged again on 2026-08-19 by the same mechanism
register Security findings
status in progress — remediated and verified 2026-08-17, regressed 2026-08-19, reopened 2026-08-24
severity High — supplied as "severity high, dated"
owner R. Chhetry
evidence Zone dumps from sky and rain, 2026-08-13, showing differing RRSIG inception and differing NSEC3 salt. SOA serial is 2026072902 on both — the re-sign never bumped it, so rain had no transfer trigger. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/zone-canon-sky-20260813.txt and raw/zone-canon-rain-20260813.txt. All four claims confirmed: SOA serial 2026072902 on both; RRSIG inception 20260812 on sky against 20260729 on rain, with rain carrying expiry 20260828; and the NSEC3 salts differ — sky A178C4BCAFEA2FA0, rain 83C78EB3B3503D4C.
acceptance test Serial bumped; transfer confirmed; RRSIG inception and NSEC3 salt identical on both servers, verified by direct query. Met in full 2026-08-17; NOT MET as of 2026-08-24 — see the regression note. Unchanged: it is the right test, and it is what caught this.
blocks
blocked_by
date_raised 2026-08-13
date_verified 2026-08-17 — superseded; re-queried 2026-08-24 and the condition no longer held

REGRESSED 2026-08-19 — reopened 2026-08-24 by direct query against both hosts

The acceptance test is not met. Queried live on 2026-08-24, dig against @localhost on each host:

Check sky rain
SOA serial 2026081702 2026081702
RRSIG inception 20260819063701 20260817170231
RRSIG expiration 20260918063701 20260916170231
NSEC3 salt CB18BB4A142868CF A178C4BCAFEA2FA0
Signing keytag 6660 6660

rain still holds exactly the values recorded as verified on 2026-08-17 — inception 20260817170231, salt A178C4BCAFEA2FA0. The 08-17 verification below was sound and is not withdrawn. sky has moved on without it.

This is the original fault, recurring — by the exact mechanism this row documents. On sky, /var/named/wdc.us.gl3.db (the source) is unchanged since Aug 17 14:01 and still carries serial 2026081701, while /var/named/wdc.us.gl3.db.signed was rewritten Aug 19 03:37 and carries 2026081702 — the same served serial rain already had. Someone re-signed on 08-19 without bumping the source serial first, -N INCREMENT derived the identical serial again, and rain — correctly — saw no reason to transfer.

The remediation note below ends with the sentence "Omitting the source bump reproduces the original fault silently." That is what happened, two days after the verification and two days before the row was recorded done.

Why this row was done for three days

The 2026-08-17 verification was real, thorough, and correctly evidenced — rain's live values still match it exactly. The row was then written up on 2026-08-21 from that verification, not from a fresh query, and the regression had landed on 08-19 in the gap between the two.

Nothing was fabricated and no evidence was wrong. The failure is that verified on the 17th was recorded as done on the 21st without re-reading live state at the moment of writing. Under the status rule as stated — verified end-to-end via the real production path against live state — the verification had gone stale and the row should not have closed.

Close this row only from a query run at the time of closing, and prefer a standing check to a point-in-time one: this zone has now diverged twice by the same mechanism, undetected both times.

The 2026-08-17 verification — sound, and NOT withdrawn

Every clause of the acceptance test was met and observed, not inferred.

Check sky rain
SOA serial 2026081702 2026081702
RRSIG inception 20260817170231 20260817170231
RRSIG expiration 20260916170231 20260916170231
NSEC3 salt A178C4BCAFEA2FA0 A178C4BCAFEA2FA0
Signing keytag 6660 6660 (unchanged)

rain flipped its salt from 83C78EB3B3503D4C to A178C4BCAFEA2FA0 — that is the proof it took the new signing rather than merely agreeing by coincidence, and it is exactly the discriminator the acceptance test was written around.

Transfer confirmed in rain's own log: zone wdc.us.gl3/IN: transferred serial 2026081702 / Transfer status: success. delv reports fully validated on both hosts. dsset md5 unchanged, so no parent DS update is needed. Five untouched zones on sky verified identical, bounding the blast radius. named PID unchanged — this was a reload, not a restart.

Root cause corrected — the recorded mechanism was the symptom

The row above states the cause as "the re-sign never bumped it, so rain had no transfer trigger." That is what was observed, not why it happened, and it is corrected here rather than left standing.

Actual mechanism: dnssec-signzone -N INCREMENT bumps the serial in the output file only, leaving the source zone untouched. Every re-sign from an unchanged source therefore reproduces the identical served serial — and a secondary that declines to transfer an unchanged serial is behaving correctly. rain was never faulty.

The 2026-07-29 and 2026-08-12 signings both derived 2026072902 from source serial 2026072901. Two signings were lost this way before anyone noticed, and nothing in the pipeline could have surfaced it: each re-sign reported success, the output file was valid, and the secondary's refusal to transfer was the correct response to an unchanged serial. Another instance of the estate's characteristic pattern — see M-01, and note this one is subtler: the control ran, did its job, and the failure was in the input.

Verified procedure — bump the source serial first, then:

cd /var/named && dnssec-signzone -3 A178C4BCAFEA2FA0 -H 10 \
  -N INCREMENT -o wdc.us.gl3 -K /var/named wdc.us.gl3.db

Omitting the source bump reproduces the original fault silently.


VLN-015 — NOPASSWD sudo on emu and ostrich; VLN-009 scope extends beyond WDC

Field Value
id VLN-015
title NOPASSWD sudo present on emu and ostrich under the gcloud-provisioned rchhetry identity — scope expansion of VLN-009 onto both DNS masters
register Security findings
status open
severity not assessed — no severity supplied. Not inherited from VLN-009 (CVSS 9.9): the hosts, identity and access path differ, and carrying the score across would be an assumption, not a measurement.
owner unassigned
evidence Enumerator 3 partial-access note, 2026-08-13.
acceptance test Full sudoers audit across the gpusa estate; VLN-009 scope updated to match what the audit finds.
blocks
blocked_by
date_raised 2026-08-13
date_verified

Both affected hosts are DNS masters. VLN-009's row above is scoped SKY/RAIN/SUN/WIND and is not edited by this row — its scope changes only once the sudoers audit in the acceptance test has run.


VLN-016 — One person holds three identity strings with inverted access across systems

Field Value
id VLN-016
title Multiple identity strings for one person across systems, with inverted access
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence access-map.md v2 (6 probes, 2026-08-13); Meraki GET /administered/identities/me. Specifics: rchhetry@greenpeace.org holds Owner on gpusa (~25 VMs including both DNS masters) while documented as SSO-only, and is denied on gpus-infra. rajesh.chhetry@greenpeace.us is the inverse. The Meraki org 395909 admin authenticates as a third string, rajesh.chhetry@greenpeace.org. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): access-map.md — the artifact this row cited, which was recorded on 2026-08-13 as not present in the portal repo. It is in the enumeration corpus, headed "Access map — Stage 1, Step 0 (v2, post-reauth)", dated 2026-08-13, with raw output at raw/access-probe-both-accounts-20260813-v2.txt. Its coverage matrix confirms the inversion: gpus-infra READABLE by rajesh.chhetry@greenpeace.us and DENIED for rchhetry@greenpeace.org; gpusa-it-infrastructure-306400 and gpus-it-infrastructure exactly the inverse. The third identity string is corroborated independently — rajesh.chhetry@greenpeace.org appears as a Meraki org admin in raw/task1-meraki-20260813.txt.
acceptance test Identity map recorded per system; a deliberate decision taken on whether to consolidate the identities or formally accept the split, and that decision encoded in the runbooks.
blocks
blocked_by
date_raised 2026-08-13
date_verified

The gap between "documented as SSO-only" and "holds Owner on ~25 VMs" is the material part: the documented access model and the effective one disagree.


VLN-017 — gcloud list verbs exit 0 on permission denial

Field Value
id VLN-017
title gcloud list verbs return exit 0 on permission denial; the error appears only on stderr
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence All 6 access probes, both directions, 2026-08-13. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/access-probe-both-accounts-20260813-v2.txt. Confirmed verbatim: every probe line reads EXIT_CODE=0, including probes 2, 3 and 4, whose bodies carry Required 'compute.instances.list' permission for 'projects/…'. A denial and a success are indistinguishable by exit code in this capture.
acceptance test A grep of gpus-infra-portals, the backup scripts and Cloud Build for any branch on $? following a gcloud list; each occurrence either fixed or documented.
blocks
blocked_by
date_raised 2026-08-13
date_verified

Same failure spine as VLN-010 and VLN-011: a control returns success while doing nothing, so every observable an operator would check comes back green.


VLN-018 — All 10 Cloud Run services in gpus-infra are ingress: all

Field Value
id VLN-018
title All 10 Cloud Run services in gpus-infra are ingress: all, including 4 backends and gpus-forms-clamav-worker
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence Stage 1 raw JSON, task4-cloudrun, 2026-08-13. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task4-cloudrun-20260813.txt — the file this row names. Date qualification: that file is dated 2026-08-13 and states in its own header "Source: raw/gpus-infra__run__rajesh.chhetry-at-greenpeace.us.json (Stage 1, generated 2026-08-12T20:28:28Z). NO new API calls." The observation is therefore 2026-08-12; 2026-08-13 is the derivation date. Service count 10 confirmed.
acceptance test Backends and the worker moved to internal-and-cloud-load-balancing, verified by an external request returning 403.
blocks
blocked_by
date_raised 2026-08-13
date_verified

VLN-019 — duck.cloud.us.gl3 resolves to an address that is not duck's

Field Value
id VLN-019
title duck.cloud.us.gl3 resolves to an address that is not duck's
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence Enumerator 3, 2026-08-13. Corroborated from a second independent source 2026-08-17 — see PRG-024: gpus-dist forward and reverse zones both say 10.1.96.40, the live VM is 10.1.96.46, and inventory.yaml says .46. Two independent sources now agree the zone data is wrong and inventory.yaml is right.
acceptance test Record corrected or removed; resolution verified by query.
blocks
blocked_by
corroborated_by PRG-024stated in the source row (B-36).
date_raised 2026-08-13
date_verified

VLN-020 — Five Puppet node definitions can never match a certname

Field Value
id VLN-020
title Five Puppet node definitions use bare names that cannot match an FQDN certname — that config has never applied to anything
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence Enumerator 2, 2026-08-13.
acceptance test Each of the five confirmed intentional, or removed.
blocks
blocked_by
date_raised 2026-08-13
date_verified

Config that has never applied to anything is the same class as VLN-010's inert rules: present, loaded, counted, and never reached. See also PRG-007 — 18 orphan Puppet node definitions, likely the same uncleaned decommission wave.


VLN-021 — 3 of 4 Meraki org admins hold orgAccess: full on the VPN-terminating org

Field Value
id VLN-021
title 3 of 4 Meraki org admins hold orgAccess: full on the organisation whose appliance terminates the WDC VPN
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence GET /organizations/395909/admins, 2026-08-13. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task1-meraki-20260813.txt. Confirmed, with the admins named: msedita@greenpeace.org, rajesh.chhetry@greenpeace.org and tanu.garg@greenpeace.org all carry orgAccess: full. All three show authMethod: Email and 2FA: True; only rajesh.chhetry@greenpeace.org has hasApiKey: True. The capture notes the querying key is itself org-full-access, authorised for GET-only use.
acceptance test Access levels reviewed and reduced to least privilege, or each full-access grant justified in writing.
blocks
blocked_by
date_raised 2026-08-13
date_verified

VLN-022 — Published forms go-live record asserts a delivery claim VLN-011 disproves

Field Value
id VLN-022
title Published forms go-live record asserts all 6 ingest addresses verified delivering and HappyFox tickets opening; the claim is false and is live on infra.greenpeace.us
register Security findings
status open
severity not assessed — no severity supplied; not inferred
owner unassigned
evidence Flag list, 2026-08-13 (priorities/flag-list-unverified-done-2026-08-13.md, Tier A2); VLN-011 finding page, which establishes 0 of 37 submissions received a ticket id and that no HappyFox client exists in the repo.
acceptance test Published record corrected; the claim replaced with an evidence-backed status.
blocks
blocked_by GOV-009 — the correction cannot be confirmed live until the portal render can be verified
date_raised 2026-08-13
date_verified

This is the only row here whose harm is currently published: the false claim is being served to readers of infra.greenpeace.us for as long as it stands.


VLN-023 — vmstorage running a degraded RAID 6 with no spare

Field Value
id VLN-023
title vmstorage (RS1619xs+, 192.168.120.51) running a degraded RAID 6 — the NFS datastore for every WDC VM
register Security findings
status in progress — remediation reported; not done, see below
severity CRITICAL (as stated)
owner R. Chhetry
evidence Synology DSM API, 2026-08-17. designedDiskCount 6, normalDevCount 5, disk_failure_number 1, spares [], repair_action degrade, data scrubbing unavailable (reason: abnormal). Three disks sit in not_usesdeb, sdec, sdfa — all reporting SMART normal. This is the NFS datastore for every WDC VM, mounted by fire, water and flower. Nothing in monitoring surfaced it and no artifact mentions it. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed: model=RS1619xs+, volume_1 … status=degrade, and the three unused disks sdeb, sdec, sdfa all reading status=not_use smart=normal. Note the capture places those three in pool RX1217rp-1 while the four normal disks sit in pool RS1619xs+ — an expansion-shelf detail not in the row text and not inferred into it.
acceptance test Array returns normal with a spare assigned and scrubbing available.
blocks
blocked_by
related_to VLN-026derived. No backup or replication compounds the degraded array.
date_raised 2026-08-17
date_verified

Spare reported assigned — remediation applied, NOT verified

R. Chhetry reports a spare assigned on vmstorage, recorded 2026-08-20. The action date itself was not stated and is not inferred here. The row is moved to in progress, not done, and the acceptance test is unchanged: array returns normal with a spare assigned and scrubbing available — three conditions, of which one is reported and none is re-observed.

Under the status rule in force, an action taken is not a state verified. That distinction is the whole reason this register exists: Tier A of the 2026-08-13 flag list is two rows that recorded an action and were read ever after as a verified outcome.

What closes it — re-run the same read-only probe that raised the finding and compare:

cd ~/estate-enum-2026-08
python3 syno_probe.py <creds-file> raw/task2-synology-<YYYYMMDD>.txt

Then confirm three things on vmstorage (192.168.120.51), all of which read the other way in the 2026-08-13 capture:

Field 2026-08-13 Required to close
volume_1 status degrade normal
spare spares: [] a spare assigned
data scrubbing unavailable (reason: abnormal) available

Note a rebuild takes time on a 26.79 TB volume — status=normal may not appear immediately, and a volume mid-rebuild is still a volume without redundancy. syno_probe.py is read-only by construction (MutatingCallBlocked guards every call), so re-running it is safe during a rebuild.

VLN-026 (no backup or replication task) and GOV-019 (snapshot config unknown, 403) are untouched by this and both still stand.

A RAID 6 down one disk with zero spares, on the datastore under every WDC VM, with three unused disks physically present and reporting healthy. The three not_use disks are why the missing spare is notable rather than merely unlucky.


VLN-024 — Plaintext vendor credentials in the Puppet control repo

Field Value
id VLN-024
title Plaintext vendor credentials for payment, payroll and CRM vendors in modules/relay/files/etc/pushtab
register Security findings
status open
severity CRITICAL (as stated)
owner R. Chhetry
evidence modules/relay/files/etc/pushtab lines 1–6. Six rows, format name user password endpoint — field 3 is a plaintext password on every row, covering payment, payroll and CRM vendors (litle, convio, adp, target). Last commit 2019-01-24. Deployed to /opt/gp/etc/pushtab on every storehouse node. The repo has 879 commits and five known working copies. Two values also transited terminal output on 2026-08-17. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage3b-credential-locations.json. Confirmed exactly: six entries, lines 1–6, all "type": "plaintext-password" at /etc/puppetlabs/code/environments/production/modules/relay/files/etc/pushtab, field 3, each "last_commit": "2019-01-24". The capture records value_length_bucket rather than the value itself — the enumeration did not copy the secrets into the artifact.
acceptance test All six rotated at the vendor, or the accounts disabled vendor-side, confirmed per vendor.
blocks
blocked_by
related_to VLN-025derived. puppetCrypt exists and pushtab bypasses it.
date_raised 2026-08-17
date_verified

DISPOSITION recorded 2026-08-20 — deferred to the Rocky migration. Read the residual risk.

Decision (R. Chhetry, recorded 2026-08-20): this closes when the servers move from RHEL to Rocky and the vendors are moved there under Terraform. Recorded as a dated decision. Status stays open — a decision is not a verification, and the acceptance test is unchanged.

The migration does not, by itself, close this row. The acceptance test is all six rotated at the vendor, or the accounts disabled vendor-side, confirmed per vendor — and that is deliberate. Standing up new Terraform-managed credentials on Rocky hosts creates a second working set; it does not invalidate the first. The six values in pushtab stay valid at litle, convio, adp and target until someone rotates or disables them at the vendor. The migration must therefore include an explicit vendor-side rotation or disablement step, per vendor, or this row survives it.

Exposure accepted by this deferral, stated so the decision is made on the full picture rather than re-litigated later:

  • The credentials have been in plaintext since 2019-01-24 — over seven years — across 879 commits and five known working copies.
  • They are deployed to /opt/gp/etc/pushtab on every storehouse node.
  • Two of the six values transited terminal output on 2026-08-17.
  • They cover payment, payroll and CRM vendors.
  • PRG-019 decommissions the code that reads the file. It does not decommission the credentials.

The exposure window is now the migration timeline. If that timeline is longer than a few weeks, the cheaper interim move is to disable vendor-side any of the six that are not currently required — which is checkable against PRG-019's finding that only 13 of 27 cron resources are potentially active, and that pushFacter, pushFPR and pushPSI transfer nothing at all.

Not closed by the Rocky migration, and not closed by deleting the file

Recorded as supplied: deleting the file does not invalidate the credentials. Five known working copies, 879 commits of history, and two values already in terminal output. The acceptance test is vendor-side rotation or disablement confirmed per vendor — nothing repo-side closes this row. Note PRG-019 disposes of the transfer stack as DECOMMISSION; that disposition does not close this row either.


VLN-025 — puppetCrypt exists and pushtab bypasses it

Field Value
id VLN-025
title puppetCrypt exists (125 GPG-encrypted secrets) and pushtab bypasses it — the capability was present and the file ignored it
register Security findings
status open
severity HIGH (as stated)
owner R. Chhetry
evidence Partial. Supplied: puppetCrypt exists and holds 125 GPG-encrypted secrets; pushtab does not use it. The count is specific, but no source artifact, path or read date was stated for this row. To complete: the puppetCrypt path and a dated listing. ⚠ Still unconfirmed after searching the enumeration corpus 2026-08-17. No artifact stating the 125-GPG-secret count was located under ~/estate-enum-2026-08; raw/stage3b-credential-locations.json records the pushtab plaintext rows only and says nothing about puppetCrypt. The count remains unsourced and this row stays partial.
acceptance test No plaintext credential in the control repo.
blocks
blocked_by
related_to VLN-024derived.
date_raised 2026-08-17
date_verified

The material point is that this was not a missing capability. Encryption was available in the same repo and the file did not use it, which makes this a process gap rather than a tooling gap.


VLN-026 — vmstorage has no backup or replication task

Field Value
id VLN-026
title vmstorage has no backup or replication task — configured-absent, not unchecked
register Security findings
status open
severity HIGH (as stated)
owner R. Chhetry
evidence Synology DSM API, 2026-08-17. SYNO.Backup.Task returns 0; SYNO.Backup.Repository repo_list is empty; CloudSync API not registered. Recorded as configured-absent, not unchecked — the queries succeeded and returned nothing, rather than failing to run. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed verbatim: Hyper Backup tasks: 0; Backup repositories: {"offset": 0, "repo_list": [], "total": 0}; CloudSync: FAILED err=103 (API not available/registered on unit). The configured-absent reading holds — these are successful queries returning empty, not failures.
acceptance test INCOMPLETE — none supplied. Not invented. Note that a close condition cannot be stated responsibly while GOV-019 stands: snapshot configuration on this unit is unknown (API 403), so what protection exists is not yet established.
blocks
blocked_by GOV-019derived. Snapshot configuration unknown blocks assessment of this row's severity.
related_to VLN-023derived. Compounds the degraded array.
date_raised 2026-08-17
date_verified

The distinction between configured-absent and unchecked is load-bearing and is preserved as supplied: this is a positive finding of absence, not a gap in the enumeration. GOV-019 is the gap in the enumeration, and it sits next to this row rather than inside it.


VLN-027 — Meraki IPsec pre-shared keys returned in plaintext

Field Value
id VLN-027
title IPsec pre-shared keys for both VPN peers returned in plaintext by the Meraki thirdPartyVPNPeers endpoint
register Security findings
status open
severity HIGH (as stated)
owner R. Chhetry
evidence Meraki thirdPartyVPNPeers endpoint returned both peers' PSKs in plaintext; the values transited terminal output on 2026-08-13.
acceptance test Both rotated, coordinated with the GCP side, tunnel verified up.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Durable handling rule, recorded as supplied

Any read of Meraki VPN config exposes PSKs. This is a standing property of the endpoint, not a one-off. It applies to every future enumeration that touches Meraki VPN configuration, including any reconciliation control built per GOV-004.


VLN-028 — SMBv1 enabled with signing off on all three Synology units

Field Value
id VLN-028
title SMBv1 enabled on all three Synology units with server signing off
register Security findings
status open
severity HIGH (as stated)
owner R. Chhetry
evidence Synology DSM API v3, 2026-08-17. smb_min_protocol = 1 on all three units — the lowest selectable value, corresponding to SMB1/NT1 — with server signing off (0) on all three. NTLMv1 is off. duck and grebe are Windows SMB proxies. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed: ntlmv1_auth=False server_signing=0 workgroup=WORKGROUP.
acceptance test Minimum protocol raised, signing enabled, verified per unit.
blocks
blocked_by
date_raised 2026-08-17
date_verified

NTLMv1 being off is recorded as supplied and is not treated as mitigating: it narrows the finding, it does not close it.


VLN-029 — Key material in a control repo scheduled for deletion

Field Value
id VLN-029
title Key material in the control repo, identified by filename only — in a project scheduled for deletion
register Security findings
status open
severity HIGH (as stated)
owner R. Chhetry
evidence Partial. Supplied by filename count only — contents never read: ca/ 81 private-key files, dnssec/ 1 .private, htpasswd/ 18 hash files. The repo lives in a project scheduled for deletion. No read date was stated for the listing. To complete: the dated directory listing the counts came from. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage3b-dist-index.json. One of three counts reproduces. dnssec .private files = 1 ✅ as stated. ⚠ Discrepancy — figure did not reproduce. The index lists 19 files under htpasswd/, not 18, and 643 files under ca/ against a stated 81 — 643 is every file under ca/ (the sample includes ca/bin/p7x and ca/certs/*.crt), so the 81 is presumably a private-key subset whose filter is not recorded. The stated figures are left unchanged; the filter needs stating before either can be confirmed.
acceptance test Each item determined live or historical; anything live rotated or relocated.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Recorded as supplied: deleting the repo copy does not revoke anything still trusted. Same shape as VLN-024 — removal of the artifact is not remediation of the secret. Note GOV-014 requires mirroring these repos before any teardown touches those projects, which this row depends on in practice.


VLN-030 — Undocumented second VPN peer

Field Value
id VLN-030
title Undocumented second VPN peer Target (206.31.252.84), IKEv1, subnet 10.5.142.0/24
register Security findings
status open
severity MEDIUM (as stated)
owner R. Chhetry
evidence Meraki API, 2026-08-13. Peer Target, 206.31.252.84, IKEv1 (deprecated), subnet 10.5.142.0/24. Appears in no artifact, no brief and no diagram.
acceptance test Owner identified; tunnel justified in writing or removed.
blocks
blocked_by
date_raised 2026-08-17
date_verified

target also appears as one of the four vendors carrying a plaintext credential in VLN-024. The two are recorded separately and no relation is asserted — the name matching is not evidence that they are the same counterparty.


VLN-031 — Zero 2FA across all three Synology units

Field Value
id VLN-031
title Zero 2FA across all three Synology units; admin account present and not renamed
register Security findings
status open
severity MEDIUM (as stated)
owner R. Chhetry
evidence Synology DSM API, 2026-08-17. 2fa_status: false on every user on every unit. The admin account is present and not renamed, though expired/disabled. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed: 2fa_status=False on every enumerated user, and the admin account present with expired=now and desc='System default user' — present and not renamed, as stated. Note rchhetry reads 2fa_status=None rather than False, i.e. unreported rather than disabled.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

VLN-032 — Meraki org 395909 has no SSO and no IdP

Field Value
id VLN-032
title Meraki org 395909: SAML disabled, zero IdPs, all four admins on email authentication
register Security findings
status open
severity MEDIUM (as stated)
owner R. Chhetry
evidence Meraki API, 2026-08-13. saml.enabled false, zero IdPs configured, all four admins authenticationMethod: Email. Only one admin holds an API key. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task1-meraki-20260813.txt. The admins table confirms authMethod: Email across the org admins and a single hasApiKey: True. The capture also notes the admins endpoint returns no saml field at all, so the saml.enabled false claim rests on a separate call, not this table.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Recorded as supplied: the W2 SSO migration has not started, rather than being partially done. That distinction is the finding — a partially-migrated org and an untouched one look similar from a status report and are not the same risk. Read alongside VLN-021: 3 of the 4 admins on this org hold orgAccess: full, and this is the org whose appliance terminates the WDC VPN.


VLN-033 — synstorage volumes at 93.1% and 82.3%, status attention

Field Value
id VLN-033
title synstorage volume_1 at 93.1% and volume_3 at 82.3%, both status attention, no documented backup
register Security findings
status open
severity MEDIUM (as stated)
owner R. Chhetry
evidence Synology DSM API, 2026-08-17. volume_1 93.1% (13.44 TB) and volume_3 82.3% (187.64 TB), both status attention. Media/Video data with no documented backup. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt covers all three Synology units in one capture.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

VLN-034 — rooster cannot compile in production: archive::backup body commented out

Field Value
id VLN-034
title rooster.cloud.us.gl3 cannot compile in production — archive::backup has its entire body commented out
register Security findings
status open
severity MEDIUM (as stated)
owner R. Chhetry
evidence Stage 3 manifest read, 2026-08-17. archive::backup body commented out at lines 5–17 in the production environment, but defined normally in seed and testing. rooster is absent from PuppetDB, which is consistent with a host that cannot compile. Children archive::backup::files and archive::backup::cron still exist and are now unreachable.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

Third instance of LOADED ≠ REACHABLE

Recorded as supplied. The other two are VLN-010 (Wazuh baseline rules loaded, valid, counted and never reached) and the transfer crons in PRG-019 firing on schedule doing nothing. See M-01 in the Method learnings section of the gap register — this is named there as the estate's characteristic failure mode.


VLN-035 — dataflow.us.gl3 resolves to a host that does not exist

Field Value
id VLN-035
title dataflow.us.gl3 resolves to merlin, which has no VM in any project and is ASSERTED-STALE
register Security findings
status open
severity MEDIUM (as stated)
owner R. Chhetry
evidence zones/us.gl3.zone:61dataflow.us.gl3 resolves to merlin. merlin has no VM in any project and is classified ASSERTED-STALE. Every live vendor script posted status there.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
evidences PRG-019stated in the source row.
date_raised 2026-08-17
date_verified

Downgraded as supplied, and the downgrade is recorded rather than applied silently. The transfer stack is DECOMMISSION per R. Chhetry 2026-08-17 (PRG-019), so this is evidence supporting the decommission rather than an active incident. Severity is carried at MEDIUM as stated; the downgrade is in the framing, not in the severity field, because no revised severity was given.


VLN-036 — wdc-wap-5 status alerting

Field Value
id VLN-036
title wdc-wap-5 (MR57, WDC) status alerting
register Security findings
status open
severity LOW (as stated)
owner R. Chhetry
evidence Meraki API, 2026-08-13.
acceptance test INCOMPLETE — none supplied. Not invented.
blocks
blocked_by
date_raised 2026-08-17
date_verified

VLN-037 — Four production security reports ran on EOL Python 3.6.8

Field Value
id VLN-037
title The weekly, monthly, newsletter and quarterly security reports executed on MAPLE's system Python 3.6.8 — end-of-life since 2021-12-23 — from deployment (~2026-04) until 2026-08-20
register Security findings
status remediated
severity HIGH
owner R. Chhetry
evidence On MAPLE, 2026-08-20: python3 --versionPython 3.6.8; ./venv/bin/python3 --versionPython 3.6.8 for a venv built with bare python3; python3.11 --versionPython 3.11.13 (present but unused by the reports). /usr/bin/python3 -m pip freeze returned the interpreter's report dependencies: reportlab==3.6.8, Pillow==8.4.0, requests==2.27.1, google-cloud-storage==2.0.0 — all 3.6-era, none installable on 3.11.
acceptance test /opt/gpus-reports/venv/bin/python3 --version reports 3.11.x, and all four SOC reports render from that venv. Both executed 2026-08-20: venv rebuilt from python3.11 (3.11.13); weekly, monthly, newsletter, quarterly each rendered OK.
blocks
blocked_by
date_raised 2026-08-20
date_verified 2026-08-20

Why this is a security finding and not a hygiene note. Python 3.6 stopped receiving security patches on 2021-12-23, so the interpreter and every one of the four pinned dependencies above had been outside their support window for over four and a half years. Pillow 8.4.0 and requests 2.27.1 in particular are image-parsing and HTTP-client code with published CVEs in the intervening releases. The exposure is bounded — these are scheduled local jobs processing data from the estate's own SOC API, not an internet-facing listener — which is why this is HIGH rather than CRITICAL. It is not LOW: the process runs as root from the root crontab, and report_generator.py parses remote JSON and renders it through reportlab and Pillow.

How it was found, which is the part worth keeping. Nothing detected this. No scanner flagged it, no dashboard showed it, and the four reports had been arriving on schedule and rendering correctly the entire time — the only observable signal was a working report. It surfaced incidentally on 2026-08-20 while building a venv for a fifth report (R2, the daily travel approval exceptions report): R2's Cloud SQL dependencies require Python ≥ 3.8, which forced the interpreter version to be looked at for the first time since deployment.

This is the estate's recurring failure mode, again. Cf. VLN-010 (Wazuh rules loaded, valid, counted, and unreachable) and VLN-011 (routing recorded success=true while dispatching nothing): a component whose every observable says healthy while the substance is wrong. An EOL interpreter produces correct output, so correct output is not evidence of a supported runtime.

Remediated 2026-08-20. All five reports now run from /opt/gpus-reports/venv, built from python3.11 (3.11.13) with a pinned requirements.txtreportlab==5.0.1, Pillow==12.3.0, requests==2.34.2, google-cloud-storage==3.13.1. report_cron.sh invokes that interpreter explicitly and refuses to fall back to system python3. The routing worker's venv was deliberately not reused, so a report dependency change cannot break mail routing.

Residual, and it is not this row's to close. MAPLE's system Python is still 3.6.8 and is still the default python3 for anything else on that host — root cron entries, ad-hoc scripts, /usr/local/bin/gpus-cloud-backup.sh. This finding covers the reports only; that wider scope was not assessed here and is not claimed to be clean.

Raised as owned work, not left as a residual field: PRG-027Enumerate every cron job, systemd unit and script on MAPLE that invokes the EOL 3.6.8 interpreter. Owner R. Chhetry, raised 2026-08-20, with an acceptance test.

A residual has no owner and no acceptance test, so it reads as closed to anyone scanning the register — which is the same failure mode as this finding itself.

VLN-038 — No Cloud Build in the estate ran unit tests

Field Value
id VLN-038
title Every Cloud Build trigger in the estate built and deployed without executing a single unit test; the only security scans were \|\| true. Controls documented as "enforced by a test" were enforced by nothing automatic.
register Security findings
status in progress — one gate added 2026-08-24, six suites and two components still ungated
severity High — argued below, not supplied
owner R. Chhetry
evidence Confirmed first-hand 2026-08-24 by reading all eleven cloudbuild.yaml files. A grep for every test runner (unittest, pytest, tox, npm test, vitest, jest, go test) across all of them returns exactly one lineforms-backend/cloudbuild.yaml:34, added the same day. Before that: zero. The only security scanning is in forms-backend and gpus-forms-clamav-worker, and every line of it ends \|\| true: bandit -q -r forms-backend/ -x forms-backend/tests -ll \|\| true and pip-audit --requirement forms-backend/requirements.txt --strict \|\| true. Meanwhile 450 tests exist and pass — forms-backend 208, gpus-forms-routing-worker 91, gpus-reports 151 — of which 5 run in CI. (Was 381 when this row was raised on 2026-08-24; gpus-reports gained 69 the same day. See the regression note.)
acceptance test scripts/check-test-gates.sh exits 0. It enumerates every directory containing test_*.py, and reports a gap when that directory has no cloudbuild.yaml, or has one that invokes no test runner, or runs only some of its suites. Re-runnable, so it confirms the gap stays closed rather than certifying a one-time review. Run 2026-08-24: 3 gaps, exit 1.
blocks
blocked_by
date_raised 2026-08-24
date_verified

Every observable was green and nothing was being verified — the VLN-010 / VLN-011 / VLN-037 shape

This is the fourth instance of one pattern this quarter, and the pattern is the finding rather than any single case:

  • VLN-010 — the rules loaded, the count was correct, validation passed. Three baseline rules had never matched an event. Only a trace revealed it.
  • VLN-011 — the dispatcher recorded success=true and continued. No HappyFox client existed anywhere in the repo; 0 of 37 submissions ever got a ticket, while the UI told submitters the team had been notified.
  • VLN-037 — four reports rendered correctly for four months on an interpreter that had been end-of-life since 2021.
  • VLN-038 (this row) — builds went green, and "203/203 tests pass" was reported repeatedly during this workstream. Both statements were true. Neither was produced by the pipeline.

In every case the control ran, reported success, and verified nothing of substance. A test suite that executes only when a person remembers to run it is indistinguishable, from outside, from a test suite that does not exist — and the thing that makes it dangerous is that it is more reassuring than having no tests, because the count is quotable.

Why High: at least one documented security control is unenforced

Severity is argued, not supplied. The case for High does not rest on the missing gate itself — it rests on what was resting on it.

gpus-reports/test_approval_report.py::TestNoApproverText carries this docstring:

"approval_steps carries a free-text column holding an approver's own words about someone's travel — often why it was declined. maple-agent can read it. NOTHING IN THE DATABASE STOPS THIS REPORT FROM PRINTING IT. These tests are what stops it. They were the condition on which the grant decision was answered 'no grant needed'; weakening one does not just fail a test, it removes the control."

A database grant was declined on the strength of that test. The test lives in gpus-reports/, which has no cloudbuild.yaml at all — it deploys by gsutil to MAPLE. So the control that replaced a grant has never been executed by anything except a person choosing to run it. The same holds for forms-backend/approval_access.py, whose module docstring names test_membership_removal_does_not_revoke as the guard on an authorization boundary that "nothing below the application" enforces.

The counter-argument, stated fairly: nothing is currently broken. All 381 tests pass today, the controls hold, and there is no exploitation path — this is a latent verification gap, not a live compromise, which is why the row is in progress rather than an incident. The reason it still reads High is that the failure mode is silent and unbounded: nothing would report the day one of those tests started failing, and the estate has now been surprised three times by exactly that.

2026-08-24 — this got WORSE, not better, on the day it was raised

R5 (the daily travel volume report) shipped to gpus-reports/ the same day this row was raised. It added 69 tests to a component with no build pipeline at all, taking it from 82 to 151 and the estate from 381 to 450. The number running in CI did not move: still 5.

The largest gap in this row grew by 84% while the row was open, and nothing about that was careless — R5 was tested more thoroughly than most things in the estate, which is precisely the point. Writing good tests into a component with no pipeline produces more unexecuted assurance, not less. The count is quotable either way.

Recorded rather than fixed, deliberately: adding a pipeline to gpus-reports is the decision described below, and taking it under schedule pressure from a report deploy is how it would get taken badly.

Consequence for the exception date — date it by the DECISION, not the difficulty

When VLN-038's exceptions are written (see the acceptance test), the instinct will be to give gpus-reports the longest date because it is the hardest gap to close. That is backwards and it should be resisted.

gpus-reports holds test_approval_report.py::TestNoApproverText — the control accepted in place of a database grant — and now 151 tests, none of which anything executes. It is simultaneously the largest exposure in this row and the hardest to close, and dating by difficulty puts the longest clock on the biggest problem. It just grew by 69 tests, which makes the mismatch worse.

The date should be set by when the DECISION is due, not by when a pipeline could be built. The decision is a narrow one and does not need a pipeline to exist first: should gpus-reports and gpus-forms-routing-worker have Cloud Build triggers at all, given both ship by gsutil and neither is a Cloud Run service? That question can be answered in an afternoon. Building whatever the answer implies can take as long as it takes, on its own row.

Conflating the two is what would produce a twelve-month exception on the estate's largest block of unexecuted tests, and it would look reasonable the whole time.

Enumeration — deliberately NOT fixed in this pass

Adding test gates to builds that have never had them will break builds, and each needs its own decision about what "green" should mean. Recorded so the decisions can be made one at a time:

Component Build Tests present Blocking checks today
forms-backend 208 field-exposure-check; travel-instructions-check (1 of 7 suites)
gpus-forms-routing-worker ❌ none 91 nothing — no pipeline exists
gpus-reports ❌ none 151 nothing — deploys by gsutil
gpus-forms-clamav-worker 0 scan only, \|\| true
forms-frontend 0 none
soc-backend 0 coverage-check
status-backend 0 coverage-check
soc-site 0 coverage-check
status-site 0 coverage-check
mkdocs-portal 0 coverage-check
security-backend 0 none
security-site 0 none
forms-backend/migrate 0 none

Two distinct problems, and they need different fixes. Six forms-backend suites are ungated inside a build that now runs one — cheap to close, and the reason the current gate is PARTIAL rather than green. Two components with 242 tests between them have no pipeline at all — that is not a cloudbuild edit, it is a decision about whether gpus-reports and the routing worker should have builds, and it is the larger of the two.

The coverage-check steps are real blocking gates and are not criticised here; they verify inventory coverage, which is a different question from whether the code works.


VLN-039 — Pre-build check steps couple every deploy to PyPI availability

Field Value
id VLN-039
title Cloud Build steps make network calls at build time with default retry settings, so a transient upstream timeout fails a build that has nothing wrong with it. Raised for pre-build pip install from PyPI; widened 2026-08-27 after the same shape appeared at the image push step, which no pip setting reaches
register Security findings
status in progress — pip half mitigated on forms-backend 2026-08-24, six other builds unchanged; the registry-push half is unmitigated everywhere (second instance 2026-08-27)
severity Medium — argued below, not supplied
owner R. Chhetry
evidence Observed, not predicted. Build 8356f08a-6bf1-49eb-b38a-1e74ee6df218 (2026-08-24, gpus-forms-backend-trigger, commit 934c37f) failed with pip._vendor.urllib3.exceptions.ReadTimeoutError: HTTPSConnectionPool(host='files.pythonhosted.org', port=443): Read timed out. in step 0, field-exposure-check, during pip install --quiet pyyaml. Only step 0 started and finished — the enumerated step list contains that one entry and no other. A re-run of the same trigger at the same commit (a6637d70) passed with no code change. Across the estate, 7 of 11 cloudbuild.yaml files run pip install in a pre-build step — 11 install commands in total: forms-backend 3 (pyyaml; requirements.txt, 21 packages; bandit + pip-audit), soc-site 2, status-site 2, and mkdocs-portal, soc-backend, status-backend, gpus-forms-clamav-worker 1 each.
acceptance test No Cloud Build step that gates a deploy fails on a transient third-party network error. Two halves, because the second instance proved one is not enough. (a) Dependency installs — every pre-build pip install in */cloudbuild.yaml carries explicit --retries and --timeout (grep, re-runnable), and no build in a 30-day window shows a ReadTimeoutError/ConnectionError from files.pythonhosted.org. (b) Registry I/O — no build in the same window fails on an i/o timeout or connection reset against *.pkg.dev, or a documented retry exists for the push/deploy steps. (b) cannot be satisfied by any pip flag and was not covered by the original test.
blocks
blocked_by
date_raised 2026-08-24
date_verified

SECOND INSTANCE 2026-08-27 — different step, and pip settings do not reach it

Build fc88f7a8-1470-4f38-bc0b-9c66e7997f7e (gpus-deploy-mkdocs-portal, commit 9c32983) failed at step 4, the image push — not at a pip install, and not at a pre-build check:

Step #4: 3a98051e9797: Pushed
Step #4: 7da9fcd3e4f6: Pushed
Step #4: 746b97c19b23: Pushed
Step #4: Head "https://us-central1-docker.pkg.dev/v2/gpus-infra/gpus-images/
         mkdocs/blobs/sha256:abf8772a…": dial tcp 142.250.125.82:443: i/o timeout
ERROR: build step 4 "gcr.io/cloud-builders/docker" failed: step exited with
       non-zero status: 1

The image built. Every layer pushed. The final manifest HEAD against Artifact Registry timed out. Re-running the trigger at the same commit succeeded with no change — the same signature as the 2026-08-24 PyPI failure, one step later in the pipeline.

This is what widens the row rather than repeating it. The finding was framed as "pre-build check steps couple every deploy to PyPI", and its acceptance test was written entirely in those terms — --retries and --timeout on every pip install. Every one of those mitigations was already in place for this build and none of them could have helped, because the call that failed was docker push to *.pkg.dev. Build-time network fragility is not confined to dependency installs; the row's original scope was the shape of the first instance, not the shape of the problem.

Note also which direction the exposure runs: this failure is safer than the pip one in that nothing shipped — but the deploy was blocked by an event with no relationship to the change, which is the row's actual harm.

The status is not the finding — reading it alone would have been wrong

Recorded because it nearly produced a false report.

The commit touched inventory.yaml, both portals, and five documents, and fired five triggers. Four succeeded; gpus-deploy-mkdocs-portal showed FAILURE. On a commit that had just added four inventory entries and a supersession block to a published doc, a mkdocs build failure reads exactly like a content regression — a malformed table, a broken link, a bad YAML front-matter block.

It was none of those. The coverage check in step 0 had already passed (Errors: 0, RESULT: PASS), MkDocs had built the site, the image was tagged, and the failure was a TCP timeout two steps later. Only the log distinguishes those two stories, and they call for opposite responses: revert-and-fix versus re-run unchanged.

This is the same lesson as the frontend test gate's first run — a green status means least exactly when a pipeline has just changed — inverted. Here a red status meant least, and for the same reason: the status is a summary of the last thing that happened, not of what the build was for.

Why this is a SEPARATE row from VLN-038 and not an update to it

The two are adjacent and they are not the same finding. VLN-038 is "controls that never execute" — a silent absence. This is "a control that executes and blocks legitimate work" — a loud presence. They have different failure modes, different acceptance tests, and different fixes.

The decisive argument is that they move in opposite directions. Closing VLN-038 means adding gates. Every gate added is another pre-build step, another network round-trip, and another way for a deploy to be blocked — so progress on VLN-038 makes VLN-039 worse. That is exactly what happened: the unit-tests step that closed VLN-038's cheap half installs 21 packages where the step beside it installs one, and it was added hours before this failure.

A row whose remediation worsens another row must not be that row. Folded together they would need one acceptance test spanning "every suite runs" and "no deploy is blocked by an outage", and no single check can say both.

Why Medium, and the argument against calling it High

The consequence is real and specific: a backend fix during an incident can be blocked by a read timeout in a test step. The gate has no knowledge of urgency, and the failure happens before any of the work being deployed is even evaluated.

But the honest read is Medium, not High, and the reason is that the mitigation is immediate and always available: a re-run at the same commit cleared it in under four minutes with no code change. Nothing is permanently blocked, no data is exposed, and an operator under pressure has a working action. High would put it beside VLN-010, VLN-011 and VLN-037 — findings where the estate was wrong about its own state for months without knowing. This is a known, visible, retryable delay.

It is not Low because the delay lands precisely when delay is most expensive, and because the number of round-trips is growing rather than shrinking.

Mitigated on forms-backend; what it does NOT solve

All three forms-backend pre-build installs now carry --retries 10 --timeout 60.

--retries 5 is already pip's default, so raising the retry count alone is close to a no-op — the lever that actually expired here is --timeout, whose default is 15 seconds. Both are now explicit rather than inherited from defaults that can change.

THIS DOES NOT DECOUPLE DEPLOYS FROM PyPI, and must not be read as though it does. forms-backend/Dockerfile:10 runs pip install --no-cache-dir -r requirements.txt to build the image that actually ships. The coupling predates every test gate; what the gates added was two earlier round-trips and a much larger package set. Lowering the probability of a transient failure is all this change does.

The rejected option, and why — a pre-baked test image

The alternative was an Artifact Registry image carrying the test dependencies, used as the step image. It removes the round-trip outright, which is strictly better than tuning timeouts. It was rejected here for two reasons.

First, it does not deliver its headline benefit. It removes one of three round-trips and leaves the Dockerfile's, so the deploy remains coupled to PyPI either way — the incident scenario above is unchanged.

Second, and decisively: it creates a second copy of the dependency set that must be kept in step with requirements.txt. This estate's characteristic failure is two copies drifting — servers.py against inventory.yaml, gpus-reports/ against soc-backend/ (99 lines), the travel form's instructions against STEP_DESCRIPTIONS (GOV-021), and the read-path argument that decided Option A in migration 022. A drifted test image is worse than a flaky build, because tests would pass against versions the shipped image does not contain, and nothing would report it.

It becomes the right answer if the flakiness persists after this change, or if the estate adopts a general staleness check for built images. Either would have to come first.

Adjacent finding — the security scan can no-op invisibly

forms-backend and gpus-forms-clamav-worker run pip install bandit pip-audit without &&, and the block's last command ends || true. So if that install fails, the step still reports success and scans nothing. Unlike the gates above it fails open rather than closed.

Not fixed here — changing it to fail closed would make a scanning outage block deploys, which is the very problem this row is about, and that trade-off deserves its own decision rather than being made in passing.


VLN-040 — The dependency scan reports 24 advisories into || true

Field Value
id VLN-040
title pip-audit runs on every forms-backend build, correctly reports 24 known vulnerabilities in 8 production dependencies, and its exit code is discarded by \|\| true. The build is green. The finding has been produced on every retained build and acted on by nothing.
register Security findings
status in progress — triaged 2026-08-24; step 1 of 6 done 2026-08-25 (flask-cors cleared, deployed and verified). 19 advisories remain across 7 packages and \|\| true is still in place.
severity High
owner R. Chhetry
evidence Build a6637d70 (2026-08-24) step 2 security-scan: Found 24 known vulnerabilities in 8 packages, build status SUCCESS. The step is quoted below. Duration: the \|\| true was present in the FIRST forms-backend commit, 802b42c (2026-04-20) — 127 days. Cloud Build log bodies expire; the earliest retained log is 2026-07-27, reporting Found 22 known vulnerabilities in 8 packages. 41 builds since then, every one green, every one carrying the finding.
acceptance test pip-audit --requirement forms-backend/requirements.txt --strict --ignore-vuln <each id in .pip-audit-exceptions with an unexpired date> exits 0. Deliberately NOT "the scan passes" — it already does that. The exception file must carry a dated entry per ignored id, per the .coverage-exceptions.yaml convention.
blocks
blocked_by
date_raised 2026-08-24
date_verified
# forms-backend/cloudbuild.yaml, unchanged since 802b42c:
        bandit -q -r forms-backend/ -x forms-backend/tests -ll || true
        pip-audit --requirement forms-backend/requirements.txt --strict || true
# gpus-forms-clamav-worker/cloudbuild.yaml:
        bandit -q -r gpus-forms-clamav-worker/ -ll || true
        pip-audit --requirement gpus-forms-clamav-worker/requirements.txt --strict || true

VLN-010's shape, not VLN-039's — and the three must not be read as one

Three findings now concern this pipeline and they are genuinely different.

  • VLN-038tests that never ran. A control absent. Fixed by adding it.
  • VLN-039a gate that blocks legitimate work. A control too coupled. Fixed by loosening it.
  • VLN-040 (this row)a control that ran, produced correct output, and was wired to nothing. This is VLN-010's shape: the Wazuh rules loaded, validated, counted, and never matched. Here the scanner installed, ran, correctly identified 24 real advisories, printed them — and \|\| true threw the exit code away, so the pipeline treated a true report as no report.

The distinction that matters operationally: VLN-038 was invisible, and this one has been on screen in every build log for at least 41 builds. Nobody was hiding it. It was reported, in full, and discarded by a two-character idiom.

Triage — 24 advisories, and most are NOT reachable. The count is not the finding.

Reachability assessed against actual estate usage, not against the package's full API. Where it could not be established it is marked unknown rather than guessed in either direction.

Package Advisories Reachable? Basis
cryptography 42.0.7 8 5 NO, 3 UNKNOWN The estate's ONLY import is AESGCM (forms-backend/crypto.py:16, worker mirror). No X.509, no key loading, no certificate verification anywhere.
python-jose 3.3.0 5 NO (all 5) See the two notes below — one blocked empirically, four need JWE, which is never used.
~~flask-cors 4.0.1~~ → 6.0.5 ~~5~~ 0 CLEARED 2026-08-25 Was: CORS(app, origins="*") at app.py:70, in the request path. Bumped in 58567f0. See the step-1 note below — including a correction to this row's own "YES": the library is in the request path, but 3 of the 4 distinct defects were not.
requests 2.32.2 2 UNLIKELY Used for JWKS + SOC API with fixed URLs. CVE-2024-47081 is a .netrc leak via malicious URLs; no .netrc, no caller-supplied URLs.
flask 3.0.3 1 NO CVE-2026-27205 is a missing Vary: Cookie on session access. The portal uses no Flask session — it is a JWT API with no cookies.
pg8000 1.31.2 1 NO CVE-2025-61385 is SQLi via pg8000.native.literal with a crafted list. The estate uses SQLAlchemy bound parameters; pg8000.native is never imported.
protobuf 4.25.9 1 UNKNOWN Transitive via google-cloud-*. Not called directly.
ecdsa 0.19.2 1 NO Minerva timing attack on P-256. Transitive via python-jose[cryptography]; the portal verifies RS256 against an RSA JWK, so no EC path is exercised.

The three cryptography UNKNOWNs are the bundled-OpenSSL advisories (GHSA-h4gh-qq45-vh27, GHSA-537c-gmf6-5ccf, PYSEC-2026-1284). Those are flaws in the OpenSSL statically linked into the wheel, not in cryptography's Python API. AESGCM does reach OpenSSL's EVP layer, so they cannot be dismissed the way the X.509 ones can — but establishing whether the specific OpenSSL defects touch AES-GCM requires reading the OpenSSL advisories of 2024-09-03 and 2026-06-09, which has not been done. Marked unknown, and it is the gap this row owes.

A finding the scan did NOT report — algorithms is taken from the attacker's own token

Assessing CVE-2024-33663 (python-jose algorithm confusion) required reading the calling code, and the calling code has a defect the scanner cannot see. auth_v2.py:163-166:

claims = jwt.decode(
    token,
    key,
    algorithms=[unverified_header.get("alg", "RS256")],

The allowlist is derived from the unverified header of the token being validated. An attacker controls alg, so they choose the algorithm their own token is checked against. That is the precise shape algorithm-confusion attacks exploit, and it is wrong independently of any CVE — an allowlist the caller supplies is not an allowlist.

It is not currently exploitable, and that was verified rather than assumed. _key_for_kid returns the Okta JWK dict (kty: "RSA"), and python-jose refuses to build an HMAC key from it. Tested against the pinned 3.3.0:

jwk.construct(RSA-JWK, 'HS256') -> JWKError: Incorrect key type. Expected: 'oct', Received: RSA
jwk.construct(RSA-JWK, 'RS256') -> RSAKey

The defence is incidental and one refactor away. It lives in python-jose's kty check, not in the portal's intent. Anyone "simplifying" _key_for_kid to return a PEM string — a natural-looking tidy-up — removes it silently, and the tests would not notice.

Fix regardless of patching: pin algorithms=["RS256"]. Cheap, independent of the version bump, and it makes the defence deliberate.

Raised as its own row — VLN-041 — and pinned there on 2026-08-25 with its own tests. This row keeps the provenance (the defect was found only because assessing CVE-2024-33663 forced a read of the calling code) and stops there: VLN-041 is a defect in this estate's own source with a different fix, a different acceptance test and a different lifecycle, and must not close or reopen with the scan-gating work tracked here. One correction to the empirical note above, made while writing VLN-041's tests: the kty check is not the only incidental defence — see that row.

Why python-jose's other four are not reachable

CVE-2024-33664 and CVE-2024-29370 (each listed twice, making four of the five) are both denial-of-service via jwe.decrypt on a high-compression JWE — a "JWT bomb". The estate never decrypts JWE. Okta issues signed (JWS) tokens; a grep for jwe across forms-backend/ and gpus-forms-routing-worker/ returns nothing. Not reachable.

Two advisories have NO PUBLISHED FIX — this decides the remediation order

  • PYSEC-2026-1325ecdsa Minerva timing attack. No fixed version.
  • PYSEC-2025-185python-jose jwe.decrypt DoS. No fixed version.

Both are assessed not reachable. But their existence means pip-audit --strict can never exit 0 by patching alone, so removing \|\| true without an exception mechanism would break the deploy pipeline permanently rather than temporarily. That is why the acceptance test above is written around a dated exception file and not around a clean scan.

Fixed versions, and where a bump will fight the pins

Package Now Fix Risk
flask-cors 4.0.1 6.0.0 Major ×2. The only REACHABLE cluster — highest value, highest churn.
cryptography 42.0.7 up to 49.0.0 Seven majors. python-jose[cryptography] depends on it; bump them together.
python-jose 3.3.0 3.4.0 Minor. Cheap.
flask 3.0.3 3.1.3 Minor.
requests 2.32.2 2.33.0 Minor.
pg8000 1.31.2 1.31.5 Patch. cloud-sql-python-connector[pg8000] pins a range — check it.
protobuf 4.25.9 5.29.6 / 6.33.5 Major, and google-cloud-* pin protobuf ranges. Most likely to conflict.
ecdsa 0.19.2 No fix. Exception.

Recommended order: patch first, then remove || true — option (a)

Argued, and the argument is that (b) inverts the pressure.

Removing \|\| true now turns 24 open advisories into a broken deploy pipeline immediately, and every one of them would have to be resolved or excepted before anything ships — including the two with no published fix, which can only ever be excepted. The estate would be patching cryptography across seven major versions and protobuf against google-cloud-* pins under a deploy freeze, which is the worst condition for a change that touches the library performing AES-GCM on every submission field. It also collides directly with VLN-039: a blocked pipeline is exactly the incident-response problem that row is about.

Under (a) each bump is verified on its own, the 264-test suite (VLN-038's gate) runs against it, and \|\| true comes off last — at which point the green build is a true statement rather than a suppressed one.

Suggested sequence, cheapest and most reachable first: 1. ~~flask-cors → 6.0.0~~ — DONE 2026-08-25, at 6.0.5 (the fixed line's current patch, not 6.0.0). See the step-1 note below. 2. python-jose → 3.4.0. The algorithms=["RS256"] pin is no longer part of this step — it shipped separately on 2026-08-25 under VLN-041, deliberately decoupled so an auth-path change was not carried by a dependency bump. test_auth_algorithm_pin.py is the regression gate for the bump: run it against 3.4.0 before merging. 3. flask, requests, pg8000 — minors and a patch. 4. cryptography → current, with python-jose. Largest jump; AESGCM is the only call site, so the blast radius is small and well-tested. 5. protobuf — last, because it is the one expected to fight the pins. 6. Except ecdsa and python-jose PYSEC-2025-185 with dates, then remove \|\| true.

What (a) does not solve: nothing is patched until someone does step 1. Until then this row's status is the honest one — the finding is known, triaged, and open.

STEP 1 DONE 2026-08-25 — flask-cors 4.0.1 → 6.0.5, deployed and verified

Commit 58567f0, build f9bb3d10, serving revision gpus-forms-backend-00108-sfw at digest sha256:1e21f227..., matching the digest that build pushed. Built is not deployed; this was checked by digest, not by tag.

The count, re-measured rather than subtracted. pip-audit --strict was re-run against the requirements file as deployed (git show 58567f0:forms-backend/requirements.txt), and independently by the build's own security-scan step on the file it shipped. Both report the same thing:

Found 19 known vulnerabilities in 7 packages     (was 24 in 8)
flask-cors advisory rows: 0

5 advisories were 4 defects. PYSEC-2024-71 and PYSEC-2024-260 are the same CVE-2024-6221 listed twice — the same double-count already noted for python-jose. Counting advisories overstates by one here, which is another reason the headline number is not the finding.

A correction to this row's own triage table. It recorded flask-cors as reachable — YES — on the grounds that CORS(app, origins="*") runs on every request. The library is in the request path; 3 of the 4 defects were not. CVE-2024-6866 (case-insensitive path match), CVE-2024-6839 (regex pattern priority) and CVE-2024-6844 (unquote_plus turning + into a space) are all path-matching defects, and each needs a per-path resources= config to have anything to mismatch between. The portal has one global policy and no per-path config. The fourth, CVE-2024-6221, sets Access-Control-Allow-Private-Network: true; the backend is a public Cloud Run service rather than a private-network target, so it was present but inert. Reachable "library" and reachable "advisory" are different claims and this table conflated them.

That does not make the bump optional — it makes its ordering load-bearing. See VLN-042: narrowing origins is normally written with a resources= dict, which is exactly the construct that makes all three path defects reachable. Bumping first is what keeps that narrowing from closing one row by opening three.

CVE-2024-6221 closed, observed on the live service rather than inferred from a version number:

OPTIONS /api/my-submissions
  Access-Control-Request-Private-Network: true
→ access-control-allow-private-network: false      (was `true` on 4.0.1)

The standard this bump set, and what the next one owes

None of the nine test suites exercises a browser preflight. Had the evidence been "282 tests pass", it would have covered none of the behaviour a CORS library bump can break. flask-cors sits in the request path on every request; the suite says nothing about that path.

What was done instead, and what the next dependency bump should be held to: the exact call the portal makesCORS(app, origins="*") — was driven through Flask's test client on both versions, and every Access-Control-* and Vary header compared across five request shapes: the SPA's real preflight (Authorization + Content-Type), a Private-Network preflight, a simple cross-origin GET on the unauthenticated /metrics, a request with no Origin header (the nginx same-origin path), and a case-variant path.

Exactly one header differed, in one caseAccess-Control-Allow-Private- Network: true → false, which is precisely CVE-2024-6221's fix. Everything else byte-identical. The deployed service was then probed directly and returned the same preflight headers the test client predicted.

That measurement is why this shipped without a staged rollout, and it is recorded here so the next bump is argued from a measurement of the changed behaviour rather than from a green suite. cryptography (step 4) is the obvious case: AESGCM is the estate's only call site, so the equivalent evidence is an encrypt/decrypt round-trip across both versions on the same ciphertext — not the suite, which stubs it.

What step 1 does NOT do, stated so the status is not misread: nothing else is patched, 19 advisories remain in 7 packages, and \|\| true is still discarding the scanner's exit code on every build. This row closes at step 6, not step 1.


VLN-041 — The JWT algorithm allowlist was taken from the token being validated

Field Value
id VLN-041
title forms-backend/auth_v2.py derived jwt.decode(algorithms=[...]) from the unverified header of the token being validated, so a caller chose which algorithm their own token was checked against. Classic JWT algorithm confusion, in the authentication path in front of every authz check and every decrypt endpoint in the forms portal.
register Security findings
status done — 2026-08-25. Both acceptance conditions met and observed; see the closure note below.
severity High (argued) — not exploitable on the pinned dependency set (see below), but the defect sits on the sole authentication path for every Phase 2 route, and the only thing preventing exploitation is third-party type checking that a routine refactor removes. Argued rather than supplied; no CVSS calculated.
owner R. Chhetry
evidence forms-backend/auth_v2.py:163-166 at commit 86c9eee, read 2026-08-25: algorithms=[unverified_header.get("alg", "RS256")], where unverified_header = jwt.get_unverified_header(token) at :150. Found while triaging CVE-2024-33663 for VLN-040pip-audit does not and cannot report it. Estate sweep, same date: one other decode site, forms-backend/auth.py:59-72, which takes algorithms=[key.get("alg", "RS256")] from the JWKS entry — server-controlled, not attacker-controlled, a materially different shape and not this defect. Nothing else in the estate decodes a JWT: gpus-forms-routing-worker uses google.auth only for signBlob; gpus-forms-clamav-worker leaves Pub/Sub push OIDC to Cloud Run's IAM layer; gpus-reports, scripts/, soc-backend, status-backend, security-backend have no jose import and no jwt requirement; the frontends do not parse tokens. Okta signing algorithm read live, not assumed: GET https://greenpeaceeu.okta.com/oauth2/v1/keys returns a single key {kty: RSA, alg: RS256, use: sig, kid: RWZkwN0Cn0ZK31J9hs2-sdu3YiKyUINyH2KJtD7RTdo}, and the discovery document advertises id_token_signing_alg_values_supported: ["RS256"].
acceptance test Two conditions, both required. (1) Live: the deployed gpus-forms-backend revision serving /api/forms is built from a commit in which python3 -m unittest test_auth_algorithm_pin passes, and a real Okta ID token still authenticates against it end-to-end — a pin that authenticates nobody is a worse defect, so the closing observation is a 200 on a real user's token, not a 401 on a forged one. (2) Structural: test_auth_algorithm_pin.TestSourceNeverReadsAlgFromTheToken is executed by a pipeline that can fail the build. Condition 2 cannot be met until VLN-038 lands; until then the guard protects a developer running the suite, not a deploy, and this row stays open on that basis rather than being closed on condition 1 alone.
blocks
blocked_by — (was VLN-038; cleared 2026-08-25 when that row's remediation put the suite behind a hard gate)
date_raised 2026-08-25
date_verified 2026-08-25
# forms-backend/auth_v2.py, before:
unverified_header = jwt.get_unverified_header(token)   # :150 — attacker-supplied
...
claims = jwt.decode(
    token,
    key,
    algorithms=[unverified_header.get("alg", "RS256")],  # :166 — and so is this
    audience=OKTA_CLIENT_ID,
    issuer=OKTA_ISSUER,
)

# after:
    algorithms=["RS256"],

CLOSURE — both conditions observed 2026-08-25, on the deployed revision

Condition 1 — a real Okta token returns 200 against the deployed revision. MET. Observed by R. Chhetry at 15:41:12 UTC: signed in at forms.greenpeace.us in a private window (empty sessionStorage, so a genuine Okta round-trip and a freshly issued token, not a rendered shell), and /my-requests listed 4 requests. Correlated in the Cloud Run request log:

TIMESTAMP  REVISION_NAME                 STATUS  REQUEST_URL
15:41:12   gpus-forms-backend-00107-n8k  200     .../api/my-submissions

/api/my-submissions is @require_auth, so that 200 is auth_v2._validate_token returning claims from a token verified against algorithms=["RS256"]. /api/me (also @require_auth) returned 200 twice in the same session, and /api/forms — the legacy auth.py path — returned 200 as well, so neither authentication path regressed.

The revision is the right one, by digest and not by tag. Serving revision gpus-forms-backend-00107-n8k, Ready=True, 100% of traffic, resolved image ...gpus-forms-backend@sha256:471a0c98dba6098b9ab7a58984b39b0e3f727a07e53c6845fac0208080876ed6 — byte-identical to the digest build df2e414c pushed for tag 1285a6b. Cloud Run pins the digest at deploy, so this cannot be a tag that moved afterwards.

Nobody was locked out — checked across all callers, not just one. Every 401 on the service since the revision went live at 14:37:46Z:

15:38:28  401  /api/my-submissions   userAgent: curl
15:38:27  401  /api/my-submissions   userAgent: curl

Both are the deliberate pre-flight probes described below. Zero 401s from any real client. That is the observation this condition was written to require: a pin that authenticates nobody is a worse defect than the one it fixes, so the closing evidence is a success, not a rejection.

Pre-flight probes against the live revision, run before the login so a failure would not be mistaken for a lockout — and worth keeping, because they exercise the rejection path end-to-end through the real deployment:

GET /health                                    -> 200
GET /api/my-submissions  (garbage bearer)      -> 401 {"error":"invalid_token",
                             "detail":"Malformed token header: ..."}
GET /api/my-submissions  (alg:HS256, fake kid) -> 401 {"error":"invalid_token",
                             "detail":"Signing key not found in JWKS"}

The second is the attack's shape, refused at the kid lookup before the algorithm check is even reached — and refused with the portal's own error contract. Under the old code the incidental defence raised JWKError, which is not a JWTError and escaped uncaught: a 500. The clean 401 is itself evidence the pinned path is what is running.

Condition 2 — the structural guard is executed by a pipeline that can fail the build. MET, and evidenced rather than argued. The unit-tests step in forms-backend/cloudbuild.yaml is a hard gate with no \|\| true, runs unittest discover, and demonstrably failed a build over this exact suite: build 34fe8dcb (8b7800c) died at step 1 with ModuleNotFoundError: No module named 'cryptography.hazmat.primitives.asymmetric', Ran 265 tests, FAILED (errors=1), build step 1 ... exited with non-zero status: 1. No build, push or deploy step ran. Build df2e414c (1285a6b) then reported Ran 282 tests ... OK with all nine suites named in the step output, test_auth_algorithm_pin among them and all 18 of its tests listed individually.

blocked_by: VLN-038 is cleared. That row's remediation is what makes this condition true, and its first real catch was a defect in this row's own commit.

The finding is that the defence was incidental — not that the attack failed

This was not exploitable on the pinned dependency set, and that was established empirically rather than argued. It is recorded as High anyway, because nothing in this repository was doing the defending. A forged HS256 token was built by hand and put through _validate_token against each shape _key_for_kid could return:

_key_for_kid returns old — algorithms=[header alg] new — algorithms=["RS256"]
JWK dict (today) JWKError: Incorrect key type. Expected: 'oct', Received: RSA rejected by the pin
PEM string JWKError: ...asymmetric key... should not be used as an HMAC secret rejected by the pin
plain secret ACCEPTED — token minted a submitter rejected by the pin

Every cell that says "rejected" in the old column is python-jose's, and none of them is the portal's. The VLN-040 note called out one such check (kty); writing the tests found a second (HMACKey's asymmetric-key guard), which is why the PEM refactor that note warned about would in fact still have been caught on 3.3.0. That does not weaken the finding — it means the estate was relying on two library behaviours it never asked for, across a dependency the same triage is scheduled to bump. The third row is the one with no library defence in it at all.

Two further observations from the same test run. JWKError is not a subclass of JWTError, which _validate_token catches — so under the old code the incidental defence escaped uncaught and would have produced a 500, not a 401. The portal's own error contract was never reached. And once pinned, python-jose raises JWTError: The specified alg value is not allowed, which the existing handler turns into a clean 401.

Why this is its own row and not an update to VLN-040

It was found inside VLN-040's triage, and the provenance is recorded there. It does not belong there.

  • VLN-040 is a pipeline defect — a scanner's exit code discarded. Its fix is cloudbuild.yaml and a dated exception file; it closes when pip-audit --strict exits 0.
  • VLN-041 is a source defect in this estate's own code. No scanner reports it, no version bump fixes it, and no exception file could ever cover it. It closes when a deployed revision carries the pin.

Different fix, different acceptance test, different owner surface, different lifecycle. Nested inside VLN-040 it would have no id to cite from a commit, no line in the Index, and it would reopen every time the dependency work regressed. The distinction is not bookkeeping: a dependency advisory is someone else's defect that the estate inherits and can defer with a dated exception, whereas this is the estate's own, and "not currently exploitable" is not a disposition that can be written against it.

What a legitimate Okta change would do to the pin, and what it would not

A pin that breaks on a legitimate change would be a different problem, not a better one, so this was checked rather than assumed.

Key rotation does not touch it. Okta rotating to a new kid with the same RS256 is the routine event, and _key_for_kid already force-refreshes the JWKS cache and retries once (auth_v2.py:129-134). The algorithm is unchanged; the pin is unaffected.

The org authorization server cannot change the algorithm. 0oazce9jd5inTjtrn417 is an app on https://greenpeaceeu.okta.com — the org AS, no /oauth2/<asId> path segment. It signs ID tokens RS256 only and exposes no per-app signing-algorithm setting.

Only a migration to a CUSTOM authorization server could, since those permit RS384/RS512/ES256 and others. That is a deliberate configuration project, not something Okta does to the estate unannounced. Handling: widen the list from the new server's own id_token_signing_alg_values_supported, in the same change as the cutover. test_there_is_exactly_one_place_to_change_on_an_okta_migration asserts there is exactly one algorithms= in the module, so that stays a one-line edit in a known place.

auth.py:68 — a known second instance. DECIDED 2026-08-25: it stays.

The Phase 1 module takes its allowlist from key.get("alg", "RS256") where key is the JWKS entry fetched from Okta over TLS. The attacker does not supply it, so this is not the defect this row is about and pinning it would close nothing.

DECIDED 2026-08-25 (R. Chhetry): auth.py:68 is not changed. Not attacker-controlled, and auth.py is scheduled for retirement under the legacy-authz work — incremental hardening of code with a removal date gets deleted along with the code.

It is recorded here so the retirement does not ship the idiom forward. That is the whole point of the entry: /api/admin/reload, /api/forms and /api/pulldowns still authenticate through auth.verify_token, and when that work migrates them to auth_v2 or to whatever replaces it, the replacement must carry algorithms=["RS256"] as a literal. Copying key.get("alg", ...) across would reintroduce a caller-supplied allowlist into new code, where the argument that it is server-controlled would have to be re-established from scratch rather than inherited. Acceptance condition on the retirement work, not on this row: when auth.py is deleted, no module that absorbs its routes reads alg from anything — neither a token header nor a JWKS entry. test_auth_algorithm_pin's AST assertion is the pattern to copy; it is written against auth_v2 and would need pointing at the replacement module.

Test coverage as shipped

forms-backend/test_auth_algorithm_pin.py — 18 tests, all passing. Real crypto: unlike the sibling suites it does not stub jose, because what python-jose actually does is the subject. No network: _key_for_kid is monkeypatched, so _get_jwks never runs.

  • Behavioural — a forged HS256 token rejected across all three key shapes in the table above; a valid RS256 token still verifies against both a JWK and a PEM and still mints a submitter; a token signed by the wrong RSA key still rejected (the pin is not a skipped signature check); an unknown kid still fails closed.
  • DiscriminatingTestOnlyThePinRejectsIt asserts that the same token and the same key python-jose accepts under ["HS256"] is rejected under ["RS256"]. This is what makes the pin load-bearing rather than decorative, and it is the assertion the other two key shapes cannot make.
  • Structural — same shape as TestNoDecryptInQueueModule: no alg read from the unverified header in comment-stripped source, every unverified_header line is the assignment or .get("kid"), exactly one algorithms=. Two AST assertions go tighter than a grep can — no ast.Constant in the module has the value "alg" (immune to spelling, and docstring prose cannot hide in an AST), and every algorithms= kwarg must be a list literal of string constants equal to ["RS256"], so algorithms=[anything_computed] fails whatever the computation is.

Mutation-proved, not assumed. The old line was reintroduced, confirmed to compile (py_compile clean) and confirmed live in the file, and the suite re-run: 9 of 18 failed — 5 structural, 4 behavioural, the latter including the two JWKError escapes described above. Restored; 18/18 pass. The eight pre-existing forms-backend suites are unaffected.

The suite passed standalone and did not run at all under the build's own command

Recorded because it is the same failure spine as VLN-010, VLN-017 and VLN-040 — a control that looks present and does nothing — and this row nearly shipped an instance of it.

forms-backend/cloudbuild.yaml runs python3 -m unittest discover -p 'test_*.py'. The pin suite was written and verified with python3 -m unittest test_auth_algorithm_pin. Under discover it failed to import: test_approval_access and test_approval_persist register in-memory placeholder modules for jose and cryptography at import time so they can run where those packages are absent, test_approval_* sorts before test_auth_*, and this is the one suite that must not be stubbed. The unit-tests step is a hard gate with no \|\| true, so the build for 8b7800c fails there and the pin does not deploy — the fix is 26021a3.

Two things worth keeping from it. The gate worked: VLN-038's remediation caught a broken test on its first real exercise, which is exactly what it was added for. And "the tests pass" is not a claim about the pipeline unless the tests were run the way the pipeline runs them; standalone and discover are different commands and only one of them gates a deploy.

STANDING HAZARD for every test added to forms-backend — not a one-off

The paragraph above describes how one suite broke. This is the general form, recorded here because there is nowhere else it currently lives and the next person to add a test file will hit it without warning.

1. A suite verified by running it directly is not verified for the pipeline. python3 -m unittest test_x and python3 -m unittest discover -p 'test_*.py' differ in more than convenience: under discover, every test module in the directory is imported into one interpreter, in alphabetical order, before any test runs. A module's import-time side effects therefore land on every module sorting after it. The only run that tells you anything about the deploy is the one the deploy performs.

2. Two suites mutate sys.modules at import time, and the mutation is undetectable by inspection. test_approval_access.py:26 and test_approval_persist.py:16 define _stub(), which registers in-memory placeholders for google.cloud.*, cryptography.* and jose.* so those suites can run where the packages are absent. Each placeholder carries

m.__getattr__ = lambda name, _m=mod: type(f"{_m}.{name}", (), {...})

so it fabricates any attribute asked of it__file__ and __spec__ included. A placeholder cannot be told from the real library by looking at it; hasattr, __file__ and __spec__ all lie. The only reliable discriminator is the module's name.

3. The stub is guarded on if mod not in sys.modules, not on whether the package is installed. It therefore fires in Cloud Build too, where the whole requirements.txt is installed — replacing real, present libraries with fakes for every module imported afterwards. That is not a local-dev-only device, and reading it as one is the trap.

What to do when adding a test to this directory. If the new suite needs a real jose, cryptography or google.cloud — that is, if the library's actual behaviour is the subject rather than an obstacle — evict the placeholders by name at the top of the module, as test_auth_algorithm_pin.py does, and do not restore them afterwards (restoring fails: both libraries import submodules lazily, so a restored placeholder is what those late imports find). If it does not need them, nothing is required. Either way, run python3 -m unittest discover -p 'test_*.py' from forms-backend/ before pushing. That is the command that gates the deploy, and it is the only one whose result means anything.


VLN-042 — origins="*" on an unauthenticated backend reachable from any origin

Field Value
id VLN-042
title forms-backend/app.py:70 configures CORS(app, origins="*"). flask-cors reflects the caller's origin rather than sending a literal *, the service is deployed --allow-unauthenticated with ingress: all, and /health, /health/deep and /metrics carry no auth decorator. Any website a staff member visits can therefore read the portal's Prometheus metrics and its DB/KMS health from that person's browser.
register Security findings
status done — 2026-08-25. Narrowed in 158f932, deployed as revision gpus-forms-backend-00110-hvk. Both conditions observed on that revision.
severity MEDIUM (argued) — bounded exposure today (see the scope note below, which is as much of this row as the exposure claim), raised above LOW for the supports_credentials edge and because the wildcard has outlived the condition its own TODO made it conditional on. Argued, not supplied; no CVSS calculated.
owner R. Chhetry
evidence forms-backend/app.py:61-70, read 2026-08-25 — the CORS(app, origins="*") call and the TODO(phase-2-followup) above it. Deployment posture: forms-backend/cloudbuild.yaml:107 --allow-unauthenticated; ingress: all per VLN-018. Unauthenticated endpoints: forms-backend/routes/health.py:31-47/health, /health/deep and /metrics carry no @require_auth or @require_role, unlike /api/admin/reload at :51. Header behaviour measured, not assumed: the exact call was driven through Flask's test client on flask-cors 4.0.1 and 6.0.5, and a simple cross-origin GET /metrics carrying Origin: https://evil.example returns Access-Control-Allow-Origin: https://evil.example with Vary: Origin on both versions — the wildcard is origin-reflection, not a literal *. Metrics exposed: forms_decrypt_total, forms_happyfox_total, forms_up and submission counters (routes/health.py:20-28).
acceptance test CORRECTED 2026-08-25 — the original could not fail. See the note below. Two conditions, BOTH required. (a) The deployed SPA at forms.greenpeace.us still works end-to-end, observed in a browser — sign in, list forms, open a submission, submit. Necessary, and not sufficient: it guards against a config error that throws at init, and nothing more. (b) A cross-origin request from a NON-LISTED origin is refused. curl -H 'Origin: https://evil.example' <backend>/metrics must come back without an Access-Control-Allow-Origin header reflecting that origin. curl, not a browser — browsers will not generate this request, which is precisely why the condition needs a tool that will. (b) is the only condition that distinguishes a working narrowing from a no-op.
blocks
blocked_by — (was VLN-040 step 1; cleared 2026-08-25 when flask-cors 6.0.5 deployed. The ordering note below is kept: the reason it blocked was technical, not procedural.)
date_raised 2026-08-25
date_verified 2026-08-25

CLOSED 2026-08-25 — condition (a) observed on the narrowed revision

Condition (a) — the deployed SPA still works, observed in a browser. MET. All four named flows exercised against revision 00110-hvk, from the Cloud Run request log:

17:35:33  200  GET   /api/forms                              ← list forms
17:35:38  200  GET   /api/forms/53f29068c37e7067             ← open a form
17:35:46  201  POST  /api/submissions                        ← the write
17:35:47  200  POST  /api/submissions/8d9578b0-…/submit      ← submit
18:06:53  200  GET   /api/my-submissions
18:06:58  200  GET   /api/submissions/c3e322ab-…/status      ← open a submission

13 browser requests on this revision, 12× 200 and 1× 201, zero errors; the /status read rendered the full page — chain, three CANCELLED steps. Every non-2xx on the revision is a curl probe of ours. Condition (b) is evidenced by the verbatim live response headers recorded below.

Why the fourth flow was nearly missed, and it is worth one line: it was blocked by a different finding. The first pass exercised three of the four. "Open a submission" did not run because the only submission to hand was a Facilities Support Request — and a non-chain form is filtered out of /my-requests (approval_queue.py:46), so there was nothing in the list to click into. VLN-044 prevented VLN-042's own acceptance test from completing, and it did so silently: the flow simply had no route in, and a green sweep over the remaining three would have read as a complete pass.

Two findings interacting is how a test silently goes unrun. Neither row predicts it. Nothing in VLN-042 mentions /my-requests scoping, and nothing in VLN-044 mentions CORS. The only thing that caught it was reading the log for the specific request the acceptance test named, rather than accepting "all four flows passed" as reported.

NARROWED 2026-08-25 — condition (b) verified on the live service

158f932, build 11bb3cba, serving revision gpus-forms-backend-00110-hvk at digest sha256:5f23ab44…, matching the digest that build pushed. CORS(app, origins="*") is replaced by a flat three-entry allowlist: the forms.greenpeace.us domain mapping and the frontend's two Cloud Run URLs (hash form and project-number form — Cloud Run serves both and the domain mapping replaces neither).

Condition (b) — an unlisted origin is refused. MET 2026-08-25. Evidence is the live service's own response headers, captured verbatim against revision 00110-hvk. Not a test client: test_client performs no preflight, and the whole point of this condition is what the deployment actually returns.

$ curl -D - -o /dev/null -H 'Origin: https://evil.example' $BACKEND/metrics
HTTP/2 200

$ curl -D - -o /dev/null -H 'Origin: https://forms.greenpeace.us' $BACKEND/metrics
HTTP/2 200
access-control-allow-origin: https://forms.greenpeace.us
vary: Origin

$ curl -D - -o /dev/null $BACKEND/metrics          # NO Origin — the nginx-proxied path
HTTP/2 200
access-control-allow-origin: https://forms.greenpeace.us
vary: Origin

$ curl -X OPTIONS -H 'Origin: https://evil.example' \
       -H 'Access-Control-Request-Method: POST' \
       -H 'Access-Control-Request-Headers: authorization, content-type' \
       -D - -o /dev/null $BACKEND/api/submissions
HTTP/2 200

$ curl -X OPTIONS -H 'Origin: https://forms.greenpeace.us' \
       -H 'Access-Control-Request-Method: POST' \
       -H 'Access-Control-Request-Headers: authorization, content-type' \
       -D - -o /dev/null $BACKEND/api/submissions
HTTP/2 200
access-control-allow-origin: https://forms.greenpeace.us
access-control-allow-headers: authorization, content-type
access-control-allow-methods: DELETE, GET, HEAD, OPTIONS, PATCH, POST, PUT
vary: Origin

The unlisted cases return no access-control-allow-origin at all. On revision 00108-sfw, the same GET /metrics returned access-control-allow-origin: https://evil.example.

THE POSITIVE CASES ARE PART OF THE EVIDENCE, NOT DECORATION. A negative result on its own cannot tell "the narrowing worked" apart from "CORS broke entirely" — both produce a missing header. The listed origin still being reflected, with its full allow-headers / allow-methods / Vary set, is what makes the negative result mean what it is claimed to mean.

The no-Origin path — the one every real request takes — was re-checked deliberately. It changed from access-control-allow-origin: * to : https://forms.greenpeace.us (flask-cors emits the first listed origin when no Origin is present). Non-browser callers ignore the header entirely, and the response is still 200. Confirmed end-to-end through the real proxy rather than only against the backend's own hostname:

$ curl -H 'Authorization: Bearer not.a.jwt' https://forms.greenpeace.us/api/my-submissions
HTTP 401  {"detail":"Malformed token header: ...","error":"invalid_token"}
$ curl https://forms.greenpeace.us/            HTTP 200   # the SPA shell

A 401 from the backend's own error contract — rather than a 502/504 — is the success signal here: it proves nginx reached the backend and the backend answered. /api/forms (legacy auth.py) and /api/me return the same. /health and /health/deep are 200.

Condition (a) — the SPA still works, observed in a browser — is NOT yet met. No browser has touched this revision. Per this row's own corrected acceptance test, (a) is necessary and insufficient, and (b) alone does not close the row either. The row stays open until both are observed.

Flat list, not resources= — and the code says so at length

forms-backend/app.py carries the reasoning inline, because the dangerous change here is a plausible-looking improvement rather than a mistake:

  • resources={r"/api/*": {...}} would reopen three CVEs. CVE-2024-6866, -6839 and -6844 are path-matching defects that need per-path config to have anything to mismatch between. A flat list has one policy for every path and cannot exercise that class at all. The tempting next step — exempting /health and /metrics from the policy — is exactly the change that would do it, and they need no exemption: they are unauthenticated and answer any caller regardless.
  • supports_credentials=True must not be added. flask-cors will not pair a literal * with credentials, but it will pair a reflected origin with them — this configuration plus credentials, on an API that returns decrypted submission fields from /approval-view.
  • This is not a load-bearing auth control. @require_auth / @require_role are the gate. CORS stops another origin reading a response and stops nothing else; non-browser callers are unaffected by any value here.
  • No localhost entry and no env var. Dev proxies /api through Vite, so dev is same-origin too. An allowlist that can be widened by untracked config is not an allowlist — a developer pointing a dev server straight at a deployed backend adds their origin in a commit.

Scope — what this does NOT expose. An overstated row is as bad as a missed one.

No PII, no submission content, no token. The API is Bearer-only: the SPA attaches Authorization: Bearer <id_token> from JavaScript (forms-frontend/src/lib/api.ts:47). The token lives in sessionStorage (src/lib/auth.ts:45), which is scoped to the SPA's own origin, and no browser attaches it to a cross-origin request — a script on another site cannot read it and cannot cause it to be sent. supports_credentials is not set, so cookies are not in play either (and the portal uses no Flask session).

The exposure is exactly and only the unauthenticated endpoints, which would answer any caller anyway — curl reaches them today. What origin-reflection adds is that the read can be performed by a staff member's browser, from a page they merely visited, which curl from the open internet cannot do: it attributes the request to an internal-looking client and needs no attacker infrastructure. That is a real difference and it is the whole of the difference.

This row is a missing defence-in-depth layer, not an open door. Recorded at that weight deliberately.

Why the original acceptance test was wrong — correct about the library, wrong about the system

The test this row was raised with read: "the deployed SPA still works, observed in a browser." It could not fail. Any narrowing at all — including one with an empty origin list, or a typo'd domain — would have passed it.

The reasoning was drawn from the code and never checked against the deployment topology. From forms-frontend/src/lib/api.ts:47, every SPA request carries an Authorization header; a request with that header is not a CORS-simple request; therefore every request is preflighted; therefore a broken origin list breaks everything and a working SPA proves the list is right. Every step of that is true of the library, and the conclusion is false of this deployment, because preflight applies only to cross-origin requests. forms-frontend/nginx.conf reverse-proxies /api/ under forms.greenpeace.us, so the browser sees same-origin requests, sends no Origin header, and never asks permission. Narrowing the list cannot break something that never consults it.

Proved, not argued. Cloud Run request logs for the browser session that verified the flask-cors bump on 2026-08-25 (revision gpus-forms-backend-00108-sfw) show two API calls from a real browser and zero OPTIONS requests:

16:02:27  200  GET  /api/my-submissions              Mozilla/5.0 (Macintosh…)
16:02:31  200  GET  /api/submissions/…/status        Mozilla/5.0 (Macintosh…)

The only OPTIONS in the window were curl probes. nginx.conf's own comment says it plainly — the proxy exists "so that the frontend and backend appear at the same origin … and CORS becomes a non-issue."

Name the mistake, because it is a different one from the register's others. VLN-010, VLN-017 and VLN-040 are all controls that ran and were wired to nothing. This is not that. This is a correct inference about a component, applied to a system whose topology makes it irrelevant — code read accurately, deployment not read at all. It produced an acceptance test that would have been signed off, on evidence that was real but measured the wrong thing.

The generalisation worth carrying: an acceptance test must be checked for whether it can fail before it is checked for whether it passes. A condition no realistic defect could violate is not a weak test, it is not a test. Ask what would have to be true for it to go red; if the answer is "nothing that could plausibly happen", the condition is decoration.

The sharp edge — one flag from wildcard to origin-reflection WITH credentials

flask-cors will not send a literal * together with credentials. Given supports_credentials=True it echoes the request origin instead, which is the genuinely dangerous configuration: any origin, credentialed, on an API that returns decrypted submission fields from /api/submissions/<id>/approval-view.

Nothing in the repository sets that flag today. The finding is that one keyword argument separates the current posture from that one, on a line whose own comment (app.py:61-69) says

"Currently wildcard so local dev + the about-to-be-deployed Cloud Run frontend both work; tighten to the specific frontend origin in a follow-up commit once that URL is known" TODO(phase-2-followup): replace "*" with explicit origins=[...]

Both URLs have been known for months. forms.greenpeace.us is live and is recorded in inventory.yaml:1044. The condition the TODO made itself conditional on was met long ago and the TODO did not fire — which is why this is a register row and not a code comment.

Why this is NOT part of VLN-040

VLN-040 is dependency advisories — someone else's defects, inherited, dispositioned by version bumps and dated exceptions. This is the estate's own configuration. No scanner reports it, no bump fixes it, and no exception file could cover it. Same distinction that separated VLN-041 from VLN-040, applied to a config defect rather than a source one.

Ordering: narrow AFTER the flask-cors bump deploys — and the reason is technical

This looks like ordinary change hygiene. It is not.

Three of the four distinct advisories against flask-cors 4.0.1 — CVE-2024-6866 (case-insensitive path matching), CVE-2024-6839 (regex pattern priority), CVE-2024-6844 (unquote_plus turning + into a space) — are path-matching defects. Every one of them needs a per-path resources= configuration to have something to mismatch between. The portal has a single global policy and no per-path config, so none of the three is reachable today.

Narrowing origins is normally written as CORS(app, resources={r"/api/*": {"origins": [...]}}). That is precisely the construct that makes all three reachable. Narrowing on 4.0.1 would close this row by opening three others. The bump (e56c646, VLN-040 step 1) must deploy first.

If the narrowing is written without a resources= dict — a flat CORS(app, origins=[...]) — the three do not become reachable. Do not rely on that: the natural next step after narrowing origins is exempting /health and /metrics from the policy, which requires per-path config.


VLN-043 — The forms SPA reads none of the three variables its env template defines

Field Value
id VLN-043
title forms-frontend/.env.local.example defines VITE_OKTA_ISSUER, VITE_OKTA_CLIENT_ID and VITE_API_BASE_URL. The SPA reads none of them — two values are hardcoded in src/lib/auth.ts and the third is read under a different name (VITE_API_BASE). The template's VITE_OKTA_CLIENT_ID is also the pre-cutover shared-portal app id, wrong since 2026-08-05.
register Security findings
status open
severity LOW (argued) — nothing is currently misrouted, because the variables are inert. The severity is about what the template would do if the inertness were ever "fixed", and about an operator trusting a value that has no effect. Argued, not supplied.
owner unassigned
evidence Read 2026-08-25. forms-frontend/.env.local.example and the working .env.local both set VITE_OKTA_ISSUER=https://greenpeaceeu.okta.com, VITE_OKTA_CLIENT_ID=0oavvg1y33wTWFsmP417, VITE_API_BASE_URL=https://forms.greenpeace.us. Against that: src/lib/auth.ts:31-32 hardcodes const ISSUER = 'https://greenpeaceeu.okta.com' and const CLIENT_ID = '0oazce9jd5inTjtrn417', and src/lib/api.ts:30 reads import.meta.env.VITE_API_BASEnot VITE_API_BASE_URL. A repo-wide grep for import.meta.env returns exactly one hit, api.ts:30. auth.ts:14-18 already documents its own half: "The frontend Cloud Run service carries no env vars, and auth.ts does not read import.meta.env … (.env.local.example defines VITE_OKTA_CLIENT_ID but nothing reads it.)" — the VITE_API_BASE_URL / VITE_API_BASE name mismatch is not documented anywhere. The clientId in the template is the shared "GPUS Internal Portals" app, which auth_v2.py:35-39 records Forms as having cut over FROM on 2026-08-05 and which status-site, soc-site and mkdocs-portal still use.
acceptance test Either the template defines only variables the code actually reads (names matching, stale values corrected or removed), or it is deleted and the hardcoded values documented as the single source. A grep of forms-frontend/ finds no VITE_ name in an env file that is absent from src/, and none present in src/ that is absent from the template.
blocks
blocked_by
date_raised 2026-08-25
date_verified

Why a dead variable is worth a row

Same class as the env-template drift that would have rerouted every forms notification on a DR restore: a configuration value that looks authoritative, is read by nothing, and is wrong. All three failure surfaces are here at once.

  • It misleads. An operator changing VITE_API_BASE_URL to point at a staging backend observes no change and starts looking for the fault somewhere else. Nothing warns them; the variable is spelled plausibly and sits under a comment block that describes it as the backend API setting.
  • It is a loaded gun for whoever "fixes" it. The obvious tidy-up — wiring auth.ts to import.meta.env so the clientId is configurable, or renaming VITE_API_BASE to match the template — would immediately point Forms at 0oavvg1y33wTWFsmP417. The backend checks aud == OKTA_CLIENT_ID, deployed as 0oazce9jd5inTjtrn417 (cloudbuild.yaml:126), so every token would fail the audience check and every user would get a 401. The template does not merely do nothing; it holds a value that would break authentication for everyone the moment it started doing something.
  • It rots silently. The clientId went stale at the 2026-08-05 cutover and nothing noticed, because nothing reads it. A value nothing reads has no feedback path and cannot self-correct.

Not fixed in this pass — deliberately. The correct fix is a decision (wire the variables up, or delete the template and document the hardcoded values), not an edit, and choosing wrongly is how the second bullet happens.


VLN-044 — The confirmation page promised an approval chain that 28 of 29 forms do not have

Field Value
id VLN-044
title forms-frontend/src/routes/Submitted.tsx rendered a "Following this request" panel — "You can check which approval step it has reached at any time — and you will be emailed when it is approved, returned or declined" — on every submission, with links to /status/<id> and /my-requests. Only travel-request-001 has an approval chain. For every other form there is no step, no outcome email, and no row in the list the second button leads to.
register Security findings
status in progress — fixed in the working tree 2026-08-25 with a both-directions test; not deployed, not verified against live state
severity MEDIUM (argued) — no data exposure and no authz failure; this is an integrity-of-record defect. It is the portal telling its user, at the moment they decide what to do next, something specific and false about what will happen to their request. Argued, not supplied; no CVSS.
owner R. Chhetry
evidence Found by submitting, 2026-08-25 17:35 UTC. A real Facilities Support Request, 8d9578b0-046a-4604-a36d-2006c6e37a69, submitted through the production SPA during VLN-042's condition (a); the confirmation page rendered the panel. Cloud Run request log: 17:35:47 200 POST /api/submissions/8d9578b0-…/submit. The gate: Submitted.tsx:179 read {result.id && (id is the submission's own UUID, returned by every successful finalize, so the condition was a tautology. The field it was supposed to read did not exist: the comment directly above claimed "routing.approval is present only when finalize ran one", but contract.ts:246-249 declared routing as {email, happyfox} and routes_phase2.py:723-735 returned exactly those two. The chain gate: routes_phase2.py:691if submission.form_id == "travel-request-001":. The chain set: approval_resolver.py:66-72, FORM_CHAINS has exactly one key, extracted from source rather than read by eye. The list scope: approval_queue.py:46-48if sub.form_id not in chain_form_ids: return False, so a non-chain submission is silently absent from /my-requests. 31 form YAML files in the repo (27 in the live catalogue); 1 has a chain.
acceptance test cd forms-backend && python3 -m unittest test_submitted_panel_gate passes, and the deployed SPA shows the panel on a travel-request confirmation and no panel on a non-travel one, observed in a browser. BOTH DIRECTIONS ARE REQUIRED IN BOTH HALVES. Deleting the panel outright removes the false statement and would pass any test that only checks "absent for a facilities request" — so the travel direction is what makes this an acceptance test rather than a regression waiting to be signed off. Re-runnable: the suite is in the hard unit-tests gate (VLN-038's remediation), so it runs on every build.
blocks
blocked_by
date_raised 2026-08-25
date_verified

The guard was designed, documented, and implemented against a field nobody built

This is not a missing check. Submitted.tsx carried, verbatim:

/*
  Only for a submission that HAS an approval chain: linking a status page
  for a form with no chain would promise a progress view that has no
  progress to show. `routing.approval` is present only when finalize ran
  one.
*/}
{result.id && (

The comment states the correct rule. The code implements a tautology. And routing.approval existed on neither side of the wire. So the author reasoned it through, wrote the reasoning down, and then wrote a condition against a contract that was never built — on either the frontend or the backend.

That is worse than an absent guard, and the reason is worth stating. An ungated block invites the question "should this be conditional?". A block whose comment says it is conditional answers that question wrongly, to every future reader, including a reviewer looking specifically for this class of defect. The comment was evidence the case had been handled. It was the opposite.

Rule taken from it: a comment asserting a guard is a claim about the code and must be checked like one. Where the two disagree, the comment is not documentation — it is a false statement in the place people look to avoid reading the code.

Third instance of the same pattern — and the first found by USING the portal

UI or copy that describes a workflow the system does not perform:

Finding How it was found
1 GOV-021travel-request-001's own instructions column described an approval order the resolver does not perform ("…then to Finance"), served live to every submitter for the life of the chain Reading — comparing the YAML to STEP_DESCRIPTIONS
2 The WITHDRAWN fallback — SubmissionStatus.tsx has no branch for it, so the submitter reads "This request is recorded with this state. Contact IT if you need more detail." (recorded in schema/021_withdraw_test_chains.sql) Reading — while verifying an unrelated deploy
3 This row Submitting a real request

The first two were found by someone reading code against other code. This one required a human to fill in a form and look at the page — and it was found on the first non-travel submission anyone has made since the panel shipped. That is the finding about the finding: the estate's verification has been strong on reading and weak on using, and this class of defect is invisible to reading alone precisely because each artifact is internally coherent. Nothing in Submitted.tsx is wrong on its own terms; it is wrong only against a fact that lives in approval_resolver.py.

The unexercised remainder is named, not implied: the revise path (/revision-source, prefill, resubmit) has never been driven through a browser end to end. It is the only piece of what shipped to staff still in that state.

Population — 6 days, 4 submissions, 3 affected, ONE PERSON (revised twice)

Stated precisely rather than as a shape, because the shape and the blast radius are very different numbers here.

The panel shipped in 1daeea0, deployed by build 3ddb9338 at 2026-08-19 18:51 UTC. Every finalize since, from the Cloud Run request log:

08-25 18:15     —    /api/submissions/b5e998f7-…/submit    ← Change of Address (AFFECTED)
08-25 18:08:24  200  /api/submissions/f8f89f9e-…/submit    ← Facilities Support (AFFECTED)
08-25 17:35:47  200  /api/submissions/8d9578b0-…/submit    ← Facilities Support (AFFECTED)
08-20 16:41:24  200  /api/submissions/e5ff2b59-…/submit    ← travel; panel truthful

e5ff2b59 persisted a chain — travel_approval.persisted submission=e5ff2b59-… outcome=PERSISTED state=IN_APPROVAL current_step=3 — so the panel was correct for it and for it alone.

Shape: 28 of 29 forms would have shown a false statement. Realized: three submissions — and, importantly, ONE PERSON.

The blast radius is one person finding this three times, not three people misled. All three affected submissions are R. Chhetry's, in a single session on 2026-08-25, made while deliberately exercising the portal. No staff member outside IT has encountered this panel on a non-chain form. That distinction is the difference between a defect with a user population and a defect with a discovery narrative, and the row would misrepresent itself by reporting "3 affected" without it.

The third instance confirms the diagnosis rather than extending it. The gate is a tautology ({result.id && (), so the panel shows on every form by construction — Change of Address is not a new class of case, it is the same case a third time. No further instances are needed or should be collected; another submission would add a row to this table and nothing to the finding.

A SECOND COUNT WAS ALSO WRONG, in the other direction. This row first said "30 of 31 forms". That counted forms/*.yaml minus pulldowns.yaml, which swept in templates.yaml and _schema.yaml — neither is a form, and yaml_loader.py:174 does not load either. The live figure is 29 forms, of which 1 has a chain: 28 affected. Corrected 2026-08-26. It changes nothing about the finding and is fixed because a register that rounds its own denominator invites the next reader to round theirs.

Still carrying the old number: forms-frontend/src/lib/contract.ts, in the RoutingApproval doc comment ("30 of 31 forms"). Comment only, no behaviour, deliberately not corrected tonight — it would trigger a frontend deploy for a comment. Fix it with the next change to that file.

THE COUNT WAS EVIDENCE AND IT CHANGED — TWICE. Recorded as revisions rather than edits. This note first read "1 affected", written at 17:50 when that was true; revised to 2 at 18:08; revised to 3 at 18:15. The population of a live defect is a measurement with a timestamp, not a property of the finding: it is correct only as of when it was taken and it grows until the fix deploys. Anything quoting an earlier figure is quoting an earlier clock, which is why each revision is dated in place rather than overwritten.

The fix — the backend reports what happened; the SPA is told, not required to know

routing.approval now exists, which is what the original comment assumed. It carries the outcome for this submission, not the form's configuration:

status meaning panel?
persisted a round exists with steps to follow and an approver to email yes
none this form has no chain — 28 of 29 forms no
blocked a chain was configured but resolved to no steps no
failed persistence raised; the submission is finalized and routed, the chain is not no

Outcome, not configuration, and that is the load-bearing choice. A form lookup would answer "travel-request-001 has a chain" for a submission whose chain failed to persist — and _run_travel_approval swallows that failure by design, because a submitted request must not 500 over an approval problem. The return value is the only place that fact can reach the page. So blocked and failed are travel-request outcomes that leave nothing to follow, and the UI treats them exactly like none.

Absent means false. A cached SPA bundle talking to an older backend gets no panel rather than the old always-on one. Guessing wrong in that direction costs a travel submitter one link that the /my-requests nav item already provides; guessing wrong in the other direction is this row.

No hedge replaces it. For a form with no chain there is genuinely nothing to follow, and the honest render of that is silence. "You may be contacted if anything is needed" would be the same defect with better manners. The headline prose already states what did happen — logged, and queued for delivery to the relevant team.

No form id in the frontend. The backend owns FORM_CHAINS; a hardcoded travel-request-001 in the SPA would be a second copy that silently stops matching the day a second chain is configured. Asserted by a test, against comment-stripped source — the code comment does name the form id when explaining where the backend gate lives, and a naive grep would have fired on the explanation.

Mutation-proved in BOTH directions, on mutants that compile

19 tests in test_submitted_panel_gate.py; the forms-backend suite is now 301 tests across 10 suites, green under unittest discover.

Mutant Compiles? Caught by
Gate back to {result.id && (, unused const removed yes, tsc exit 0 6 tests — …is_not_gated_on_something_always_true, …gate_reads_the_approval_outcome, …absent_status_is_treated_as_no_chain, and 3 more
Panel deleted outright ("fixed" by removal) yes test_the_panel_still_exists_for_a_chain_bearing_submission

The two mutants are caught by disjoint tests, which is the property that makes this bidirectional rather than merely thorough. The first mutant was initially written without removing the now-unused hasApprovalChain, which made tsc fail for an incidental reason — re-run with the const removed so the mutant genuinely compiles and the test is demonstrably what catches it.


VLN-045 — Every submission's status page is headed "Travel request."

Field Value
id VLN-045
title forms-frontend/src/routes/SubmissionStatus.tsx:90 renders a hardcoded <h1>Travel request.</h1>. The page serves any submission the caller owns, so a Facilities Support Request's own status page is titled with a different form's name. Two sibling views carry the same hardcoded string and are correct only by accident of scoping.
register Security findings
status in progress — fixed 2026-08-26 in ee2da8a; not yet verified in a browser
severity LOW (argued) — no disclosure and no authz failure. It is the fourth instance of the estate's recurring pattern: an artifact that is internally coherent and false against a fact living elsewhere. Argued, not supplied.
owner unassigned
evidence Read 2026-08-25 while investigating a report of "the approver in the other forms". SubmissionStatus.tsx:90<h1 style={{ fontSize: 32, marginBottom: 6 }}>Travel request.</h1>, outside every conditional. The route is /status/<submission_id> and its endpoint authorizes per submission with no form filter (routes_phase2.py:1223, resolve_submitter_access), so it renders for any owned submission. Two siblings carry the same literal and are currently correct: MyRequests.tsx:148 (the list is scoped to chain-bearing forms, approval_queue.py:46) and ApprovalDetail.tsx:117 (the approver queue is chain-only). Both become wrong the day a second chain is configured — they are latent, not safe.
acceptance test /status/<id> for a non-travel submission shows that form's own name, observed in a browser; and /status/<id> for a travel request still shows "Travel request." Both directions, for the same reason as VLN-044: replacing the heading with a generic "Request." would fix the false statement and pass a one-directional test while making the travel page worse. The form name is already available — GET /api/submissions/<id>/status returns form_id, and the SPA has the catalogue from /api/forms.
blocks
blocked_by
date_raised 2026-08-25
date_verified

What this row is NOT — the investigation that produced it

Raised out of a report that non-travel forms were "showing the approver", which would have been a far larger finding: chain data rendered for a submission that has none, or read from another submission. Neither is happening. Three checks, all negative:

  • No non-travel form claims an approval chain. All 31 form YAMLs were grepped for approver / approval / SMT / sign-off / supervisor / manager wording. Seven HR forms carry fields asking for a manager or timesheet approver (Manager, TimeApproval, CurrentManager — "Current time card approver"). Those are input questions the submitter answers, not statements about routing, and they are correct. travel-request.yaml is the only form whose text describes an approval chain.
  • Nothing approver-shaped renders for a chainless submission. SubmissionStatus.tsx:98 gates the whole chain block on {approval.steps.length > 0 && <Chain steps={approval.steps} />}, and the endpoint returns steps: [] when there is no round. total_steps is returned as the travel constant 3 even with no chain, but it is read only inside the STATE_IN_APPROVAL branch, which state === null cannot reach. stateLabel(null) correctly returns "Submitted". No approver name, no step, no roster reaches a chainless page.
  • No path reads another submission's chain. Every chain read is keyed on submission_id: load_current_round and load_steps (approval_access.py:376,414) both filter on it, load_submission_ref is a primary-key get, and _decrypt_submission_fields (routes_phase2.py:852) filters SubmissionField.submission_id. The one query that is not per-submission is load_current_rounds() (approval_queue.py:317), which loads every round to build the queue — and it attributes steps to submissions by a (submission_id, version) dict key, never positionally. A chainless submission has no row in submission_approvals, so it cannot appear in that result set at all, with or without steps.

What the reporter most likely saw is either the HR forms' legitimate manager fields, or the Approvals queue they opened at 17:36:55 — which shows approvers because it is the approver queue and is chain-only. The one thing genuinely wrong on a non-travel page is this heading.

FIXED 2026-08-26 — the name is data, in all three views

ee2da8a. QueueSubmission carries form_title (one batch lookup, not one per row); both row builders emit it; /status and approval-view emit it. A shared submissionHeading() renders "<name>." and falls back to "Request."deliberately generic, because a form name as the default is the defect whichever form it names.

Two explicit key-allowlist tests fired on the new row key and were extended deliberately rather than loosened. That is exactly what they exist for: TestQueueRowShape.ROW_KEYS and TestMySubmissionsRowShape.ROW_KEYS pin the row's ENTIRE key set so that no key can be added by accident, and form_title is not field content — it is the public name of a form whose id the caller already holds.

And a test fake was wrong in a way worth recording. _FakeSession.get ignored its model argument and returned the submission for every lookup. Harmless while the handler fetched one model; approval-view now also reads the Form row, so the fake handed back a Submission, .name raised, and eight unrelated tests failed at once as a 500. A stub that is loose about what it is asked for does not fail where the looseness is — it fails somewhere else, later, in bulk.

Fourth instance of the pattern, and the one that dates it

UI or copy describing a workflow the system does not perform: GOV-021's served instructions text, the WITHDRAWN fallback, VLN-044's panel, and this. All four share a shape — an artifact that is correct on its own terms and false against a fact that lives in another file. Nothing about <h1>Travel request.</h1> is wrong in SubmissionStatus.tsx; it is wrong only because resolve_submitter_access does not filter by form.

This one dates the pattern rather than just repeating it. The heading was written when /status served travel requests and nothing else, and it was true then. It became false when the page began serving every form — a change made elsewhere, which no test and no reviewer connected back to a literal string three files away. The defect was introduced by a change that did not touch it.

What a test pinning the copy would NOT have caught — and what would

Asked for explicitly, because this row is a different class from its three siblings and the difference decides what to build.

GOV-021, the WITHDRAWN copy and VLN-044 were all wrong when written. A test asserting the copy at authoring time — "the panel says X", "the instructions say Y" — would have been written against the same wrong understanding and would have passed, then defended the error. Snapshot tests on copy are worth little against them.

This heading was TRUE when written. /status served travel requests and nothing else. It was made false by a change three files away that did not touch it: the page began serving every form the caller owns, because resolve_submitter_access is per-submission and has no form filter. A snapshot test would have passed before the change and passed after it, since the literal never moved.

What catches this class is asserting the relationship, not the value:

  • "No view hardcodes a form name" — a structural assertion over source, which fails the moment a literal is reintroduced anywhere, and which cannot be satisfied by a literal that happens to be right today.
  • "The heading is a function of the payload" — the view must read form_title from its own data. A view that renders a constant fails whatever the constant says.
  • Both directions, so replacing one literal with a generic literal fails too — which is mutant M2 below.

That is the general form: when a view's correctness depends on a fact owned by another module, pin the dependency, not the output. A value can be right for the wrong reason and stay right until the reason changes.


VLN-046 — "Ticket: NOT CREATED" asserts a negative the portal cannot verify

Field Value
id VLN-046
title Submitted.tsx:47 renders an amber Not created pill against a Ticket row whenever a form declares any happyfox_template action and no ticket id came back. It states as fact that no ticket exists. The portal cannot know that: the same submission's email_template legs deliver to gpus-*@greenpeace.org addresses, at least one of which is proven to auto-open a HappyFox ticket by email ingest.
register Security findings
status in progress — copy fixed 2026-08-26 in ee2da8a; not yet verified in a browser
severity LOW (argued) — no disclosure, no data loss. It is the inverse of VLN-044: that one is reassuring and false, this one is alarming and unverifiable. Argued, not supplied.
owner unassigned
evidence Observed 2026-08-25 on the Change of Address Notification confirmation (b5e998f7), and absent from both Facilities Support confirmations. The trigger is form shape, not submission state: Submitted.tsx:47const ticketNotCreated = happyfoxCount > 0 && !result.happyfox_ticket_id; where happyfoxCount is routing.happyfox.count, computed in routes_phase2.py by counting the form's happyfox_template actions. change-of-address-notification.yaml declares 2 (destination: 92, destination: 96); facilities-support-request.yaml declares 0 — which is the whole of why one showed the row and the other did not. No ticket is ever created by the API path: gpus-forms-routing-worker/routing_worker.py:683-694 writes status='deferred' for every happyfox_template action and dispatches nothing, pending 2.5(e). But the email path is a different story, and it is documented: architecture/forms-phase2.5c-design.md:157 records the G4.4 finding of 2026-06-12 — "some email_template destinations are HappyFox email-ingest addresses — the Gate-4 close-out email to gpus-it-support@greenpeace.org auto-opened ticket #USITS00357259 via email-to-ticket. The worker dispatched no HappyFox action (verified)." Change of Address's two email legs go to gpus-finance-support@ and gpus-people@, from the same seven-address gpus-*@greenpeace.org family.
acceptance test For each of the seven email destinations, it is recorded whether that inbox opens a HappyFox ticket on receipt — the T3 per-form deconfliction answer. Then the confirmation page says only what follows from it. A page that cannot determine ticket state says nothing about tickets; it must not assert a negative. Re-runnable: a submission on a form whose email leg IS an ingest inbox must not render "Not created".
blocks
blocked_by CLEARED 2026-08-26. The row was blocked on T3 while the copy tried to report the OUTCOME (does a ticket exist?), which needs a per-mailbox answer. It is unblocked by no longer making that claim: the copy now describes the MECHANISM and the portal's own blindness, both of which are knowable today. T3 still gates 2.5(e) and still gates ever showing a ticket NUMBER.
date_raised 2026-08-26
date_verified

CONFIRMED, then FIXED — 2026-08-26

The finding is no longer inferential. Ticket #USFIN00367708 opened in US - Finance Support for change-of-address submission b5e998f7, delivered by email, Requestor: Alerts - alerts@greenpeace.us. A ticket existed while the confirmation page said one had not been created. The routing is correct — the body concerns reimbursement addresses and Finance is the right queue. Only the copy was wrong.

Note the requestor: alerts@greenpeace.us, the relay-rewritten sender (the send-as decision). The ticket is attributed to the alerts mailbox rather than to the submitter — a separate consequence of that decision, not tracked here, and worth knowing before anyone tries to correlate tickets to people.

Fixed in ee2da8a. The row now reads:

TicketNot shown here · the team's system assigns it when the email arrives

It states what the portal did and what it cannot see, and asserts nothing about whether a ticket exists — the portal cannot observe that either way. Neutral, not amber: the warning colour did as much of the work as the word "not", and there is no fault here to warn about.

Also gated on an email leg existing, which is not cosmetic. Of the 12 live forms declaring a happyfox_template action, new-employee-notification-contractor-intern declares no email action at all — nothing is dispatched and nothing is emailed, so no mailbox can open anything, and wording about email ingest would be false for exactly that form. The headline prose already tells that submitter their request has not been sent to a team, which is the honest place for it.

Not overcorrected. "A ticket has been created" would have been the mirror-image defect: the portal cannot see that either. The copy describes the mechanism and its own blindness, both of which are knowable today, and neither of which is a claim about this particular ticket.

The assertion the portal is not entitled to make

Three separate things are true at once, and the row collapses them into one false sentence:

Fact Confidence
No ticket was created by API dispatch Certain. The worker defers every happyfox_template and dispatches nothing.
A ticket may have been created by email ingest Unknown per address, and demonstrated for at least one. gpus-it-support@ provably opens tickets (#USITS00357259).
The email was handed off for delivery True, and it is what the small print beside the pill already says.

"Not created" reports the first as though it were the whole picture. The honest statement is "the portal cannot tell you the ticket number" — which is a statement about the portal, not about the ticket.

This is not pedantry about wording. The submitter's decision is whether to chase. Told a ticket was not created, the reasonable action is to contact IT and open one — and if the ingest path did its job, they have now caused a duplicate for the queue to close, which is precisely the double-ticket failure T3 exists to prevent. The copy can generate the outcome it is warning about.

RESIDUAL — 17 forms are told nothing about a ticket that probably exists

Recorded because the 2026-08-26 fix corrected the copy and did not change the rule that decides who sees it. Anyone touching that gate next should know what it actually keys on.

The gate is happyfoxCount > 0 && !happyfox_ticket_id && hasEmailLeg. In words: this form declares a HappyFox action. It is NOT a ticket can result from this submission. Those are different questions, and the second one is the one a submitter is asking.

Counted from the 29 live form definitions on 2026-08-26 (forms/*.yaml less pulldowns.yaml, templates.yaml and _schema.yaml, which yaml_loader.py:174 does not load):

forms ticket row a ticket can result?
HF action + email leg 11 shown yes — by ingest
HF action, no email leg 1 hidden (correctly) no — nothing is sent
email leg, no HF action 17 hidden yes — by ingest, identically
29

The 17 are the residual. Their email goes to the same gpus-*@greenpeace.org family — gpus-facilities@, gpus-it-accounts@, gpus-data-request@ and the rest — and ingest is what opens a ticket, not the happyfox_template action, which dispatches nothing at all. A Facilities Support Request very likely opens a ticket by exactly the mechanism that opened #USFIN00367708, and its submitter is told nothing about tickets whatsoever.

This is the same asymmetry as this row, pointing the other way. The original defect was a form being told something false about a ticket. This is a form being told nothing about one that probably exists. Both come from the same root: the row is gated on a declared action rather than on whether a ticket can result.

Not a new finding and not wrong. Silence is not a false statement, and saying nothing is strictly better than the "Not created" it replaced. It is logged here rather than raised because the fix is the same fix — and it is the same blocked answer: which of the seven destination mailboxes are HappyFox-watched. Once T3 answers that, the honest gate is "this submission's email goes to a watched mailbox", which would show the row on the 17 and could remove it from any of the 11 whose mailbox is not watched.

Do not tighten the gate before T3. Showing the row on all 28 email-bearing forms would assert the ingest mechanism for mailboxes nobody has confirmed are watched — trading a silence for a guess.

Coverage is inverted — the form most likely to have opened a ticket says nothing

Worth stating plainly because it is the opposite of what the row's presence implies.

  • Change of Address — 2 email legs + 2 happyfox_template legs → shows "Not created". It has the higher chance of a ticket existing, because it has two email legs into the gpus-* family.
  • Facilities Support — 1 email leg, 0 happyfox_template legs → shows nothing about tickets at all. Its email to gpus-facilities@greenpeace.org may equally have opened one, and the submitter is told nothing either way.

The row keys off the unbuilt API path while the thing that actually creates tickets today — email ingest — is invisible to it. A submitter reading both confirmations would conclude the facilities request is the cleaner outcome. Nothing supports that.

Why its OWN row and not part of VLN-044

Both are on the same page, in the same component, found in the same session. They are still different findings.

  • VLN-044 is a promise about a workflow that does not exist, gated by a tautology. Fix: gate it on a fact the backend already has. Owned entirely within this repo; shipped the same day it was found.
  • VLN-046 is an assertion about a downstream system's state that this repo cannot determine from anything it holds. No gate fixes it. The fix requires an answer from outside the code — which of seven mailboxes are HappyFox-watched — and that answer is T3, which also blocks 2.5(e).

Different fix, different acceptance test, different blocking dependency, and one closes while the other waits on a mailbox audit. Folding this into VLN-044 would have closed it by association when the panel shipped — exactly the reason VLN-041 was separated from VLN-040.

The T3 block is cleared, and how is the point. This row was blocked_by T3 while the copy tried to report the outcome — does a ticket exist? — which genuinely needs a per-mailbox answer nobody has. It is unblocked not by getting that answer but by no longer making that claim. The mechanism ("the team's system assigns it when the email arrives") and the portal's blindness ("not shown here") are both knowable today.

T3 still gates 2.5(e), and still gates ever showing a ticket number. What it does not gate is telling the truth about what the portal can see — which was available the whole time.


VLN-047 — The revise path shipped complete and unreachable, and the page it should start from says it does not exist

Field Value
id VLN-047
title Migration 022's revise path — /revision-source, prefill, parent-linked submit, chain rendering — is deployed and working, and no UI anywhere constructs the URL that starts it. /status/<id>, the one page a returned request lands on, asserts in bold that the request "cannot be edited" and offers a button that produces an unlinked resubmission. Two defects: an absent entry point, and a surface that actively contradicts the feature while silently generating the wrong data shape.
register Security findings
status done — CLOSED 2026-08-26. Both halves of the acceptance test met and observed in the database; see the closure note. Frontend revision gpus-forms-frontend-00021-dlw, worker bc0c84b, both content-verified
severity MEDIUM (argued) — no disclosure and no authz failure; resolve_revision_access is sound and was never bypassed. Argued up from LOW on a ground the copy findings lack: a blocked staff member who reported it rather than working around it — a live, staff-facing feature was unreachable to every user for a week while the page a returned traveller lands on denied it existed. Stated precisely: no real trip is known to have been blocked. The submission that surfaced it, 81c74ae9-…, was returned by IT with the comment "I am returning this as a test"; what was real was the user, who could not find the control and asked. The defect is the unreachability, which was total; the discovery was a test return that a real user then walked into. A second ground, "one orphan resubmission already in submissions", was asserted here on 2026-08-26 and WITHDRAWN the same day — see the population note; it was an inference from timing that the field data contradicts. Argued, not supplied.
owner unassigned
evidence Raised 2026-08-26 from Tanu Garg's report on 81c74ae9-22b0-451e-be8c-6123b39c0a18 (RETURNED_FOR_REVISION, returned at step 2, 14:33:24 UTC), verbatim: "how do I edit the same response? I have an option to submit a new request, but do not see re-submit the request." Both halves of that sentence are the two defects. (a) grep -rn "revise" forms-frontend/src returns seven hits: six consume the parameter (FormFill.tsx:29,48,67,69,85,90), one is a stale comment (SubmissionStatus.tsx:30), zero construct it. No template literal, no searchParams.set, no navigate() carrying it. The URL is reachable only by typing it with a UUID already in hand. FormFill.tsx:29 documents a caller that was never written: "?revise=<parent_id> is how the SPA carries the parent, set by the link on /my-requests." There is no such link — MyRequests.tsx contains no RETURNED branch at all, and its whole card is one <Link to={/status/…}>. (b) SubmissionStatus.tsx:152-154 renders <strong>This request cannot be edited and will not move any further.</strong> followed by the only control on the panel, <Link to={TRAVEL_FORM_PATH}>'/forms/travel-request', bare, no ?revise=. Commit 65350f8 ("wire the revise path", 2026-08-25) touched four frontend files — api.ts, contract.ts, FormFill.tsx, MyRequests.tsx. SubmissionStatus.tsx is not in its file list. The +45 lines to MyRequests.tsx are ChainLine and its call site: rendering of an existing chain, not creation of one.
acceptance test Two halves, both required. (1) A submitter reaches the revise path from the UI alone — from /status/<id> on a returned request, without typing a URL — and the prefilled form opens carrying the returning approver's comment. (2) The submission that results carries a non-NULL parent_submission_id, observed in the database. The second half is what distinguishes a fix from a link that produces orphans: a revise button pointing at the bare form would pass half 1 and reproduce the exact defect the row exists for. Half 2 additionally proves the chain restarts at step 1 and ChainLine renders on both ends.
blocks
blocked_by
date_raised 2026-08-26
date_verified 2026-08-26

CLOSED 2026-08-26 — the column was written by the real path, by the person who reported it

Both halves, observed in gpus_forms rather than inferred from a green deploy. The verification was Tanu Garg's, which is correct: she found it.

17:08:58  submission_viewed + decrypt_success   on 81c74ae9…  actor tgarg
17:09:03  submission_viewed + decrypt_success   on 81c74ae9…  actor tgarg
17:09:33  35750caf-80b7-4879-adf7-ecf174ce1bcf submitted
          parent_submission_id = 81c74ae9-22b0-451e-be8c-6123b39c0a18

Half 2 — the parent is recorded. SELECT count(*) FROM submissions WHERE parent_submission_id IS NOT NULL went from 0 to 1. That column has existed since migration 022 on 2026-08-25 and had never been written by any path; this is the first time, and it is what separates a fix from a link that produces orphans.

Half 1 — reached from the UI. The audit pair at 17:08:58 and 17:09:03 is GET /revision-source, the portal's fourth decrypt path, which is called only when ?revise=<parent_id> is present — and /status/<id> is the only thing in the SPA that constructs that URL. Stated precisely: the audit log proves the endpoint was exercised, not which control produced the URL, since a hand-typed URL is indistinguishable here. Typing it is not a plausible reading — she had reported two and a half hours earlier that she could not find any way to do this — but the distinction is recorded rather than glossed.

The chain restarted and completed. The revision's steps are 1 SKIPPED_DUPLICATE (Kevin Toruno) · 2 APPROVED (Rajesh Chhetry) · 3 APPROVED (Kevin Toruno), round state APPROVED. A fresh round from step 1, exactly as designed — no special case, because a revision routes, finalizes and dispatches through the same code as any other submission.

What this does not close. VLN-049: the form's own instructions went on telling every submitter a returned request "cannot be edited" for a further day after this row closed. The control existed and worked; the page introducing the form still said it did not.

Population — ZERO. This note asserted one orphan on 2026-08-26 and was WRONG.

Corrected 2026-08-26, hours after it was written, before anyone acted on it. The retraction is kept in place rather than edited away, because the row's whole subject is a claim that outlived its evidence.

What the row originally said. That 1b4dfd50-fe3c-4cfc-b3e0-c8dada16c713 was an orphan resubmission of 6418f160-…: Tanu Garg was returned, followed the "Submit a new request" button, and filed a request the database could not connect to the one it replaced. The reasoning was a timeline:

08-19 16:29:00  6418f160-…  tgarg  submitted
08-19 16:30:49  6418f160-…         RETURNED at step 1
08-19 17:44:55  1b4dfd50-…  tgarg  submitted   ← 74 min later, parent NULL
08-19 17:46:16  1b4dfd50-…         step 1 APPROVED

Same person, same form, 74 minutes after a return. It reads as obvious.

The field data says otherwise, and it was never checked. The return comment on 6418f160-… step 1 asks for exactly one thing:

Lodging is too much

6418f160-… (returned) 1b4dfd50-… (the supposed revision)
Lodging 1000 1000 — unchanged
Airfare 400 150
Depart 2026-10-10 2026-09-10
Return 2026-12-10 2026-10-20

The one field the approver asked to change is identical. Every field they did not mention changed. A revision answering "lodging is too much" that leaves lodging untouched and moves both dates by a month is not a revision. These are two different trips.

So the population is ZERO — no orphan, and no known instance of anyone following the button into an unlinked resubmission. All 11 travel-request-001 rows have parent_submission_id IS NULL, and SELECT count(*) … WHERE parent_submission_id IS NOT NULL returns 0 across the whole table. That is fully explained by "no revision has ever been created" — the entry point's absence stated as data — and needs no second explanation.

R5's August trip count is therefore NOT off by one. 1b4dfd50-… is its own trip and counting it as one is correct.

What this cost, and why it is recorded here rather than quietly fixed. The claim shipped in two commit messages (01957b4, bc0c84b), a test docstring, and this row's severity argument, which cited "a realized data defect" as a ground for MEDIUM. It was an inference from timing, written with the confidence of a query result, in a row whose own thesis is that a well-argued statement defends a stale conclusion better than a bare one. The register reproduced the defect it was describing, in the paragraph describing it. The evidence that settles it — one comment column and four searchable_values keys — was one query away the entire time and was not run, because the timeline already looked like an answer.

Severity is unchanged at MEDIUM, re-argued. The ground that survives is the realized one: a live, staff-facing feature was unreachable for a week and the page a returned traveller lands on actively denied it existed, with a confirmed blocked user who reported it. The orphan ground is withdrawn.

Nothing to link retroactively. The Part-4 data decision this row was carrying — whether to backfill parent_submission_id on 1b4dfd50-… — is moot, and backfilling it would now be the error: it would assert a supersession between two unrelated trips, inside a live approval record currently sitting at step 2. Closed as "no action, premise false" rather than as "decided against".

TWO defects in one row, and why they are not the same defect

Distinguished because the fixes are different and half a fix is worse than none.

  • (a) No entry point. The feature is complete and has no front door. Everything from /revision-source through parent-linked submit to ChainLine shipped and works; nothing calls it. Fix: construct the URL somewhere a submitter will find it.
  • (b) An active contradiction. /status/<id> does not merely omit the control — it denies the control exists, in bold, and offers a button that produces an orphan. Fix: delete the false claim and remove the orphan-producing button.

(b) is the more serious half and it is not a copy defect. A page that silently says nothing about revising leaves the submitter stuck. A page that says "cannot be edited" and hands them a button routes them into producing the wrong data shape — which is what happened on 08-19. The false sentence is not the harm; the button beside it is.

Which is why the fix must not leave both controls standing. Two buttons, one of which produces an orphan, is worse than the current state, because today's single wrong button is at least unambiguous.

The mechanism — TRUE when written, made false by a change that never touched it

This is the part worth the row, and it is the same class as VLN-045 rather than a repeat of the copy findings.

SubmissionStatus.tsx's comment was correct on the day it was written, 1daeea0, 2026-08-20:

 * READ ONLY. There is no edit and no resubmit-in-place, because the revision
 * path is blocked on an unsettled schema question about how a revised
 * submission chains to the original (ASVS §2.4). The page therefore never
 * offers a control that does not exist — on a returned request it says a NEW
 * request is required and links the form, which is the same thing the outcome
 * email says, so the two cannot tell the traveller different stories.

Every clause was true. ASVS §2.4 did record the revise endpoint as "NOT BUILT, PROVISIONAL". The schema question was genuinely unsettled. And the reasoning is good — it refuses to name a control that does not exist, precisely so as not to send a traveller hunting for a missing button. The inline comment at the panel says it again:

      NO EDIT PATH EXISTS. Saying "update your request" would send the
      traveller looking for a button that is not there, and someone who
      believes they have edited a request that never changed is worse off
      than someone told plainly to start again. Same wording as the outcome
      email, deliberately.

Migration 022 settled the schema question on 2026-08-25 — Option A, a revision is a new row with parent_submission_id. 65350f8 built everything on top of it the same day. The refusal outlived its reason by one day and now says the opposite of the truth in the same words.

The care is what makes it dangerous. Both comments explain why the control is absent so convincingly that a reader arriving later — including the author of 65350f8 — is told, in the file, that this page is correct as it stands. A well-argued comment defends a stale conclusion better than a bare one does.

Three rows in this family now split two ways

Asked for explicitly, because the split decides what test is worth writing.

Wrong when written — GOV-021's served instructions, the WITHDRAWN fallback, and VLN-044's panel. Each was authored against a mistaken understanding of the system. A copy snapshot test written at authoring time would have been written against the same mistake, passed, and then defended the error.

True when written, made false elsewhere — VLN-045 and this row. VLN-045's <h1>Travel request.</h1> was true while /status served travel requests only, and was falsified by resolve_submitter_access gaining no form filter three files away. This row's "cannot be edited" was true while no revise path existed, and was falsified by migration 022 in a different directory. In both, a snapshot test passes before the change and passes after it, because the literal never moves.

The difference between the two halves of the family: for the first, the fix is to correct a statement. For the second, the fix is to pin the relationship rather than the value — assert that the view's claim is a function of the fact it depends on, so it cannot survive the fact changing. That is what the structural test below does, and it is the only test here that would have caught this before Tanu Garg did.

Register mechanics — VLN-044 predicted this exactly, one day early, and nothing acted on it

Recorded as a finding about the register itself, not as a footnote.

VLN-044 already contained this, written 2026-08-26, the day before the report:

The unexercised remainder is named, not implied: the revise path (/revision-source, prefill, resubmit) has never been driven through a browser end to end. It is the only piece of what shipped to staff still in that state.

It named the exact path, the exact reason, and the exact risk — three lines below VLN-044's own lesson that "the estate's verification has been strong on reading and weak on using." It was right. It changed nothing.

Because it was prose inside another row. It had no id, so nothing could reference it; no owner, so nobody held it; no acceptance test, so no build and no review could fail on it; and no status, so it could not be open. It could not appear in the summary table, could not be triaged, and closed silently when VLN-044 closed — which is the failure mode VLN-046's own "Why its OWN row" note warns about, reached from the other direction.

A prediction without an id is a note. The register's rule should be read as: if a sentence in a row identifies unexercised behaviour in shipped code, it is a row. Naming a risk inside another finding feels like diligence and produces none of a row's mechanics.

This is also the second time the estate has been told and not heard: ASVS §2.4 recorded the endpoint as NOT BUILT, PROVISIONAL, and that line went stale on 2026-08-25 without anyone revisiting the surfaces that cite it.

Why tsc and eslint both passed, and what does catch it

65350f8's verification line reads "tsc and eslint clean on the four frontend files touched." True, and structurally incapable of finding this.

A query parameter that is consumed but never produced is perfectly typed. searchParams.get('revise') returns string | null whether or not anything in the universe ever sets it. null is a legal, handled value — FormFill branches on it correctly and renders a normal blank form. There is no unused symbol for eslint, no type error for tsc, no dead code: every line of the revise path is reachable, just never reached. The feature's absence has the same type signature as its presence.

Compounding it: forms-frontend has no test runner at all. No vitest, no jest, zero test files, and forms-frontend/cloudbuild.yaml invokes no test step. check-test-gates.sh cannot see this, because it enumerates directories containing test_*.py — a component with no Python tests is invisible to the gate that exists to find ungated components. VLN-038 closed the backend half; the SPA was never in scope.

What catches it is an assertion over the relationship between two files: every query parameter read by a route must be constructed somewhere in the SPA. It is cheap, it needs no DOM, and it fails on exactly this shape — a reader with no writer — without knowing anything about revisions.


VLN-048 — The portal-coverage gate passes every cloud_services entity automatically, forever

Field Value
id VLN-048
title validate_portal_presence() computes present = (sentinel in blob) or (eid.lower() in blob), and both portals carry the cloud-services-render sentinel in a comment. The or is therefore satisfied for every cloud_services entity by a string that has nothing to do with that entity, so the gate passes with zero portal edits and cannot distinguish "this entity is rendered" from "a render path exists somewhere on this page". It is the only automated enforcement of the three-portal rule for cloud services, and it enforces nothing.
register Security findings
status open
severity MEDIUM (argued) — no disclosure and no runtime impact. Argued at MEDIUM rather than LOW because this is a control that reports success without testing anything, and the estate has already acted on its green: the Component Coverage Standard tells contributors the pipeline enforces portal presence, CLAUDE.md repeats it, and two live internet-facing Cloud Run services sat unrendered on both portals for three weeks with the gate green throughout. A check whose pass carries no information is worse than an absent check, because an absent check does not get cited. Argued, not supplied.
owner unassigned
evidence scripts/check-component-coverage.py:557,564,584-608. The sentinel constant is PORTAL_RENDER_SENTINEL = "cloud-services-render" and the presence expression is present = (sentinel in blob) or (eid.lower() in blob). _portal_source_blob() concatenates every *.html/*.js in the portal directory including comments, so the sentinel is matched inside an HTML comment. Both portals carry one — status-site/index.html:318 and soc-site/index.html:1215, each reading <!-- cloud-services-render \| T5: replace this hardcoded block with an inventory.json render … -->. Demonstrated, not theorised: gpus_forms_frontend and gpus_forms_backend have been registered in inventory.yaml since 2026-08-05 and appear as a literal data-svc-id on neither portal; they passed every coverage run for three weeks purely on the sentinel. Confirmed by scanning both portals for each ID on 2026-08-26 — 7 of 9 entries rendered, 2 absent, gate green.
acceptance test Re-runnable, and it must fail on a sentinel match — that is the whole point, so the test is written against data-svc-id attributes rather than raw substrings:

python3 - <<'EOF'
import re, sys, yaml
cs = yaml.safe_load(open("inventory.yaml"))["inventory"]["cloud_services"]
gaps = [f"{d}: {e}" for d in ("status-site","soc-site")
for e, v in cs.items()
if v.get("monitoring_status") != "decommissioned"
and e not in set(re.findall(r'data-svc-id="([^"]+)"', open(f"{d}/index.html", errors="replace").read()))]
print("\n".join(gaps) or "PASS"); sys.exit(1 if gaps else 0)
EOF

Both directions required. (1) It must exit 0 on the current tree. (2) It must exit non-zero on the tree as it stood before 2026-08-26, where it reports gpus_forms_frontend and gpus_forms_backend missing from both portals — the two entries validate_portal_presence() passes today. A replacement gate that cannot fail on that input has reproduced this finding. Verified in both directions on 2026-08-26.
blocks
blocked_by
date_raised 2026-08-26
date_verified

The docstring overclaims in a second, independent way

PORTAL_DIRS = ("status-site", "soc-site")

def validate_portal_presence(...):
    """Every non-decommissioned cloud_services entity must be present in
    BOTH status-site and soc-site render sources (literal ID or generic
    render sentinel). Makes the three-portal coverage requirement an
    enforced gate for cloud_services entities."""

PORTAL_DIRS has two entries. The "three-portal coverage requirement" is enforced across two portals. The third, mkdocs-portal, is not checked here at all — it is covered by validate_documented_in(), which is a materially weaker check:

Condition Severity
documented_in is empty error — fails the build
a listed path does not exist warning
a listed path exists but never mentions the entity warning

So for the third portal, listing one path that exists and says nothing is a pass. There are 122 such warnings in the current run and they have never failed anything.

Two separate overclaims, then, in eight lines: the sentinel makes the two-portal check unconditional, and the docstring calls a two-portal check a three-portal one. Neither is visible to anyone reading the summary line RESULT: PASS.

Why this outranks the task that found it

Raised while propagating the travel approval workflow into inventory.yaml and both portals. That work is four entries and eight portal rows. This is the reason such work can be skipped and still look done.

The propagation rule is real and is stated in three places — the Component Coverage Standard, CLAUDE.md's "Adding Infrastructure" section, and the quick-start. All three tell a contributor that the pre-build check enforces it and that drift fails Cloud Build. For cloud_services and portal presence specifically, that is not true and has not been true since the sentinel was introduced. A contributor who adds an entry and no portal row gets a green build and a documented assurance that green means covered.

The estate's recurring failure mode is an artifact that is internally coherent and false against a fact living elsewhere. This is that shape in a control: validate_portal_presence is correct on its own terms — it genuinely tests what its expression says — and false against the claim the standard makes about it.

The sentinel is not a mistake, and it should not simply be deleted

Recorded because the obvious fix is wrong.

The sentinel exists for a real reason, stated at its definition: the T5 refactor will replace the hardcoded portal blocks with an inventory.json render, and after that refactor no literal entity ID will appear in the portal source at all — the rows will be generated at runtime from window._inventory. A gate that demanded literal IDs would fail on the day the render path became correct, which is exactly backwards. The or was put there so the check would survive that transition.

So the fix is not "drop the sentinel". It is to make each branch prove something:

  • Pre-T5 (today). Require a per-entity marker — data-svc-id="<id>" — rather than any substring. That is already the convention every rendered row follows, so the check costs nothing and the acceptance test above passes on the current tree.
  • Post-T5. The sentinel branch is the right idea but must be paired with evidence that the render path actually enumerates the inventory — e.g. asserting the block iterates cloud_services and that the baked inventory.json contains the entity. Presence of a comment is not that evidence.

Two more sentinel occurrences were added on 2026-08-26 by the very commit that fixed the underlying coverage gap (status-site/index.html:335, soc-site/index.html:1234), marking the new approval-workflow tables as T5-replaceable alongside the existing ones. Correct on its own terms, and it also means the free pass is now four comments wide. That is how a check like this erodes — not by anyone disabling it, but by ordinary correct work adding more of the thing it accepts.

Scope — what this does and does not cover

  • validate_portal_presence also checks compute_instances and any asset tagged project: gpus-it-infrastructure, against COMPUTE_RENDER_SENTINEL = compute-instances-render. Same defect, same shape — both portals carry that sentinel too, so all nine compute instances pass unconditionally. They happen to be genuinely rendered today; nothing checks that they stay so.
  • The IAR sub-check inside the same function (eid.lower() not in iar_blob) is a real check — it matches the entity ID against the file's text and can fail. It applies only to gpus-it-infrastructure assets, so no gpus-infra cloud service is covered by it.
  • network_devices, hypervisors, power_devices, storage and linux_hosts have no portal-presence check at all, by design. Not part of this finding; noted so a reader does not infer coverage from its absence.

VLN-049 — The form's own instructions still deny the revise path, and a test keeps them that way

Field Value
id VLN-049
title VLN-047 fixed /status/<id> and the outcome email. It did not reach the third surface that carries the same false sentence — the forms.instructions column served at the top of the travel form itself, which reads "If your request is returned, it cannot be edited. Read the comment, then submit a new request." This is the same column GOV-021 was raised and closed about, wrong again for a different reason. A passing test requires it to say this, so correcting it fails the suite.
register Security findings
status done — CLOSED 2026-08-27. All four surfaces fixed in the version: 3 bump and the served column verified against gpus_forms, which is this row's acceptance test and GOV-021's standard
severity MEDIUM (argued) — same class and the same realized-harm ground as VLN-047, argued at the same level rather than lower because this surface is read earlier. /status/<id> is seen after a return; the instructions are seen by every submitter before they fill the form in, so this is where the expectation is set. A traveller told at submission time that returns cannot be edited will not go looking for the control when one is returned — the fix shipped, and the form still teaches people it does not exist.
owner unassigned
evidence Live, not just in the repo. SELECT current_version, md5(instructions), length(instructions), (instructions LIKE '%cannot be edited%') FROM forms WHERE id='travel-request-001' returned 2 \| e9b6e1acc6bd0b0bf3ce8f2aacc1bee9 \| 570 \| t on 2026-08-27 — byte-identical to forms/travel-request.yaml, so the served text is exactly the wrong text. Four places carry it: (1) the YAML instructions block; (2) forms-backend/test_travel_instructions.py::test_returned_requests_are_described_as_a_new_submission, which asserts the text must say a returned request needs a new submission — a green test pinning the false copy, the same shape as the outcome-mail guard evolved in bc0c84b; (3) forms-backend/approval_decide.py:128-130, whose comment says "What is still absent is a submitter STATUS view and an edit path" — the status view shipped 2026-08-19 and the revise path 2026-08-25, so both clauses are false and the test above cites this comment as its contract; (4) approval_decide.py:138, which still describes a request moving "from Admin Ops to Finance" — step 3 is Final approval, settled in 07d5766 and recorded in the platform doc.
acceptance test Two halves. (1) The live forms.instructions column for travel-request-001 describes the revise path — verified by query against gpus_forms, not by reading the YAML and not by a green build, per GOV-021's own closing standard and VLN-014's worked example of why. (2) test_travel_instructions passes with an assertion that the text describes revising, and fails if the text says a returned request cannot be edited — the guard must be inverted, not deleted, or the copy is left unprotected in both directions.
blocks
blocked_by
date_raised 2026-08-27
date_verified 2026-08-27

CLOSED 2026-08-27 — verified in the database, not in the YAML

Both halves of the acceptance test.

(1) The served column. Read back from gpus_forms after the v3 reload:

current_version | md5(instructions)                | length | cannot_be_edited | says_revise
              3 | 721984482ad6de97b0fa8465561add09 |    780 | f                | t

The md5 and length were recomputed from forms/travel-request.yaml locally and match exactly, so the text staff are served is the text that was written — the same check GOV-021 closed on, for the same reason: reading the YAML proves what was intended, not what is being served.

(2) The guard is inverted, not deleted. test_returned_requests_are_described_as_revisable now asserts "revise" present and "cannot be edited" absent, and test_a_revision_is_still_never_described_as_an_in_place_edit keeps the half that survives. Mutation-proved both ways: restoring the old sentence fails the first, and phrasing revise as "just edit your request and resubmit" fails the second.

All four surfaces landed in the same version: 3 bump as the funding-source fields, so the correction cost one line and no additional deploy — the sequencing note below, acted on rather than recorded and deferred.

THE DURABLE LESSON — VLN-047's sweep grepped SOURCE; this lives in DATA

The reason this survived, stated as the class rather than the instance.

VLN-047 enumerated every surface that denied the revise path and found two — SubmissionStatus.tsx and approval_render.py. Both are code, and the search that found them was a grep over source files for the sentence and for revise. It was a thorough search of the wrong space.

The instructions are data. They live in a YAML block scalar (forms/travel-request.yaml) and, after yaml_loader._upsert_form runs on the next container boot, in a database column (forms.instructions). Neither is a source file in the sense the sweep meant, and a grep of *.py / *.ts / *.tsx cannot see either.

This is the same blind spot that made GOV-021 possible, on the same column. That row's whole finding was that the form's served instructions described a workflow the resolver does not perform, and it survived because "the instructions and STEP_DESCRIPTIONS are both machine-readable and nothing compared them". GOV-021 was closed by fixing the sentence and adding test_travel_instructions.py. What was not done was widening the definition of "a surface" for the next sweep — so the next time a claim went stale, the same column was missed again, by a search that was otherwise careful.

What a sweep for this class would have to cover. Not "grep harder" — a different and enumerable set of places prose about system behaviour is served from:

Class Where Reachable by a source grep?
Component source *.tsx, *.ts, *.py yes — this is what VLN-047 searched
Assembled mail bodies approval_render.py, approval_dispatch.py yes, the literals are in source
Form definitions forms/*.yamlinstructions, every field description, every pulldowns.yaml value no — data, not code
The live forms table forms.instructions, fields.description no — and can drift from the YAML independently
Email templates forms/templates.yaml → the templates table no
Published docs mkdocs-portal/docs/** yes
The registers themselves this file, gap-register.md yes, and they go stale too — see VLN-047's own withdrawn population note

The two rows in bold are the ones both sweeps missed, and they are precisely the surfaces a submitter reads. The estate's own guides (docs/guides/travel-*.md) are a further copy of the same claims, correct today only because someone remembered them.

So the standard a claim-staleness sweep has to meet: enumerate by audience, not by file extension. Ask who is told this, and then find every place they are told it — which for a submitter means the form's instructions and field descriptions before it means any .py file. A grep over source is a search of the developer-facing surfaces only.

A green test is holding the false copy in place

test_returned_requests_are_described_as_a_new_submission asserts "new request" in text or "submit a new" in text. It was written against a correct understanding — no edit path existed — and it is now the reason the copy cannot be corrected without a red suite. Anyone fixing the sentence sees a failing test and has to decide whether they have broken something.

Third instance of this exact pattern in four days. bc0c84b evolved the outcome-mail guard (test_returned_says_a_NEW_request_is_required), 01957b4 evolved two more in the frontend, and this is the same shape again: a guard written against a true fact, outliving the fact, and defending the error it was meant to prevent.

The fix follows the precedent already set — evolve, do not delete. The half that survives is real: a revision is a NEW submission (parent_submission_id, migration 022), so the instructions must still not imply this request is edited in place. "Revise" is now correct; "edit your request" is still wrong.

FIXED 2026-08-27 — four surfaces, one version bump

Folded into the same version: 3 bump as the funding-source fields, on the sequencing argument below: the reload happens regardless, so the correction cost one line and no extra deploy.

# Surface Change
1 forms/travel-request.yaml instructions "it cannot be edited … submit a new request" → tells the submitter they can revise, from /my-requests, with answers prefilled, and that approval restarts at the first approver
2 test_travel_instructions.py guard evolved, not deleted — now asserts "revise" present and "cannot be edited" absent, plus a second test keeping the half that survives
3 approval_decide.py:128-130 "what is still absent is a submitter STATUS view and an edit path" — both shipped (2026-08-19, 2026-08-25); corrected in place with what changed and what did not
4 approval_decide.py:138 "from Admin Ops to Finance""to final approval", matching STEP_DESCRIPTIONS[3] and 07d5766

Mutation-proved, each mutant parsing first. Restoring the old sentence fails test_returned_requests_are_described_as_revisable; rephrasing "revise" as "just edit your request and resubmit" fails test_a_revision_is_still_never_described_as_an_in_place_edit. A guard that only fires in one direction would have accepted the second, which is the in-place edit the original comment was right to forbid.

Cheap now, dearer later — a sequencing note, not a severity argument

travel-request-001 is being bumped to version: 3 for the funding-source fields. That bump reloads forms.instructions from the YAML on the next container boot regardless. Correcting the sentence in the same version costs one line and no additional deploy; deferring it costs a second version bump and a second reload, and leaves the wrong text served in the interval. Recorded as a fact about sequencing so the decision is made deliberately rather than by default.


Lynis scan findings

Daily Lynis scans run at 03:00 on all 4 servers. See Lynis Scan Results for the latest hardening index and warnings per server.

Current hardening indices (2026-03-16):

Server Hardening Index Warnings Suggestions
SKY 79/100 1 (kernel reboot) 23
RAIN 79/100 1 (kernel reboot) 24
SUN 76/100 1 (kernel reboot) 27
WIND 76/100 1 (kernel reboot) 28

Remediation guidance

VLN-001 — SSH MFA

# Install Google Authenticator PAM module
dnf install -y google-authenticator pam

# Configure for each service account
su - dnsadmin -c "google-authenticator -t -d -f -r 3 -R 30 -w 3"

# Add to /etc/pam.d/sshd
echo "auth required pam_google_authenticator.so" >> /etc/pam.d/sshd

# Enable ChallengeResponseAuthentication in sshd_config
sed -i 's/ChallengeResponseAuthentication no/ChallengeResponseAuthentication yes/' /etc/ssh/sshd_config
systemctl restart sshd

VLN-006 — DNS recursive query restriction

# Add ACL to /etc/named.conf
# acl "internal" { 192.168.120.0/23; 192.168.124.0/24; 172.16.0.0/24; };
# options { allow-recursion { internal; }; };
sudo rndc reload


VLN-050 — The HappyFox ticket body is a third render surface, and the v3 bump left it enumerating a 13-field form

Field Value
id VLN-050
title forms/templates.yamltravelrequest001 is the body of the only ticket a travel request raises (queue 45). It enumerates its fields as literal {{ Placeholder }} tokens rather than iterating them. The version: 3 bump of 2026-08-27 added FundingSource, GrantName and CostCenter to travel-request.yamla different file — and the template carries no version of its own, so nothing updated it and nothing failed. The portal and the approval mail show the funding fields; the ticket does not.
register Security findings
status CLOSED 2026-09-01 — both halves met. The template + guard closed 2026-08-31; the second half closed on af7ea350 (Jack Sundius), submitted 13:28 UTC on 2026-09-01, the first submission ever to reach the corrected template. Ticket #USITS00368946 read in queue 45 and confirmed line by line. c84bfb45's short ticket is not repaired and is not a condition of closure — see the closing note
severity LOW → MEDIUM (argued), raised 2026-08-31 — the original LOW rested on realized harm being zero. It is no longer zero: it is one real travel request, c84bfb45, submitted by Madison Carter, whose queue-45 ticket omits the funding attribution she supplied. Still not HIGH — no disclosure, no authz failure, and the approval mail and portal both carried the fields correctly, so the approver was never misinformed. But a durable IT record of a real person's travel is missing two values she entered, and "argued on zero harm" is no longer available as the argument. No disclosure, no authz failure, no data loss: the ticket under-reports fields the triager can see in the portal. Deliberately not argued at MEDIUM like VLN-049, whose surface was read by every submitter before filling the form in. This one is read by one triager, after the fact, beside a portal page that is correct. What raises it above trivia is not this instance — it is that the surface was uncounted, and would have taken a fourth cost field down with it silently.
owner unassigned
evidence Live, not just the repo — the VLN-049 standard. Queried gpus_forms as maple-agent@gpus-infra.iam on 2026-08-31: SELECT id, length(body), md5(body), body LIKE '%FundingSource%' … FROM templates WHERE id='travelrequest001' returned travelrequest001 \| 531 \| 5016c0c4b24e1cb1672b3fcbdef230e1 \| f \| f \| f \| t — no funding fields, CostAirfare present. The md5 and byte length were recomputed from forms/templates.yaml locally and matched exactly, so the served body is the written body. Meanwhile SELECT id, current_version FROM forms WHERE id='travel-request-001' returned 3. Parsed diff: the form defines 16 field keys, the template rendered 13 placeholders; IN FORM, NOT IN TEMPLATE = ['FundingSource','GrantName','CostCenter'], IN TEMPLATE, NOT IN FORM = []. One action only: SELECT form_id, action_order, action_type, destination, template_id FROM actions WHERE template_id='travelrequest001' → a single email_template to gpus-it-support@greenpeace.org, so there is no per-queue fan-out to make partiality correct. travelrequest001 has been edited exactly once in its life78cdd23, 2026-08-18, the commit that created the travel form. It predates version: 2 and version: 3 entirely.
acceptance test forms-backend/test_travel_ticket_template.py — asserts the template's placeholder list equals the form's field-key list, in order. A relationship, not a literal, per this family's own diagnosis. Mutation-proved both directions on 2026-08-31: run against the pre-fix templates.yaml, 4 failures; against the corrected file, 4 passed. Second half, and the one that actually closes it: a real submission's ticket body in queue 45 carries the three funding lines — the YAML is what was intended, the ticket is what a triager reads. NOT MET. The first real submission — c84bfb45 — was rendered by the old template 6m06s before the fix reached the templates row, so its ticket is evidence the defect was real, not evidence it is fixed. The next travel submission is the first that can meet this half.
blocks
blocked_by
date_raised 2026-08-31
date_verified Template + guard 2026-08-31. Ticket-body half MET 2026-09-01 on af7ea350 / #USITS00368946 — submitted 13:28:23 UTC, routed 13:28:26, against a templates row reloaded 12:31:36

CLOSED 2026-09-01 — on a real ticket, read line by line, not on the YAML

af7ea350, Jack Sundius, submitted 13:28:23 UTC, routed 13:28:26 — the first submission ever to reach the corrected template. The templates row was reloaded at 12:31:36 by a container boot after the 12:50 report deploy window, so the render is 57 minutes downstream of the fix rather than 6m06s upstream of it, which is what closed c84bfb45's ticket short.

Ticket #USITS00368946, queue 45, read by R. Chhetry. The four lines appear consecutively and in form order immediately after Department:

How is this travel funded?:     Cost Center
Grant name:
Cost centre:                    20100 — INFORMATION TECHNOLOGY
SMT member you report to:       Kevin Toruno

— and Event registration fee (USD): 0 is present, the v4 field rendering in a real ticket for the first time, two days after it shipped.

WHICH RENDERING SHIPPED, and why the question was asked. GrantName is legitimately blank here: funding is Cost Center, and the field's own description says to leave it blank in that case. Two renderings were possible and they mean different things:

ships when reads as
Grant name: + nothingTHIS ONE the SPA stored an empty-string row a question asked and answered "none"
Grant name: [GrantName: —] no submission_fields row was stored template_render.VisibleUndefined — "the template names a key the form did not supply"

The empty-string rendering shipped. Confirmed in the ticket, and it matches the database: GrantName appears exactly once in the whole submission_fields table — this submission — at ciphertext length 16, the global minimum across every row. All 17 of 17 declared fields have stored rows. So the SPA writes a row for a touched-but-empty optional text field, and {{ GrantName }} is defined and empty rather than undefined.

That distinction is recorded because the marker would also have been a pass under this row's acceptance test — the label survives either way — while meaning something quite different about the SPA. A future reader seeing [GrantName: —] on some later ticket should know it is a change in behaviour, not the normal blank.

What could NOT have failed, stated so the test is not over-credited. The labels are literal static text in the template (Grant name:\t{{ GrantName }}). A label cannot go missing unless the whole line is absent from the template — which is the original defect, and is now held by test_travel_ticket_template asserting the placeholder list equals the form's field-key list in order. What this ticket proves is the part the test cannot reach: that the deployed templates row, the live renderer and the queue-45 delivery path all agree with the YAML.

c84bfb45's ticket is still short, and that is deliberately not a condition of closure

Madison Carter's ticket remains missing the funding attribution she supplied. The correct action is unchanged — a note on the existing queue-45 ticket carrying the two values, not a re-route: c84bfb45 is IN_APPROVAL with a real approver holding step 2, there is no re-route endpoint, and inventing one to re-render a ticket would be a state change on a live travel authorisation.

This row closes without it because the row is about a render surface, and the surface is now proven correct end to end on a real submission. The outstanding item is a single historical record repair, owned by Rajesh, tracked in the workstream close-out rather than by holding a verified finding open. Holding it open would misreport the estate's state: the defect is fixed and cannot recur.

SUPERSEDED 2026-08-31 — TRUE WHEN WRITTEN, falsified six minutes later. Read the reopening note below it.

CORRECTION to the raising brief — the affected population was ZERO at the time of writing

This row was raised on the understanding that "every HappyFox ticket raised from a travel request since 2026-08-27 is missing the funding fields" and that queue 45's triager "has been reading an incomplete record for four days." Checked, and it did not happen.

all_travel_subs | since_v3 | most_recent
             12 |        0 | 2026-08-26 17:09:33.222181+00

Twelve travel submissions all-time, none since the v3 bump, the most recent one predating it by a day. So no incomplete ticket was ever raised and nobody has read one. That is consistent with M-01 — the same reason count(*) WHERE searchable_values ? 'FundingSource' is zero.

The defect was real and the harm was not. Recorded as a correction rather than quietly downgrading the severity, because the difference between "this happened twelve times" and "this has never happened" is the difference between a remediation and a near-miss, and the register should be able to tell them apart afterwards.

It also inverts the sequencing argument in the useful direction. The first travel submission to hit this template would have been the v3 exercise itself — the one submission the estate is running specifically to prove the funding fields reach the surfaces that consume them. Fixing the template before that submission is not tidiness; it is the difference between the exercise proving three surfaces and proving two.

REOPENED — the population became ONE, 6m06s before the fix landed, and it is a real person

c84bfb45-27e2-4250-aa67-62b42b215947, Madison Carter (mcarter@), submitted 2026-08-31 15:46:12.683Z. The first genuine end-user travel request since the v3 bump, from someone who did not know any of this had changed.

Time (UTC, 2026-08-31) Event
15:46:12.683 Madison submits. form_version = 3 — served all 16 fields, answers 15
15:46:15.668 Step 1 dispatched to Kalina Francis
15:46:30.259 Approval mail sent — carries Funded by: and Cost centre: correctly
15:46:32.088 Ticket email rendered and sent to queue 45 — OLD 13-placeholder template
15:46:32.412 submission_routed
15:48:38 The fix commit is pushed; Cloud Build triggers fire
15:52:38.981 templates.travelrequest001 upserted with the corrected 640-byte body

Her ticket missed the fix by 6 minutes and 6 seconds.

What queue 45 is holding. Two values she actually entered are absent — not blank, absent:

  • FundingSource = Cost Center
  • CostCenter = 23051 — COMMUNICATIONS CORE

GrantName is not among them: she chose the Cost Center branch and left it blank, so it has no row in submission_fields at all (15 rows, not 16). A correct rendering would have shown it as a present label with an empty value.

The blank-vs-missing distinction — the whole risk — resolves cleanly here, by luck of this template's shape. A missing placeholder removes the entire line, label included. So the ticket contains no How is this travel funded? line and no Cost centre line anywhere. A legitimately blank value would have shown the label followed by nothing. A triager cannot confuse the two in this instance because the labels themselves are gone — but that is a property of how this template is written, not a guarantee. A template that emitted labels unconditionally would have produced the ambiguous case, and the next one might.

What was NOT affected, so the blast radius is not overstated: the approval mail Kalina Francis received carries both funding lines and a correct Total: 500.00 (lodging 0 rendered as 0.00 and kept, so the total is complete, not partial); the portal renders all 15 answered fields in field order; and searchable_values stored the cost centre byte-identical to the pulldown option — md5 dd61e655b30d38f828c126cccfd9f445 on both sides, 29 bytes over 27 characters. Nothing was lost from the database and no approver was misinformed. Only the ticket is short.

Repair is by hand, and must not touch the request. c84bfb45 is IN_APPROVAL with a real approver holding it. There is no re-route endpoint, and inventing one to re-render a ticket would be a state change on a live travel authorisation for a real trip. The correct action is a note on the existing queue-45 ticket carrying the two values.

M-14, applied to this row by events rather than by a sweep

The superseded note above said the affected population was zero. It was true when written — 12 submissions, none since the v3 bump — and a user falsified it six minutes later by filling in a form.

This is M-14 with a new operand, and it is worth recording because the previous three instances were all spatial: a contradicting claim elsewhere on the page, in another file, in the database. This one is temporal. No sweep could have found it; at the moment of writing there was nothing to find.

"Realized harm is zero" is a measurement, not a property. It carries a timestamp whether or not one is written, and it decays. A row argued down to LOW on the basis that nothing has happened yet must record when that was measured and must be re-measured before the row closes — because the interval between raising a defect and fixing it is precisely the interval in which the thing is still broken.

The superseded note argued that fixing before the exercise was "the difference between the exercise proving three surfaces and proving two." The reasoning was right and the timing beat it. The estate did not get to choose which submission was the first to hit this, and the one that arrived was better evidence than the planned test — and a live problem for someone's actual trip.

THE CLASS — 'true when written, made false elsewhere', third instance

Filed against the split this tracker already drew for VLN-045 and VLN-049.

travelrequest001's body was complete and correct on 2026-08-18 against a 13-field form. It was falsified on 2026-08-27 by an edit to travel-request.yaml, in a different file, that never touched it. As with the other two, a snapshot test on the template passes before the change and after it, because the literal never moves.

What is new here, and why this is worth its own row rather than a footnote on VLN-049: the other two were surfaces someone had at least thought about and got wrong. This one was not counted at all. The v3 work identified "both renderers" — approval_dispatch.py and approval_render.py — and updated both. There were three.

Why the version column could not catch it. version: is a property of the form (forms.current_version), and it advanced 2 → 3 correctly. The template is a separate row in a separate table with no version of its own, and yaml_loader._load_templates upserts on the template YAML's content — which had not changed. So the bump was well-formed, the loader was correct, and the drift is invisible to both.

What a v-bump checklist would have to cover to catch the class

The trigger is not "a YAML changed". It is "the field set changed" — and the file that changes is never the file that goes stale. So a checklist keyed on the edited file cannot work. It has to be keyed on the field set, and it has to enumerate the consumers.

The discriminator is ENUMERATES vs ITERATES. A surface that walks the fields it is given adapts for free. A surface that names them in a literal list must be edited by hand on every field-set change, and nothing tells anyone. Only the second kind belongs on the checklist:

Surface Shape v3 outcome
forms/templates.yaml → the ticket body enumerates missed — this row
mkdocs-portal/docs/guides/travel-request-submitting.md enumerates (numbered table) missed — see below
approval_dispatch.pyCOST_FIELDS + literal sv.get() enumerates ✅ updated
approval_render.pyREAD_KEYS, OUTCOME_READ_KEYS, COST_FIELDS enumerates ✅ updated, and the only one with a structural guard
gpus-reports/summary_report.py — R1's four key constants enumerates ✅ updated
forms-frontend/.../ApprovalDetail.tsxfields.map(...) iterates ✅ safe by construction

So: four enumerating surfaces were in scope for v3 and two were missed — both of them outside forms-backend/, which is where the attention was.

The durable form of the checklist is not a checklist. It is the property approval_render.py already has and the other three did not: TestReadKeysIsLoadBearing compares the declared key set against the keys the source actually reads, and fails in both directions. That is why approval_render.py was updated and the template was not — one of them had a test that could notice. test_travel_ticket_template.py gives the template the same property. R1's constants and the docs guide still do not have it.

What cannot be generalised, stated so nobody tries. "Every form's template must cover every field" is false for this estate and would fail 61 of 67 form/template pairs. Templates fan out per queue: a multi-action form renders a different template per destination, each scoped to what that queue needs (89ebf11). Partial is correct there. The strict rule only holds for a form whose sole action is one email_template, which is why the new test asserts that premise first and fails if travel ever gains a second destination.

A FOURTH surface, found while writing this row, and fixed with it

mkdocs-portal/docs/guides/travel-request-submitting.md — the staff-facing guide, published on infra.greenpeace.us — carried three claims the same v3 bump falsified:

  • "Thirteen questions. All of them are required." Sixteen, and two are not.
  • A numbered table of 13 questions, correct through question 2 and wrong from question 3 on, because the funding fields sit between Department and SMTApprover in the served order.
  • "Question 3 matters more than the others" and "the SMT member you named in question 3" — SMT approver is now question 6. Two co-located cross-references to a number the same edit moved.

Corrected here rather than deferred, because deferring it would repeat the exact mechanism this row is about. Also added: an explicit note that nothing validates questions 3–5 against each other, which is the documented consequence of the platform having no conditional display and no server-side cross-field validation — a submitter should be told that before they fill the form in, not discover it.

Two more stale surfaces the sweep found, and what was done with each

Applying M-14 to this row's own fix — grep the estate for the claim you just corrected — turned up two more artifacts the v3 bump falsified, neither of them a render surface:

  • forms-frontend/src/routes/ApprovalDetail.tsx:295,347 — two comments asserting "All thirteen fields" and "scrolled past thirteen fields". Fixed. The component maps over what the API gives it, so it was never wrong in behaviour; the number in the prose was a claim nothing checked. It now states no count, and says why.
  • forms-backend/test_approval_access.py:495 — a module constant THIRTEEN listing the pre-v3 field set, consumed by test_returns_all_thirteen_fields_decrypted and by test_submission_status.py. Not changed. It is a fixture, not an assertion about the live form: the test proves the endpoint returns what it was handed, which is still true and still worth proving. But the name and the list now describe a form that no longer exists, and a reader would reasonably take it for the travel field set. Renaming it and extending the fixture to sixteen is correct and is deliberately left as its own change, because it touches two test modules and this row is closed on a template.

Not in scope, and named so it is not mistaken for clean

16 of the 17 single-action forms in forms/ show the same drift today. The travel form is where this class was caught, not the only place it lives. Widening the guard is real work with a real backlog behind it and it is not this row; raising it here so the next person does not read test_travel_ticket_template.py's narrow scope as evidence the rest are fine.


VLN-051 — The approval decision audit row was built for non-repudiation and reaches no consumer

Field Value
id VLN-051
title approval_decision_recorded and approval_notification_sent are values 32 and 33 of models.py AUDIT_ACTIONS, and appear in neither ship_actions nor excluded_actions in soc/forms-authz-detection/soc_shipper_allowlist.json — 31 of 33 declared, and the two undeclared are exactly the approval ones. soc-log-shipper builds its filter from ship_actions, so no approval decision in this estate has ever reached the SOC stream.
register Security findings
status open — raised 2026-08-31. No change made to ship_actions or the Wazuh ruleset in this pass, deliberately
severity MEDIUM (argued) — argued on what the row was built to be, not on a hypothetical attack. approval_decide.py writes the audit row inside the decision's transaction, and its own comment states the reasoning: audit_log has UPDATE/DELETE revoked from PUBLIC and from every application role, while approval_steps is UPDATE-able by forms_app and carries no immutability constraint (ASVS fact G). So the immutable audit row is the thing the non-repudiation guarantee actually rests on — §2.2 asserts non-repudiation as the endpoint's product. The row exists, it is correct, it is written atomically with the decision, and nothing consumes it. A control that is implemented and unread is this estate's characteristic failure, and it is the same shape as VLN-011 and VLN-040: the mechanism runs and the result goes nowhere.
owner unassigned
evidence From source, 2026-08-31. forms-backend/models.py:204-222AUDIT_ACTIONS is a 33-tuple; the comment at :201-203 records it reaching 33 "after 020 added approval_notification_sent" and requires any enum-altering migration to update it. soc/forms-authz-detection/soc_shipper_allowlist.jsonship_actions has 17 entries, excluded_actions 14; 17 + 14 = 31. Set difference computed directly: AUDIT_ACTIONS − ship_actions − excluded_actions = ['approval_decision_recorded', 'approval_notification_sent'], and (ship_actions ∪ excluded_actions) − AUDIT_ACTIONS = []. The file's own header is the standard it fails: "Do NOT hand-edit… Regenerate via gen_authz_detection.py; a new enum value forces a ship decision." Two enum values were added — migrations 018 and 020 — and neither forced one. The write site is forms-backend/routes_phase2.py:1189-1204 (audit_row, success: True, inside the decision transaction) and the durability reasoning is approval_decide.py:21-30.
acceptance test Two halves, both able to fail. (1) Structural, and it fails today: a test asserting set(AUDIT_ACTIONS) == set(ship_actions) | set(excluded_actions) — i.e. every enum value carries an explicit ship decision, in one direction or the other. This is the file's stated contract and nothing currently enforces it; it belongs beside gen_authz_detection.py so a future enum value cannot be added without one. Note it must assert declaredness, not shipping — "excluded, with a reason" is a valid answer for a high-volume action and the test must accept it. (2) Live, and it is the one that closes the row: after a ship decision is taken and applied, a real approval decision produces a matching event in Wazuh — observed in the SOC stream by submission_id, not inferred from config. Per VLN-010's worked example, a rule present in a file is not a rule that fires.
blocks
blocked_by
date_raised 2026-08-31
date_verified

What this row does NOT claim

It does not claim the audit row is missing, wrong, or losable. It is written in the same transaction as the decision, so a decision without its audit row cannot arise from the application path — which is precisely why FP-IR-06 can use audit_id IS NULL on a decided step as a tampering signal. The durable record is sound.

The finding is that the record is unread. Detection and forensics are different things: the row is excellent forensics and there is no detection, because nothing is watching. Whether the right answer is ship_actions, a scheduled reconciliation query, or an explicit excluded_actions entry with a reason is a ship decision, and taking it inside a playbook edit or a register row would be exactly the hand-edit the allowlist header forbids.

Deliberately not fixed in this pass

ship_actions feeds a live log shipper and the Wazuh ruleset feeds live alerting. Changing either is a detection-config change with a blast radius on the SOC's alert volume, and it should be made as its own change with its own verification — not as a side effect of writing up the finding. Recorded, not actioned.


VLN-052 — A denied approval attempt is logged at level 3 and pages nobody

Field Value
id VLN-052
title approval_access's denial reasons are absent from soc/forms-authz-detection/authz_reasons.json, so a denied approval-decision attempt matches none of rules 100031100036 and falls through to the parent rule 100030 at level 3 — below the level ≥ 10 threshold that auto-creates a SOC ticket. The denial is recorded and is shipped; it simply arrives beneath the floor at which anyone is told. An attempt to forge an approval decision is logged and nobody is paged.
register Security findings
status open — raised 2026-08-31. No change made to authz_reasons.json or the ruleset in this pass, deliberately
severity MEDIUM (argued) — argued at the same level as VLN-051 and for the complementary reason. VLN-051 is the successful decision that is never seen; this is the failed one. Together they mean the approval endpoint — the control approval_access.py exists to enforce, on a live financial authorisation — produces no SOC signal in either outcome. Not argued higher because the denial is genuinely recorded and queryable, 403 is uniform and the endpoint is not an enumeration oracle, and the authorisation itself holds: this is a monitoring gap, not an access-control failure.
owner unassigned
evidence From source, 2026-08-31. The write site is forms-backend/routes_phase2.py:1154-1166: on step is None or not authorizes_decision(reason) it logs approval_decision.denied and writes _audit_v2(session, "auth_failure", …, details={"reason": reason, "endpoint": "approval_decision"}, success=False). auth_failure is in ship_actions, so the event reaches Wazuh. soc/forms-authz-detection/forms_authz_rules.xml:6-11 — rule 100030, <action>auth_failure</action>, level 3, no reason field. Rules 100031100036 each add <if_sid>100030</if_sid> plus <field name="details.reason">^…$</field>, anchored exact-match. authz_reasons.json lists six reasons — auto_provision_denied, inactive_user, insufficient_role, not_owner, unknown_required_role, user_resolution_failed — all sourced from models.py AUTHZ_DENIAL_REASONS, which is auth_v2's vocabulary. approval_access has a different and disjoint vocabulary (approval_access.py:91 — the reason constants behind authorizes_read / authorizes_decision), none of which appears in that file. With ^…$ anchoring, a non-matching reason cannot fall into 100031100036; it stops at the parent. The ≥ 10 threshold is the SOC ticketing rule already cited in FP-IR-02.
acceptance test Two halves. (1) Structural, fails today: approval_access's denial-reason constants are the single source for a generated tier in authz_reasons.json, exactly as AUTHZ_DENIAL_REASONS already is for auth_v2 — with a test asserting the generated file covers every constant, so a new reason cannot be added without a level decision. This is the same relationship-not-literal shape as TestReadKeysIsLoadBearing. (2) Live, and it closes the row: a deliberate denied approval attempt against a non-approver identity produces a Wazuh alert at the chosen level, observed via wazuh-logtest-legacy -v and in the live archive. Per VLN-010, a rule that loads is not a rule that fires — that row's whole mechanism was rules that parsed cleanly and were unreachable.
blocks
blocked_by
date_raised 2026-08-31
date_verified

The level is a decision, not an oversight to be corrected upward

A denied approval attempt is not automatically level 10. not_owner earns 10 because cross-user ID-walking is an IDOR probe; insufficient_role earns 6 because a role probe is common and mostly benign. The approval_access reasons are not uniform either — round_not_in_approval fires whenever an approver opens a request someone else has already decided, which is normal behaviour, while no_matching_step on a DISPATCHED round is the one that looks like a forgery attempt.

So the fix is a per-reason tier, not a blanket promotion, and tiering them is the ship decision this row asks for. Setting them all to 10 would produce exactly the alert-fatigue that makes a SOC stop reading, which is a worse outcome than level 3.

Read with VLN-051 — together they are the whole endpoint

Neither row is interesting alone. Together: the approval endpoint records a successful decision in a table nothing ships, and a denied one at a level nothing tickets. approval_access.py exists because unauthorised state transitions on a travel authorisation are the threat, and the SOC currently cannot see either side of it.

This is why FP-IR-06 states its detection as "a human noticing" and labels its queries forensics rather than monitors. That playbook is honest about the gap; these two rows are the gap.