Vulnerability Tracker¶
Classification: CONFIDENTIAL — Internal Use Only Document:
security/vuln/tracker.md· v1.34 · 2026-09-01 · GPUS-IT
Summary¶
| Severity | Open | In Progress | Remediated |
|---|---|---|---|
| Critical (CVSS ≥ 9.0) | 1 | 0 | 0 |
| High (CVSS 7.0–8.9) | 2 | 1 | 1 |
| Medium (CVSS 4.0–6.9) | 2 | 2 | 0 |
| Low (CVSS < 4.0) | 2 | 0 | 1 |
What this table counts — and a recount, 2026-09-08
Scope: VLN-001–VLN-012 only — the two tables immediately below. It
does not cover the Security findings register (VLN-013 onward),
which holds the majority of the estate's open findings. A reader taking
these twelve as the total will understate the backlog by a wide margin.
Recounted 2026-09-08 against the rows, and it had drifted. The previous
version totalled 14 across tables holding 12 rows: High read
2 / 2 / 1 and Medium 3 / 2 / 0, neither of which matched. The counts
above are derived from the rows rather than maintained by hand.
VLN-010 and VLN-011 carry no CVSS by design (footnotes 1 and 2) and
are counted under their argued severity, High. A CVSS-banded table
cannot place a row that deliberately has no score; that is a limitation of
the banding, not a missing score.
Two register-hygiene defects — raised 2026-08-10, DECIDED and applied 2026-08-13
Read this first if you followed a reference to VLN-009 and did not find
what you expected. Both defects below were recorded on 2026-08-10 and left
open for an owner decision. That decision was taken on 2026-08-13 and is
applied in this version (v1.7). This note is retained, not deleted, so that
references written between 2026-03-16 and 2026-08-13 still resolve.
Defect 1 — VLN-009 was used twice. It was carried by both the open
Critical passwordless-sudo finding and the remediated chronyd/NTP entry.
DECIDED 2026-08-13:
VLN-009stays with the passwordless-sudo finding (VLN-009) — its row, page, nav entry and all cross-references are unchanged. The chronyd/NTP entry is reassigned toVLN-012and now appears under that id in Remediated vulnerabilities below.If you arrived here from a pre-2026-08-13 reference that meant "VLN-009 = chronyd / NTP / timestamp skew", the row you want is now
VLN-012. Nothing about that finding changed — same CVSS 2.6, same remediation date 2026-03-10, same conclusion. Only the identifier moved.
Defect 2 — VLN-011 had a published finding page but no row here. It
was registered in the nav and under security/ but never entered in this
table; adding it was out of scope for the VLN-010 publication pass.
DECIDED 2026-08-13: the missing row is added to Open vulnerabilities below, pointing at the existing finding page. The page and nav entry were already correct and are unchanged.
No id other than the chronyd entry was renumbered. VLN-010 and VLN-011
are untouched, and no new id was minted for the passwordless-sudo finding.
Open vulnerabilities¶
| ID | Component | Description | CVSS | Severity | Affected | Target | Status |
|---|---|---|---|---|---|---|---|
| VLN-009 | Sudo / SSH | Passwordless root (NOPASSWD: ALL) for the backend SSH account on all four WDC hosts; key is mounted into three internet-facing Cloud Run services |
9.9 | Critical | SKY/RAIN/SUN/WIND | Q3 2026 | Open |
| VLN-011 | Forms / Routing worker | Forms HappyFox Dispatch Never Ran — routing_worker.py:680-694 matched every happyfox_template action, wrote status='deferred', recorded success=true and continued; no HappyFox client existed anywhere in the repo (HAPPYFOX_URL declared at config.py:78, never read). 0 of 37 submissions had ever received a HappyFox ticket id; 19 sat at happyfox_status='deferred'. Two forms carried no email action, so they delivered nothing while the UI told submitters "The relevant team has been notified." Confirmed 2026-08-06 read-only against live Cloud SQL + repo. it-support-request CLOSED end to end — fix deployed and a real ticket confirmed opened (finding §11). new-employee-notification-contractor-intern still HELD; contract-extension-notification UNREVIEWED. Full detail: VLN-011 |
n/a² | High | Forms portal (Cloud Run + MAPLE routing worker) | Q3 2026 | In Progress — 1 of 2 affected forms verified closed; second form held |
| VLN-001 | SSH Config | No MFA — password auth disabled but no TOTP/U2F second factor | 7.5 | High | ALL | Q2 2026 | Open |
| VLN-002 | Network | No VLAN segmentation between server roles | 7.2 | High | ALL | Q3 2026 | Open |
| VLN-003 | GCP IAM | Service account key stored on-disk — no Workload Identity Federation | 6.5 | Medium | GCP | Q2 2026 | In progress |
| VLN-004 | ESXi 6.7 | EOL hypervisor — no vendor security patches since Oct 2023. ⚠ SCOPE DISPUTED 2026-08-13 — see VLN-013: authenticated vSphere API shows the 6.7 builds are on fire (17700523, 6.7 U3) and flower (8169922, 6.7 GA 2018, never patched); water is on 8.0.3 (24022510) and is not EOL. The Affected cell below is believed wrong. It is left unchanged pending VLN-013's acceptance test rather than corrected here |
6.3 | Medium | WATER hypervisor (disputed — see VLN-013) | Q4 2026 | Open |
| VLN-005 | Identity | No SSO — separate credentials per service, no central identity | 5.4 | Medium | ALL | Q3 2026 | Open |
| VLN-006 | DNS | Recursive resolver accepts queries from all internal hosts | 4.9 | Medium | SKY/RAIN | Q2 2026 | In progress |
| VLN-007 | Logging | No centralized SIEM alerting — manual log review only | 3.8 | Low | WIND | Q3 2026 | Open |
| VLN-008 | Backup | No automated backup integrity verification or restore testing | 3.1 | Low | ALL | Q2 2026 | Open |
¹ VLN-010 carries no CVSS deliberately. CVSS scores a weakness that an attacker exploits against a system; this is a detection control that does not run. There is no attack vector, no privilege requirement and no impact metric that describes it honestly, and any score produced would be false precision in a register people make prioritisation decisions from. Severity is assessed as High on operational grounds: the control is the primary alert-fatigue remediation for the estate, it is documented as delivered, and 99.0% of level-12 volume is the noise it was written to remove. Full detail: VLN-010.
² VLN-011 carries no CVSS. The source finding
(security/finding-2026-08-06-forms-happyfox-never-dispatched.md, v1.1) assesses
severity as High — silent loss of user-submitted requests and states no CVSS
vector. None is invented here. As with VLN-010, this is a delivery control that
did not run rather than a weakness an attacker exploits, and a score produced
after the fact would be false precision in a register used for prioritisation.
Severity High is carried through from the finding page as written.
Remediated vulnerabilities¶
| ID | Component | Description | CVSS | Remediated | Notes |
|---|---|---|---|---|---|
| VLN-012 | NTP | chronyd not verified — potential log timestamp skew | 2.6 | 2026-03-10 | chronyd verified and syncing on all 4 servers. Formerly VLN-009; reassigned 2026-08-13 to resolve a duplicate ID. The original VLN-009 designation is retained by the passwordless-sudo finding. |
| VLN-010 | Wazuh / Detection | Baseline rules 100016/100017/100018 were loaded, valid and counted but never evaluated — Wazuh sorts siblings by level and stops at first match, so the level-3 baselines sat behind the level-12 catch-all 100015 and could not be reached. The tuning written to suppress ~184 daemon alerts/day never took effect, and every observable was green: the rules loaded, the count was correct, validation passed. Six hypotheses were eliminated, all of which assumed rejection at load; none was |
n/a¹ | 2026-09-08 | Fixed by re-parenting the three baselines as children of 100015 via <if_sid>, verified before deploy on 4.14.4-rc2 and deployed 2026-08-11 18:29:50 UTC. All of D1–D6 measured and passed. D6(a) mail down ~95% at the inbox, counted on the real recipient to=<rajesh.chhetry@greenpeace.us>: 351 across the fix boundary, then 19 / 18 / 25 / 7. Two earlier needles (ossec, wazuh) both returned 0 and both were meaningless — postfix logs addresses, not the sending application. Report as a RATE: email_maxperhour=12 capped Wazuh at 288/day against ~184/day pre-fix, so 351 is a throttled sample, not an alert count — "351 became 19" is wrong. D6(b) indexed volume flat at 182–184/day vs 183–184/day pre-fix — flat IS the pass, since log_alert_level=3 keeps every suppressed event indexed; the fix removes mail, not documents. D6(c) 100016/100017/100018 non-zero and exclusively level 3. Compound failure (silence mistaken for suppression) excluded by measurement, not assumption. BT-003 proven in production: 100015 fired 15 times since 2026-08-12, every one at level 12 — the daemon noise went without touching the detection. Reproduction scripts retained on MAPLE. See the finding |
Security findings register — schema v1 (opened 2026-08-13)¶
This section uses the schema below; the tables above do not. The legacy tables (
VLN-001–VLN-012) keep their original columns and are unchanged. New findings from 2026-08-13 onward are recorded here against the full schema:id · title · register · status · severity · owner · evidence · acceptance test · blocks · blocked_by · date_raised · date_verified.Status vocabulary:
open·in progress·blocked·done.donemeans verified end-to-end via the real production path against live state. Authored, present, loaded or validated is not done.One row in this section is
done:VLN-041, closed 2026-08-25. It is the first, and it is closed against a live observation on the deployed revision — a real Okta token returning 200 on the endpoint that runs the changed code, correlated to that revision's digest in the Cloud Run request log. Not "authored", not "merged", not "the build was green".
VLN-014is why that bar is written the way it is. It was recordeddoneon 2026-08-21 against a verification performed 2026-08-17; a re-query on 2026-08-24 found the condition no longer held, and the row is back toin progress. See the regression note on that row — it is a worked example of why this rule says verified end-to-end against live state and not verified once.VLN-041carries the same exposure: it is closed against today's revision, and a future revision that reintroduces the defect would reopen it. What makes that unlikely rather than merely hoped-for is condition 2 — the guard is executed by a gate that has demonstrably failed a build over exactly this suite.
evidencerecords what was read, where, and on what date.acceptance teststates in advance the observable condition that closes the row. A row missing either is marked INCOMPLETE and is not filled with a plausible guess. Fields recorded asnot assessedwere not supplied and were not inferred.
Index¶
| ID | Title | Status | Severity | Blocked by |
|---|---|---|---|---|
| VLN-013 | ESXi 6.7 EOL is on fire and flower — VLN-004 misattributes it to water | open | not assessed | — |
| VLN-014 | sky and rain serve different DNSSEC signings — converged 2026-08-17, diverged again 2026-08-19 | in progress | High (as supplied) | — |
| VLN-015 | NOPASSWD sudo on emu and ostrich — VLN-009 scope extends beyond WDC | open | not assessed | — |
| VLN-016 | One person holds three identity strings with inverted access across systems | open | not assessed | — |
| VLN-017 | gcloud list verbs exit 0 on permission denial |
open | not assessed | — |
| VLN-018 | All 10 Cloud Run services in gpus-infra are ingress: all |
open | not assessed | — |
| VLN-019 | duck.cloud.us.gl3 resolves to an address that is not duck's |
open | not assessed | — |
| VLN-020 | Five Puppet node definitions can never match a certname | open | not assessed | — |
| VLN-021 | 3 of 4 Meraki org admins hold orgAccess: full on the VPN-terminating org |
open | not assessed | — |
| VLN-022 | Published forms go-live record asserts a delivery claim VLN-011 disproves | open | not assessed | GOV-009 (correction cannot be verified live) |
| VLN-023 | vmstorage running a degraded RAID 6 with no spare — NFS datastore for every WDC VM |
in progress | CRITICAL | — |
| VLN-024 | Plaintext vendor credentials in the Puppet control repo (pushtab) |
open | CRITICAL | — |
| VLN-025 | puppetCrypt exists and pushtab bypasses it |
open | HIGH | — |
| VLN-026 | vmstorage has no backup or replication task |
open | HIGH | GOV-019 (derived) |
| VLN-027 | Meraki IPsec pre-shared keys returned in plaintext | open | HIGH | — |
| VLN-028 | SMBv1 enabled with signing off on all three Synology units | open | HIGH | — |
| VLN-029 | Key material in a control repo scheduled for deletion | open | HIGH | — |
| VLN-030 | Undocumented second VPN peer (Target, IKEv1) |
open | MEDIUM | — |
| VLN-031 | Zero 2FA across all three Synology units | open | MEDIUM | — |
| VLN-032 | Meraki org 395909 has no SSO and no IdP | open | MEDIUM | — |
| VLN-033 | synstorage volumes at 93.1% and 82.3%, status attention |
open | MEDIUM | — |
| VLN-034 | rooster cannot compile in production — archive::backup body commented out |
open | MEDIUM | — |
| VLN-035 | dataflow.us.gl3 resolves to a host that does not exist |
open | MEDIUM | — |
| VLN-036 | wdc-wap-5 status alerting |
open | LOW | — |
| VLN-037 | Four production security reports ran on EOL Python 3.6.8 for ~4 months | remediated | HIGH | — |
| VLN-038 | No Cloud Build in the estate ran unit tests — controls implemented as tests were never enforced | in progress | High (argued) | — |
| VLN-039 | Pre-build check steps couple every deploy to PyPI availability — a gate can block an incident fix | in progress | Medium (argued) | — |
| VLN-040 | The dependency scan runs, finds 24 advisories, and its exit code is discarded — green on every build since April | in progress | High | — |
| VLN-041 | auth_v2 took the JWT algorithm allowlist from the unverified header of the token being validated |
done 2026-08-25 | High (argued) | — |
| VLN-042 | CORS(app, origins="*") on an unauthenticated, ingress: all backend — any site a staff member visits can read /metrics and /health/deep from their browser |
done 2026-08-25 | MEDIUM (argued) | — |
| VLN-043 | The forms SPA reads none of the three variables its own env template defines, and the template's Okta clientId is the pre-cutover one | open | LOW (argued) | — |
| VLN-044 | The Submitted page promised every submitter an approval step to watch and an outcome email — true for 1 form of 29. Gate written against a field that never existed | in progress — fixed, not yet deployed | MEDIUM (argued) | — |
| VLN-045 | /status/<id> is headed "Travel request." for every form — a Facilities Support Request's own status page names a workflow it is not in |
in progress — fixed 2026-08-26, not yet verified in a browser | LOW (argued) | — |
| VLN-046 | The confirmation page asserted "Ticket — NOT CREATED" while a ticket existed — CONFIRMED by #USFIN00367708. Copy fixed 2026-08-26 | in progress — fixed, not yet verified in a browser | LOW (argued) | — |
| VLN-047 | Migration 022's revise path is deployed and no UI constructs its URL — and /status/<id> asserts the request "cannot be edited" while offering a button that produces an unlinked resubmission |
done 2026-08-26 — acceptance test met; first revision 35750caf links to its parent |
MEDIUM (argued) | — |
| VLN-048 | The portal-coverage gate passes every cloud_services entity automatically — present = (sentinel in blob) or (eid.lower() in blob) and both portals carry the sentinel in a comment. Two live Cloud Run services sat unrendered for three weeks, gate green |
open | MEDIUM (argued) | — |
| VLN-049 | The travel form's served instructions told every submitter a returned request "cannot be edited" — the VLN-047 claim, live at version: 2, on the same column GOV-021 was raised about. A green test pinned it |
done 2026-08-27 — 4 surfaces fixed in the v3 bump; served column verified against gpus_forms | MEDIUM (argued) | — |
| VLN-050 | The HappyFox ticket body (travelrequest001) is a third, uncounted render surface that enumerates its fields — the v3 bump added three and never touched it. One real ticket (c84bfb45, Madison Carter) was rendered by the broken template 6m06s before the fix landed |
CLOSED 2026-09-01 — verified on a real ticket, #USITS00368946 |
MEDIUM (argued, raised from LOW) | — |
| VLN-051 | approval_decision_recorded / approval_notification_sent are in AUDIT_ACTIONS but in neither ship_actions nor excluded_actions — 31 of 33 declared. The non-repudiation audit row reaches no consumer |
open | MEDIUM (argued) | — |
| VLN-052 | approval_access's denial reasons are absent from authz_reasons.json, so a denied approval attempt falls to rule 100030 at level 3 — below the ≥ 10 auto-ticket threshold. A forged-approval attempt pages nobody |
open | MEDIUM (argued) | — |
| VLN-053 | CEDAR's log indexer serves the entire security-alert corpus with no authentication — Elasticsearch 8.19.13 on 172.16.0.13:9200, xpack.security.enabled: false, plain HTTP, no TLS. Anonymous _cat/_count/_search over 155 daily wazuh-alerts-* indices back to 2026-04-07 (339 indices total). Reachability CORRECTED 2026-09-08 (v1.1) — v1.0 said "anything that can reach the subnet" and counted 113 WDC workstations, reasoning from the GCP rule cedar-ingress alone. CEDAR's host firewall is narrower and is the effective control: 9200 is permitted from 172.16.0.12 (MAPLE) and 10.8.0.0/28 only; the WDC LAN gets 5140/tcp and ICMP, not 9200. The 10.8.0.0/28 end is the Serverless VPC connector, enumerated rather than assumed: 5 of 10 Cloud Run services attach to it (gpus-forms-backend, gpus-forms-clamav-worker, gpus-security-backend, gpus-soc-backend, gpus-status-backend), all ingress: all — so the distance from the internet to the corpus is one SSRF/RCE in any of five services, with no credential to steal. The cloud layer is still open to the 113 workstations and only firewalld withholds them, which is a latent trap if the host firewall is ever stopped. No audit logging is possible with security off, so the access is unattributable and the exposure cannot be scoped retrospectively. Write/delete are almost certainly equally unauthenticated — reasoned from the absent auth layer, deliberately not tested |
open | High (argued) | — |
| VLN-054 | Webmin on CEDAR and MAPLE is reachable only from the Cloud Run VPC connector — miniserv.pl runs as root, active and enabled at boot on both hosts, bound 0.0.0.0:10000 + [::]:10000. The host firewall on each permits 10000/tcp from 10.8.0.0/28 and nothing else — the Serverless VPC connector, where no administrator sits and 5 of 10 ingress: all Cloud Run services egress. The management network 192.168.124.0/24 gets 22/tcp (and Grafana on MAPLE) but not 10000, so no operator can reach Webmin over the network at all; real use goes via ssh -L to loopback, which needs no rule. The exposure therefore buys nothing and removing it costs nothing. Auth is real but single-layered: TLS enforced (ssl_enforce=2, HSTS, no SSLv2/3), session-based, passdelay=1, blockhost_failures=5 — but blockhost_time=60 is sixty seconds, there is no 2FA, and no allow= IP allow-list, so firewalld is again the only control. Also contradicts defense-in-depth.md, which claims all admin interfaces are management-network-only: Prometheus, Kibana and Webmin are not reachable from that network, and all four are reachable from a connector the doc never mentions. CEDAR and MAPLE are absent from access-control/roles.md entirely. Exactly one account exists, it is root, and its password field is x — Unix/PAM auth against /etc/shadow, confirmed. So the Webmin login is the host's root password: no lower-privilege account to compromise first, no second factor, and blockhost_failures=5/blockhost_time=60 permits ~5 guesses/minute indefinitely, offered to the origin five ingress: all Cloud Run services egress from. Established by printing field lengths and then a single character — no credential material left the host. …and on CEDAR that login cannot succeed. passwd -S root returns LK — the root password is locked in /etc/shadow, so pam_unix refuses every attempt. The sole account cannot authenticate: nobody can log into Webmin on CEDAR. Severity DOWNGRADED High → Medium 2026-09-08 on that check, after three revisions had published High. The chain x → "the login is the root password" was correct at every step and still wrong, because it stopped one layer short of the credential store — the same read-one-layer-and-stop shape recorded against this author in VLN-053 §4. What remains is a root-owned miniserv.pl, bound 0.0.0.0 + [::], enabled at boot, needlessly reachable from the connector, whose only live attack surface is pre-authentication — and a locked account is no defence there at all. Pre-auth question RESOLVED 2026-09-08 against the vendor advisory list: none published for 2.621. Every later fix (2.640, 2.641, 2.652, 2.653 — incl. CVE-2026-42210/56022 2FA bypass, CVE-2026-22678 XSS, an SSRF and a privesc) requires an authenticated session, so none is reachable while root cannot log in. The Critical trigger did not fire. What keeps it at Medium rather than Low: the mitigation is a state of /etc/shadow, not a control in Webmin. Setting a root password for any unrelated reason — recovery, vendor instruction, console rescue — silently re-opens a root login on a connector-facing port with no config change, no firewall change and no alert, and makes that entire authenticated-CVE backlog live in the same moment. "Safe because root happens to be locked" is not "safe because Webmin cannot authenticate that way". 2.621 is ≥4 security releases behind — unreachable here, live on any host with a usable login. MAPLE measured, not assumed, and it matches: passwd -S root → LK and cut -d: -f1,2 miniserv.users → root:x. All three facts hold on both hosts, so nobody can log into Webmin on either. Measured rather than inferred because a locked root only yields an unusable login if root is the sole account and delegates to Unix auth — a second account, or root with a real stored hash, would authenticate inside Webmin and never consult /etc/shadow. Parity on one input is not parity on the conjunction. Both hosts rate Medium. If §7.2 ever comes true and a root password is set, MAPLE's consequence is the worse of the two: it is the Wazuh manager, so a root compromise there also compromises the pipeline meant to notice a compromise. Remediation: disable the service (systemctl disable --now webmin), on three independent lines of evidence — nobody can log in (locked root, sole account), nothing in the repo consumes port 10000 while the other connector rules do have consumers, and miniserv.log records no logins at all (its only three entries are the author's own 401 probes, which also resolved a miniserv.conf mtime that had been flagged as possible third-party activity). Disabling also converts §7.2's state into a real control: with the service gone, setting a root password later cannot silently re-open anything. Check MAPLE's own log first rather than reusing CEDAR's |
open | Medium (both hosts; was High) | — |
VLN-013 — ESXi 6.7 EOL is on fire and flower; VLN-004 misattributes it to water¶
| Field | Value |
|---|---|
| id | VLN-013 |
| title | ESXi 6.7 EOL on fire and flower — VLN-004 misattributes this to water |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | Authenticated vSphere API, 2026-08-13: fire build 17700523 (6.7 U3); flower build 8169922 (6.7 GA, 2018, never patched); water build 24022510 (8.0.3). fire hosts sky, rain, sun and wind. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-esxi-inventory-20260813.txt. All three build numbers confirmed present in the capture: 17700523, 8169922, 24022510. |
| acceptance test | WIDENED 2026-08-17 per GOV-017 (B-22). No artifact asserts that water is ESXi 6.7 or that water hosts VMs — covering (a) VLN-004's row above, (b) inventory.yaml:146-152, and (c) any downstream generated output derived from either. (Original test, 2026-08-13: "VLN-004 corrected to name fire and flower; water removed from its scope." Superseded — it closed on one artifact and left the other two asserting the same wrong thing.) |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Note: flower is on the 2018 GA build and has never been patched. VLN-004 above
still reads WATER hypervisor; a scope-dispute marker has been added to that row
pointing here. VLN-004's text is otherwise unchanged pending this row's closure.
RESOLVED 2026-08-17 — acceptance test widened
Raised 2026-08-13: the same misattribution was found in a second place while
evidencing PRG-010 — inventory.yaml:149 records hypervisor: VMware
ESXi 6.7 for water, and :152 records vms: [ocean] on a host that
hosts zero VMs. The acceptance test then named only VLN-004, so
correcting VLN-004 would have closed this row while inventory.yaml — the
artifact three governance documents call the single source of truth — kept
asserting the same wrong thing. It was flagged rather than widened
unilaterally, and left as open item 2 in the 2026-08-13 session log.
Decided 2026-08-17: widened. The test now reads no artifact asserts
water is ESXi 6.7 or hosts VMs, covering VLN-004, inventory.yaml:146-152
and downstream generated output. This closes open item 2 from the
2026-08-13 session log.
Justification is GOV-017, which found the disagreement is broader than
water: flower is declared as ESXi 6.8, a release that does not exist,
and fire is declared as hosting nothing while carrying all four core WDC
servers. Those are separate rows and are not folded into this one.
VLN-014 — sky and rain serve different DNSSEC signings — REGRESSED 2026-08-19, REOPENED 2026-08-24¶
| Field | Value |
|---|---|
| id | VLN-014 |
| title | rain served a 2026-07-29 DNSSEC signing while sky served 2026-08-12; converged 2026-08-17, then diverged again on 2026-08-19 by the same mechanism |
| register | Security findings |
| status | in progress — remediated and verified 2026-08-17, regressed 2026-08-19, reopened 2026-08-24 |
| severity | High — supplied as "severity high, dated" |
| owner | R. Chhetry |
| evidence | Zone dumps from sky and rain, 2026-08-13, showing differing RRSIG inception and differing NSEC3 salt. SOA serial is 2026072902 on both — the re-sign never bumped it, so rain had no transfer trigger. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/zone-canon-sky-20260813.txt and raw/zone-canon-rain-20260813.txt. All four claims confirmed: SOA serial 2026072902 on both; RRSIG inception 20260812 on sky against 20260729 on rain, with rain carrying expiry 20260828; and the NSEC3 salts differ — sky A178C4BCAFEA2FA0, rain 83C78EB3B3503D4C. |
| acceptance test | Serial bumped; transfer confirmed; RRSIG inception and NSEC3 salt identical on both servers, verified by direct query. Met in full 2026-08-17; NOT MET as of 2026-08-24 — see the regression note. Unchanged: it is the right test, and it is what caught this. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | 2026-08-17 — superseded; re-queried 2026-08-24 and the condition no longer held |
REGRESSED 2026-08-19 — reopened 2026-08-24 by direct query against both hosts
The acceptance test is not met. Queried live on 2026-08-24, dig against
@localhost on each host:
| Check | sky | rain |
|---|---|---|
| SOA serial | 2026081702 |
2026081702 |
| RRSIG inception | 20260819063701 |
20260817170231 |
| RRSIG expiration | 20260918063701 |
20260916170231 |
| NSEC3 salt | CB18BB4A142868CF |
A178C4BCAFEA2FA0 |
| Signing keytag | 6660 |
6660 |
rain still holds exactly the values recorded as verified on 2026-08-17 —
inception 20260817170231, salt A178C4BCAFEA2FA0. The 08-17 verification
below was sound and is not withdrawn. sky has moved on without it.
This is the original fault, recurring — by the exact mechanism this row
documents. On sky, /var/named/wdc.us.gl3.db (the source) is unchanged
since Aug 17 14:01 and still carries serial 2026081701, while
/var/named/wdc.us.gl3.db.signed was rewritten Aug 19 03:37 and carries
2026081702 — the same served serial rain already had. Someone re-signed
on 08-19 without bumping the source serial first, -N INCREMENT derived
the identical serial again, and rain — correctly — saw no reason to transfer.
The remediation note below ends with the sentence "Omitting the source bump
reproduces the original fault silently." That is what happened, two days
after the verification and two days before the row was recorded done.
Why this row was done for three days
The 2026-08-17 verification was real, thorough, and correctly evidenced — rain's live values still match it exactly. The row was then written up on 2026-08-21 from that verification, not from a fresh query, and the regression had landed on 08-19 in the gap between the two.
Nothing was fabricated and no evidence was wrong. The failure is that
verified on the 17th was recorded as done on the 21st without
re-reading live state at the moment of writing. Under the status rule as
stated — verified end-to-end via the real production path against live
state — the verification had gone stale and the row should not have closed.
Close this row only from a query run at the time of closing, and prefer a standing check to a point-in-time one: this zone has now diverged twice by the same mechanism, undetected both times.
The 2026-08-17 verification — sound, and NOT withdrawn
Every clause of the acceptance test was met and observed, not inferred.
| Check | sky | rain |
|---|---|---|
| SOA serial | 2026081702 |
2026081702 |
| RRSIG inception | 20260817170231 |
20260817170231 |
| RRSIG expiration | 20260916170231 |
20260916170231 |
| NSEC3 salt | A178C4BCAFEA2FA0 |
A178C4BCAFEA2FA0 |
| Signing keytag | 6660 |
6660 (unchanged) |
rain flipped its salt from 83C78EB3B3503D4C to A178C4BCAFEA2FA0 — that
is the proof it took the new signing rather than merely agreeing by
coincidence, and it is exactly the discriminator the acceptance test was
written around.
Transfer confirmed in rain's own log: zone wdc.us.gl3/IN: transferred serial
2026081702 / Transfer status: success. delv reports fully validated
on both hosts. dsset md5 unchanged, so no parent DS update is needed.
Five untouched zones on sky verified identical, bounding the blast radius.
named PID unchanged — this was a reload, not a restart.
Root cause corrected — the recorded mechanism was the symptom
The row above states the cause as "the re-sign never bumped it, so rain had no transfer trigger." That is what was observed, not why it happened, and it is corrected here rather than left standing.
Actual mechanism: dnssec-signzone -N INCREMENT bumps the serial in the
output file only, leaving the source zone untouched. Every re-sign from an
unchanged source therefore reproduces the identical served serial — and a
secondary that declines to transfer an unchanged serial is behaving
correctly. rain was never faulty.
The 2026-07-29 and 2026-08-12 signings both derived 2026072902 from source
serial 2026072901. Two signings were lost this way before anyone
noticed, and nothing in the pipeline could have surfaced it: each re-sign
reported success, the output file was valid, and the secondary's refusal to
transfer was the correct response to an unchanged serial. Another instance of
the estate's characteristic pattern — see M-01, and note this one is
subtler: the control ran, did its job, and the failure was in the input.
Verified procedure — bump the source serial first, then:
cd /var/named && dnssec-signzone -3 A178C4BCAFEA2FA0 -H 10 \
-N INCREMENT -o wdc.us.gl3 -K /var/named wdc.us.gl3.db
Omitting the source bump reproduces the original fault silently.
VLN-015 — NOPASSWD sudo on emu and ostrich; VLN-009 scope extends beyond WDC¶
| Field | Value |
|---|---|
| id | VLN-015 |
| title | NOPASSWD sudo present on emu and ostrich under the gcloud-provisioned rchhetry identity — scope expansion of VLN-009 onto both DNS masters |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied. Not inherited from VLN-009 (CVSS 9.9): the hosts, identity and access path differ, and carrying the score across would be an assumption, not a measurement. |
| owner | unassigned |
| evidence | Enumerator 3 partial-access note, 2026-08-13. |
| acceptance test | Full sudoers audit across the gpusa estate; VLN-009 scope updated to match what the audit finds. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Both affected hosts are DNS masters. VLN-009's row above is scoped
SKY/RAIN/SUN/WIND and is not edited by this row — its scope changes only
once the sudoers audit in the acceptance test has run.
VLN-016 — One person holds three identity strings with inverted access across systems¶
| Field | Value |
|---|---|
| id | VLN-016 |
| title | Multiple identity strings for one person across systems, with inverted access |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | access-map.md v2 (6 probes, 2026-08-13); Meraki GET /administered/identities/me. Specifics: rchhetry@greenpeace.org holds Owner on gpusa (~25 VMs including both DNS masters) while documented as SSO-only, and is denied on gpus-infra. rajesh.chhetry@greenpeace.us is the inverse. The Meraki org 395909 admin authenticates as a third string, rajesh.chhetry@greenpeace.org. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): access-map.md — the artifact this row cited, which was recorded on 2026-08-13 as not present in the portal repo. It is in the enumeration corpus, headed "Access map — Stage 1, Step 0 (v2, post-reauth)", dated 2026-08-13, with raw output at raw/access-probe-both-accounts-20260813-v2.txt. Its coverage matrix confirms the inversion: gpus-infra READABLE by rajesh.chhetry@greenpeace.us and DENIED for rchhetry@greenpeace.org; gpusa-it-infrastructure-306400 and gpus-it-infrastructure exactly the inverse. The third identity string is corroborated independently — rajesh.chhetry@greenpeace.org appears as a Meraki org admin in raw/task1-meraki-20260813.txt. |
| acceptance test | Identity map recorded per system; a deliberate decision taken on whether to consolidate the identities or formally accept the split, and that decision encoded in the runbooks. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
The gap between "documented as SSO-only" and "holds Owner on ~25 VMs" is the material part: the documented access model and the effective one disagree.
VLN-017 — gcloud list verbs exit 0 on permission denial¶
| Field | Value |
|---|---|
| id | VLN-017 |
| title | gcloud list verbs return exit 0 on permission denial; the error appears only on stderr |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | All 6 access probes, both directions, 2026-08-13. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/access-probe-both-accounts-20260813-v2.txt. Confirmed verbatim: every probe line reads EXIT_CODE=0, including probes 2, 3 and 4, whose bodies carry Required 'compute.instances.list' permission for 'projects/…'. A denial and a success are indistinguishable by exit code in this capture. |
| acceptance test | A grep of gpus-infra-portals, the backup scripts and Cloud Build for any branch on $? following a gcloud list; each occurrence either fixed or documented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Same failure spine as VLN-010 and VLN-011: a control returns success while doing nothing, so every observable an operator would check comes back green.
VLN-018 — All 10 Cloud Run services in gpus-infra are ingress: all¶
| Field | Value |
|---|---|
| id | VLN-018 |
| title | All 10 Cloud Run services in gpus-infra are ingress: all, including 4 backends and gpus-forms-clamav-worker |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | Stage 1 raw JSON, task4-cloudrun, 2026-08-13. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task4-cloudrun-20260813.txt — the file this row names. Date qualification: that file is dated 2026-08-13 and states in its own header "Source: raw/gpus-infra__run__rajesh.chhetry-at-greenpeace.us.json (Stage 1, generated 2026-08-12T20:28:28Z). NO new API calls." The observation is therefore 2026-08-12; 2026-08-13 is the derivation date. Service count 10 confirmed. |
| acceptance test | Backends and the worker moved to internal-and-cloud-load-balancing, verified by an external request returning 403. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
VLN-019 — duck.cloud.us.gl3 resolves to an address that is not duck's¶
| Field | Value |
|---|---|
| id | VLN-019 |
| title | duck.cloud.us.gl3 resolves to an address that is not duck's |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | Enumerator 3, 2026-08-13. Corroborated from a second independent source 2026-08-17 — see PRG-024: gpus-dist forward and reverse zones both say 10.1.96.40, the live VM is 10.1.96.46, and inventory.yaml says .46. Two independent sources now agree the zone data is wrong and inventory.yaml is right. |
| acceptance test | Record corrected or removed; resolution verified by query. |
| blocks | — |
| blocked_by | — |
| corroborated_by | PRG-024 — stated in the source row (B-36). |
| date_raised | 2026-08-13 |
| date_verified | — |
VLN-020 — Five Puppet node definitions can never match a certname¶
| Field | Value |
|---|---|
| id | VLN-020 |
| title | Five Puppet node definitions use bare names that cannot match an FQDN certname — that config has never applied to anything |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | Enumerator 2, 2026-08-13. |
| acceptance test | Each of the five confirmed intentional, or removed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
Config that has never applied to anything is the same class as VLN-010's inert rules: present, loaded, counted, and never reached. See also PRG-007 — 18 orphan Puppet node definitions, likely the same uncleaned decommission wave.
VLN-021 — 3 of 4 Meraki org admins hold orgAccess: full on the VPN-terminating org¶
| Field | Value |
|---|---|
| id | VLN-021 |
| title | 3 of 4 Meraki org admins hold orgAccess: full on the organisation whose appliance terminates the WDC VPN |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | GET /organizations/395909/admins, 2026-08-13. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task1-meraki-20260813.txt. Confirmed, with the admins named: msedita@greenpeace.org, rajesh.chhetry@greenpeace.org and tanu.garg@greenpeace.org all carry orgAccess: full. All three show authMethod: Email and 2FA: True; only rajesh.chhetry@greenpeace.org has hasApiKey: True. The capture notes the querying key is itself org-full-access, authorised for GET-only use. |
| acceptance test | Access levels reviewed and reduced to least privilege, or each full-access grant justified in writing. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-13 |
| date_verified | — |
VLN-022 — Published forms go-live record asserts a delivery claim VLN-011 disproves¶
| Field | Value |
|---|---|
| id | VLN-022 |
| title | Published forms go-live record asserts all 6 ingest addresses verified delivering and HappyFox tickets opening; the claim is false and is live on infra.greenpeace.us |
| register | Security findings |
| status | open |
| severity | not assessed — no severity supplied; not inferred |
| owner | unassigned |
| evidence | Flag list, 2026-08-13 (priorities/flag-list-unverified-done-2026-08-13.md, Tier A2); VLN-011 finding page, which establishes 0 of 37 submissions received a ticket id and that no HappyFox client exists in the repo. |
| acceptance test | Published record corrected; the claim replaced with an evidence-backed status. |
| blocks | — |
| blocked_by | GOV-009 — the correction cannot be confirmed live until the portal render can be verified |
| date_raised | 2026-08-13 |
| date_verified | — |
This is the only row here whose harm is currently published: the false claim
is being served to readers of infra.greenpeace.us for as long as it stands.
VLN-023 — vmstorage running a degraded RAID 6 with no spare¶
| Field | Value |
|---|---|
| id | VLN-023 |
| title | vmstorage (RS1619xs+, 192.168.120.51) running a degraded RAID 6 — the NFS datastore for every WDC VM |
| register | Security findings |
| status | in progress — remediation reported; not done, see below |
| severity | CRITICAL (as stated) |
| owner | R. Chhetry |
| evidence | Synology DSM API, 2026-08-17. designedDiskCount 6, normalDevCount 5, disk_failure_number 1, spares [], repair_action degrade, data scrubbing unavailable (reason: abnormal). Three disks sit in not_use — sdeb, sdec, sdfa — all reporting SMART normal. This is the NFS datastore for every WDC VM, mounted by fire, water and flower. Nothing in monitoring surfaced it and no artifact mentions it. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed: model=RS1619xs+, volume_1 … status=degrade, and the three unused disks sdeb, sdec, sdfa all reading status=not_use smart=normal. Note the capture places those three in pool RX1217rp-1 while the four normal disks sit in pool RS1619xs+ — an expansion-shelf detail not in the row text and not inferred into it. |
| acceptance test | Array returns normal with a spare assigned and scrubbing available. |
| blocks | — |
| blocked_by | — |
| related_to | VLN-026 — derived. No backup or replication compounds the degraded array. |
| date_raised | 2026-08-17 |
| date_verified | — |
Spare reported assigned — remediation applied, NOT verified
R. Chhetry reports a spare assigned on vmstorage, recorded 2026-08-20.
The action date itself was not stated and is not inferred here. The row
is moved to in progress, not done, and the acceptance test is
unchanged: array returns normal with a spare assigned and scrubbing
available — three conditions, of which one is reported and none is
re-observed.
Under the status rule in force, an action taken is not a state verified.
That distinction is the whole reason this register exists: Tier A of the
2026-08-13 flag list is two rows that recorded an action and were read ever
after as a verified outcome.
What closes it — re-run the same read-only probe that raised the finding and compare:
Then confirm three things on vmstorage (192.168.120.51), all of which
read the other way in the 2026-08-13 capture:
| Field | 2026-08-13 | Required to close |
|---|---|---|
volume_1 status |
degrade |
normal |
| spare | spares: [] |
a spare assigned |
| data scrubbing | unavailable (reason: abnormal) |
available |
Note a rebuild takes time on a 26.79 TB volume — status=normal may not
appear immediately, and a volume mid-rebuild is still a volume without
redundancy. syno_probe.py is read-only by construction (MutatingCallBlocked
guards every call), so re-running it is safe during a rebuild.
VLN-026 (no backup or replication task) and GOV-019 (snapshot config
unknown, 403) are untouched by this and both still stand.
A RAID 6 down one disk with zero spares, on the datastore under every WDC VM,
with three unused disks physically present and reporting healthy. The
three not_use disks are why the missing spare is notable rather than merely
unlucky.
VLN-024 — Plaintext vendor credentials in the Puppet control repo¶
| Field | Value |
|---|---|
| id | VLN-024 |
| title | Plaintext vendor credentials for payment, payroll and CRM vendors in modules/relay/files/etc/pushtab |
| register | Security findings |
| status | open |
| severity | CRITICAL (as stated) |
| owner | R. Chhetry |
| evidence | modules/relay/files/etc/pushtab lines 1–6. Six rows, format name user password endpoint — field 3 is a plaintext password on every row, covering payment, payroll and CRM vendors (litle, convio, adp, target). Last commit 2019-01-24. Deployed to /opt/gp/etc/pushtab on every storehouse node. The repo has 879 commits and five known working copies. Two values also transited terminal output on 2026-08-17. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage3b-credential-locations.json. Confirmed exactly: six entries, lines 1–6, all "type": "plaintext-password" at /etc/puppetlabs/code/environments/production/modules/relay/files/etc/pushtab, field 3, each "last_commit": "2019-01-24". The capture records value_length_bucket rather than the value itself — the enumeration did not copy the secrets into the artifact. |
| acceptance test | All six rotated at the vendor, or the accounts disabled vendor-side, confirmed per vendor. |
| blocks | — |
| blocked_by | — |
| related_to | VLN-025 — derived. puppetCrypt exists and pushtab bypasses it. |
| date_raised | 2026-08-17 |
| date_verified | — |
DISPOSITION recorded 2026-08-20 — deferred to the Rocky migration. Read the residual risk.
Decision (R. Chhetry, recorded 2026-08-20): this closes when the servers move
from RHEL to Rocky and the vendors are moved there under Terraform. Recorded
as a dated decision. Status stays open — a decision is not a
verification, and the acceptance test is unchanged.
The migration does not, by itself, close this row. The acceptance test
is all six rotated at the vendor, or the accounts disabled vendor-side,
confirmed per vendor — and that is deliberate. Standing up new
Terraform-managed credentials on Rocky hosts creates a second working set;
it does not invalidate the first. The six values in pushtab stay valid at
litle, convio, adp and target until someone rotates or disables them at the
vendor. The migration must therefore include an explicit vendor-side
rotation or disablement step, per vendor, or this row survives it.
Exposure accepted by this deferral, stated so the decision is made on the full picture rather than re-litigated later:
- The credentials have been in plaintext since 2019-01-24 — over seven years — across 879 commits and five known working copies.
- They are deployed to
/opt/gp/etc/pushtabon every storehouse node. - Two of the six values transited terminal output on 2026-08-17.
- They cover payment, payroll and CRM vendors.
PRG-019decommissions the code that reads the file. It does not decommission the credentials.
The exposure window is now the migration timeline. If that timeline is
longer than a few weeks, the cheaper interim move is to disable vendor-side
any of the six that are not currently required — which is checkable against
PRG-019's finding that only 13 of 27 cron resources are potentially
active, and that pushFacter, pushFPR and pushPSI transfer nothing at
all.
Not closed by the Rocky migration, and not closed by deleting the file
Recorded as supplied: deleting the file does not invalidate the
credentials. Five known working copies, 879 commits of history, and two
values already in terminal output. The acceptance test is vendor-side
rotation or disablement confirmed per vendor — nothing repo-side closes
this row. Note PRG-019 disposes of the transfer stack as DECOMMISSION; that
disposition does not close this row either.
VLN-025 — puppetCrypt exists and pushtab bypasses it¶
| Field | Value |
|---|---|
| id | VLN-025 |
| title | puppetCrypt exists (125 GPG-encrypted secrets) and pushtab bypasses it — the capability was present and the file ignored it |
| register | Security findings |
| status | open |
| severity | HIGH (as stated) |
| owner | R. Chhetry |
| evidence | ⚠ Partial. Supplied: puppetCrypt exists and holds 125 GPG-encrypted secrets; pushtab does not use it. The count is specific, but no source artifact, path or read date was stated for this row. To complete: the puppetCrypt path and a dated listing. ⚠ Still unconfirmed after searching the enumeration corpus 2026-08-17. No artifact stating the 125-GPG-secret count was located under ~/estate-enum-2026-08; raw/stage3b-credential-locations.json records the pushtab plaintext rows only and says nothing about puppetCrypt. The count remains unsourced and this row stays partial. |
| acceptance test | No plaintext credential in the control repo. |
| blocks | — |
| blocked_by | — |
| related_to | VLN-024 — derived. |
| date_raised | 2026-08-17 |
| date_verified | — |
The material point is that this was not a missing capability. Encryption was available in the same repo and the file did not use it, which makes this a process gap rather than a tooling gap.
VLN-026 — vmstorage has no backup or replication task¶
| Field | Value |
|---|---|
| id | VLN-026 |
| title | vmstorage has no backup or replication task — configured-absent, not unchecked |
| register | Security findings |
| status | open |
| severity | HIGH (as stated) |
| owner | R. Chhetry |
| evidence | Synology DSM API, 2026-08-17. SYNO.Backup.Task returns 0; SYNO.Backup.Repository repo_list is empty; CloudSync API not registered. Recorded as configured-absent, not unchecked — the queries succeeded and returned nothing, rather than failing to run. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed verbatim: Hyper Backup tasks: 0; Backup repositories: {"offset": 0, "repo_list": [], "total": 0}; CloudSync: FAILED err=103 (API not available/registered on unit). The configured-absent reading holds — these are successful queries returning empty, not failures. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. Note that a close condition cannot be stated responsibly while GOV-019 stands: snapshot configuration on this unit is unknown (API 403), so what protection exists is not yet established. |
| blocks | — |
| blocked_by | GOV-019 — derived. Snapshot configuration unknown blocks assessment of this row's severity. |
| related_to | VLN-023 — derived. Compounds the degraded array. |
| date_raised | 2026-08-17 |
| date_verified | — |
The distinction between configured-absent and unchecked is load-bearing and
is preserved as supplied: this is a positive finding of absence, not a gap in
the enumeration. GOV-019 is the gap in the enumeration, and it sits next to
this row rather than inside it.
VLN-027 — Meraki IPsec pre-shared keys returned in plaintext¶
| Field | Value |
|---|---|
| id | VLN-027 |
| title | IPsec pre-shared keys for both VPN peers returned in plaintext by the Meraki thirdPartyVPNPeers endpoint |
| register | Security findings |
| status | open |
| severity | HIGH (as stated) |
| owner | R. Chhetry |
| evidence | Meraki thirdPartyVPNPeers endpoint returned both peers' PSKs in plaintext; the values transited terminal output on 2026-08-13. |
| acceptance test | Both rotated, coordinated with the GCP side, tunnel verified up. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Durable handling rule, recorded as supplied
Any read of Meraki VPN config exposes PSKs. This is a standing property
of the endpoint, not a one-off. It applies to every future enumeration that
touches Meraki VPN configuration, including any reconciliation control built
per GOV-004.
VLN-028 — SMBv1 enabled with signing off on all three Synology units¶
| Field | Value |
|---|---|
| id | VLN-028 |
| title | SMBv1 enabled on all three Synology units with server signing off |
| register | Security findings |
| status | open |
| severity | HIGH (as stated) |
| owner | R. Chhetry |
| evidence | Synology DSM API v3, 2026-08-17. smb_min_protocol = 1 on all three units — the lowest selectable value, corresponding to SMB1/NT1 — with server signing off (0) on all three. NTLMv1 is off. duck and grebe are Windows SMB proxies. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed: ntlmv1_auth=False server_signing=0 workgroup=WORKGROUP. |
| acceptance test | Minimum protocol raised, signing enabled, verified per unit. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
NTLMv1 being off is recorded as supplied and is not treated as mitigating: it narrows the finding, it does not close it.
VLN-029 — Key material in a control repo scheduled for deletion¶
| Field | Value |
|---|---|
| id | VLN-029 |
| title | Key material in the control repo, identified by filename only — in a project scheduled for deletion |
| register | Security findings |
| status | open |
| severity | HIGH (as stated) |
| owner | R. Chhetry |
| evidence | ⚠ Partial. Supplied by filename count only — contents never read: ca/ 81 private-key files, dnssec/ 1 .private, htpasswd/ 18 hash files. The repo lives in a project scheduled for deletion. No read date was stated for the listing. To complete: the dated directory listing the counts came from. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/stage3b-dist-index.json. One of three counts reproduces. dnssec .private files = 1 ✅ as stated. ⚠ Discrepancy — figure did not reproduce. The index lists 19 files under htpasswd/, not 18, and 643 files under ca/ against a stated 81 — 643 is every file under ca/ (the sample includes ca/bin/p7x and ca/certs/*.crt), so the 81 is presumably a private-key subset whose filter is not recorded. The stated figures are left unchanged; the filter needs stating before either can be confirmed. |
| acceptance test | Each item determined live or historical; anything live rotated or relocated. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Recorded as supplied: deleting the repo copy does not revoke anything still
trusted. Same shape as VLN-024 — removal of the artifact is not remediation
of the secret. Note GOV-014 requires mirroring these repos before any
teardown touches those projects, which this row depends on in practice.
VLN-030 — Undocumented second VPN peer¶
| Field | Value |
|---|---|
| id | VLN-030 |
| title | Undocumented second VPN peer Target (206.31.252.84), IKEv1, subnet 10.5.142.0/24 |
| register | Security findings |
| status | open |
| severity | MEDIUM (as stated) |
| owner | R. Chhetry |
| evidence | Meraki API, 2026-08-13. Peer Target, 206.31.252.84, IKEv1 (deprecated), subnet 10.5.142.0/24. Appears in no artifact, no brief and no diagram. |
| acceptance test | Owner identified; tunnel justified in writing or removed. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
target also appears as one of the four vendors carrying a plaintext credential
in VLN-024. The two are recorded separately and no relation is asserted —
the name matching is not evidence that they are the same counterparty.
VLN-031 — Zero 2FA across all three Synology units¶
| Field | Value |
|---|---|
| id | VLN-031 |
| title | Zero 2FA across all three Synology units; admin account present and not renamed |
| register | Security findings |
| status | open |
| severity | MEDIUM (as stated) |
| owner | R. Chhetry |
| evidence | Synology DSM API, 2026-08-17. 2fa_status: false on every user on every unit. The admin account is present and not renamed, though expired/disabled. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt. Confirmed: 2fa_status=False on every enumerated user, and the admin account present with expired=now and desc='System default user' — present and not renamed, as stated. Note rchhetry reads 2fa_status=None rather than False, i.e. unreported rather than disabled. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
VLN-032 — Meraki org 395909 has no SSO and no IdP¶
| Field | Value |
|---|---|
| id | VLN-032 |
| title | Meraki org 395909: SAML disabled, zero IdPs, all four admins on email authentication |
| register | Security findings |
| status | open |
| severity | MEDIUM (as stated) |
| owner | R. Chhetry |
| evidence | Meraki API, 2026-08-13. saml.enabled false, zero IdPs configured, all four admins authenticationMethod: Email. Only one admin holds an API key. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task1-meraki-20260813.txt. The admins table confirms authMethod: Email across the org admins and a single hasApiKey: True. The capture also notes the admins endpoint returns no saml field at all, so the saml.enabled false claim rests on a separate call, not this table. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Recorded as supplied: the W2 SSO migration has not started, rather than being
partially done. That distinction is the finding — a partially-migrated org and
an untouched one look similar from a status report and are not the same risk.
Read alongside VLN-021: 3 of the 4 admins on this org hold orgAccess: full,
and this is the org whose appliance terminates the WDC VPN.
VLN-033 — synstorage volumes at 93.1% and 82.3%, status attention¶
| Field | Value |
|---|---|
| id | VLN-033 |
| title | synstorage volume_1 at 93.1% and volume_3 at 82.3%, both status attention, no documented backup |
| register | Security findings |
| status | open |
| severity | MEDIUM (as stated) |
| owner | R. Chhetry |
| evidence | Synology DSM API, 2026-08-17. volume_1 93.1% (13.44 TB) and volume_3 82.3% (187.64 TB), both status attention. Media/Video data with no documented backup. Artifact located and re-verified 2026-08-17 (enumeration corpus now reachable at ~/estate-enum-2026-08): raw/task2-synology-20260813.txt covers all three Synology units in one capture. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
VLN-034 — rooster cannot compile in production: archive::backup body commented out¶
| Field | Value |
|---|---|
| id | VLN-034 |
| title | rooster.cloud.us.gl3 cannot compile in production — archive::backup has its entire body commented out |
| register | Security findings |
| status | open |
| severity | MEDIUM (as stated) |
| owner | R. Chhetry |
| evidence | Stage 3 manifest read, 2026-08-17. archive::backup body commented out at lines 5–17 in the production environment, but defined normally in seed and testing. rooster is absent from PuppetDB, which is consistent with a host that cannot compile. Children archive::backup::files and archive::backup::cron still exist and are now unreachable. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
Third instance of LOADED ≠ REACHABLE
Recorded as supplied. The other two are VLN-010 (Wazuh baseline rules
loaded, valid, counted and never reached) and the transfer crons in
PRG-019 firing on schedule doing nothing. See M-01 in the Method
learnings section of the gap register — this is named there as the estate's
characteristic failure mode.
VLN-035 — dataflow.us.gl3 resolves to a host that does not exist¶
| Field | Value |
|---|---|
| id | VLN-035 |
| title | dataflow.us.gl3 resolves to merlin, which has no VM in any project and is ASSERTED-STALE |
| register | Security findings |
| status | open |
| severity | MEDIUM (as stated) |
| owner | R. Chhetry |
| evidence | zones/us.gl3.zone:61 — dataflow.us.gl3 resolves to merlin. merlin has no VM in any project and is classified ASSERTED-STALE. Every live vendor script posted status there. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| evidences | PRG-019 — stated in the source row. |
| date_raised | 2026-08-17 |
| date_verified | — |
Downgraded as supplied, and the downgrade is recorded rather than applied
silently. The transfer stack is DECOMMISSION per R. Chhetry 2026-08-17
(PRG-019), so this is evidence supporting the decommission rather than an
active incident. Severity is carried at MEDIUM as stated; the downgrade is in
the framing, not in the severity field, because no revised severity was given.
VLN-036 — wdc-wap-5 status alerting¶
| Field | Value |
|---|---|
| id | VLN-036 |
| title | wdc-wap-5 (MR57, WDC) status alerting |
| register | Security findings |
| status | open |
| severity | LOW (as stated) |
| owner | R. Chhetry |
| evidence | Meraki API, 2026-08-13. |
| acceptance test | ⚠ INCOMPLETE — none supplied. Not invented. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-17 |
| date_verified | — |
VLN-037 — Four production security reports ran on EOL Python 3.6.8¶
| Field | Value |
|---|---|
| id | VLN-037 |
| title | The weekly, monthly, newsletter and quarterly security reports executed on MAPLE's system Python 3.6.8 — end-of-life since 2021-12-23 — from deployment (~2026-04) until 2026-08-20 |
| register | Security findings |
| status | remediated |
| severity | HIGH |
| owner | R. Chhetry |
| evidence | On MAPLE, 2026-08-20: python3 --version → Python 3.6.8; ./venv/bin/python3 --version → Python 3.6.8 for a venv built with bare python3; python3.11 --version → Python 3.11.13 (present but unused by the reports). /usr/bin/python3 -m pip freeze returned the interpreter's report dependencies: reportlab==3.6.8, Pillow==8.4.0, requests==2.27.1, google-cloud-storage==2.0.0 — all 3.6-era, none installable on 3.11. |
| acceptance test | /opt/gpus-reports/venv/bin/python3 --version reports 3.11.x, and all four SOC reports render from that venv. Both executed 2026-08-20: venv rebuilt from python3.11 (3.11.13); weekly, monthly, newsletter, quarterly each rendered OK. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-20 |
| date_verified | 2026-08-20 |
Why this is a security finding and not a hygiene note. Python 3.6 stopped
receiving security patches on 2021-12-23, so the interpreter and every one of
the four pinned dependencies above had been outside their support window for
over four and a half years. Pillow 8.4.0 and requests 2.27.1 in
particular are image-parsing and HTTP-client code with published CVEs in the
intervening releases. The exposure is bounded — these are scheduled local jobs
processing data from the estate's own SOC API, not an internet-facing
listener — which is why this is HIGH rather than CRITICAL. It is not LOW: the
process runs as root from the root crontab, and report_generator.py
parses remote JSON and renders it through reportlab and Pillow.
How it was found, which is the part worth keeping. Nothing detected this. No scanner flagged it, no dashboard showed it, and the four reports had been arriving on schedule and rendering correctly the entire time — the only observable signal was a working report. It surfaced incidentally on 2026-08-20 while building a venv for a fifth report (R2, the daily travel approval exceptions report): R2's Cloud SQL dependencies require Python ≥ 3.8, which forced the interpreter version to be looked at for the first time since deployment.
This is the estate's recurring failure mode, again. Cf. VLN-010 (Wazuh
rules loaded, valid, counted, and unreachable) and VLN-011 (routing recorded
success=true while dispatching nothing): a component whose every observable
says healthy while the substance is wrong. An EOL interpreter produces
correct output, so correct output is not evidence of a supported runtime.
Remediated 2026-08-20. All five reports now run from
/opt/gpus-reports/venv, built from python3.11 (3.11.13) with a pinned
requirements.txt — reportlab==5.0.1, Pillow==12.3.0, requests==2.34.2,
google-cloud-storage==3.13.1. report_cron.sh invokes that interpreter
explicitly and refuses to fall back to system python3. The routing
worker's venv was deliberately not reused, so a report dependency change
cannot break mail routing.
Residual, and it is not this row's to close. MAPLE's system Python is
still 3.6.8 and is still the default python3 for anything else on that
host — root cron entries, ad-hoc scripts, /usr/local/bin/gpus-cloud-backup.sh.
This finding covers the reports only; that wider scope was not assessed here
and is not claimed to be clean.
Raised as owned work, not left as a residual field:
PRG-027— Enumerate every cron job, systemd unit and script on MAPLE that invokes the EOL 3.6.8 interpreter. Owner R. Chhetry, raised 2026-08-20, with an acceptance test.A residual has no owner and no acceptance test, so it reads as closed to anyone scanning the register — which is the same failure mode as this finding itself.
VLN-038 — No Cloud Build in the estate ran unit tests¶
| Field | Value |
|---|---|
| id | VLN-038 |
| title | Every Cloud Build trigger in the estate built and deployed without executing a single unit test; the only security scans were \|\| true. Controls documented as "enforced by a test" were enforced by nothing automatic. |
| register | Security findings |
| status | in progress — one gate added 2026-08-24, six suites and two components still ungated |
| severity | High — argued below, not supplied |
| owner | R. Chhetry |
| evidence | Confirmed first-hand 2026-08-24 by reading all eleven cloudbuild.yaml files. A grep for every test runner (unittest, pytest, tox, npm test, vitest, jest, go test) across all of them returns exactly one line — forms-backend/cloudbuild.yaml:34, added the same day. Before that: zero. The only security scanning is in forms-backend and gpus-forms-clamav-worker, and every line of it ends \|\| true: bandit -q -r forms-backend/ -x forms-backend/tests -ll \|\| true and pip-audit --requirement forms-backend/requirements.txt --strict \|\| true. Meanwhile 450 tests exist and pass — forms-backend 208, gpus-forms-routing-worker 91, gpus-reports 151 — of which 5 run in CI. (Was 381 when this row was raised on 2026-08-24; gpus-reports gained 69 the same day. See the regression note.) |
| acceptance test | scripts/check-test-gates.sh exits 0. It enumerates every directory containing test_*.py, and reports a gap when that directory has no cloudbuild.yaml, or has one that invokes no test runner, or runs only some of its suites. Re-runnable, so it confirms the gap stays closed rather than certifying a one-time review. Run 2026-08-24: 3 gaps, exit 1. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-24 |
| date_verified | — |
Every observable was green and nothing was being verified — the VLN-010 / VLN-011 / VLN-037 shape
This is the fourth instance of one pattern this quarter, and the pattern is the finding rather than any single case:
- VLN-010 — the rules loaded, the count was correct, validation passed. Three baseline rules had never matched an event. Only a trace revealed it.
- VLN-011 — the dispatcher recorded
success=trueand continued. No HappyFox client existed anywhere in the repo; 0 of 37 submissions ever got a ticket, while the UI told submitters the team had been notified. - VLN-037 — four reports rendered correctly for four months on an interpreter that had been end-of-life since 2021.
- VLN-038 (this row) — builds went green, and "203/203 tests pass" was reported repeatedly during this workstream. Both statements were true. Neither was produced by the pipeline.
In every case the control ran, reported success, and verified nothing of substance. A test suite that executes only when a person remembers to run it is indistinguishable, from outside, from a test suite that does not exist — and the thing that makes it dangerous is that it is more reassuring than having no tests, because the count is quotable.
Why High: at least one documented security control is unenforced
Severity is argued, not supplied. The case for High does not rest on the missing gate itself — it rests on what was resting on it.
gpus-reports/test_approval_report.py::TestNoApproverText carries this
docstring:
"
approval_stepscarries a free-text column holding an approver's own words about someone's travel — often why it was declined.maple-agentcan read it. NOTHING IN THE DATABASE STOPS THIS REPORT FROM PRINTING IT. These tests are what stops it. They were the condition on which the grant decision was answered 'no grant needed'; weakening one does not just fail a test, it removes the control."
A database grant was declined on the strength of that test. The test
lives in gpus-reports/, which has no cloudbuild.yaml at all — it
deploys by gsutil to MAPLE. So the control that replaced a grant has never
been executed by anything except a person choosing to run it. The same holds
for forms-backend/approval_access.py, whose module docstring names
test_membership_removal_does_not_revoke as the guard on an authorization
boundary that "nothing below the application" enforces.
The counter-argument, stated fairly: nothing is currently broken. All
381 tests pass today, the controls hold, and there is no exploitation path —
this is a latent verification gap, not a live compromise, which is why the
row is in progress rather than an incident. The reason it still reads High
is that the failure mode is silent and unbounded: nothing would report
the day one of those tests started failing, and the estate has now been
surprised three times by exactly that.
2026-08-24 — this got WORSE, not better, on the day it was raised
R5 (the daily travel volume report) shipped to gpus-reports/ the same day
this row was raised. It added 69 tests to a component with no build
pipeline at all, taking it from 82 to 151 and the estate from 381 to
450. The number running in CI did not move: still 5.
The largest gap in this row grew by 84% while the row was open, and nothing about that was careless — R5 was tested more thoroughly than most things in the estate, which is precisely the point. Writing good tests into a component with no pipeline produces more unexecuted assurance, not less. The count is quotable either way.
Recorded rather than fixed, deliberately: adding a pipeline to
gpus-reports is the decision described below, and taking it under
schedule pressure from a report deploy is how it would get taken badly.
Consequence for the exception date — date it by the DECISION, not the difficulty
When VLN-038's exceptions are written (see the acceptance test), the
instinct will be to give gpus-reports the longest date because it is the
hardest gap to close. That is backwards and it should be resisted.
gpus-reports holds
test_approval_report.py::TestNoApproverText — the control accepted in
place of a database grant — and now 151 tests, none of which anything
executes. It is simultaneously the largest exposure in this row and the
hardest to close, and dating by difficulty puts the longest clock on
the biggest problem. It just grew by 69 tests, which makes the mismatch
worse.
The date should be set by when the DECISION is due, not by when a
pipeline could be built. The decision is a narrow one and does not need a
pipeline to exist first: should gpus-reports and
gpus-forms-routing-worker have Cloud Build triggers at all, given both
ship by gsutil and neither is a Cloud Run service? That question can be
answered in an afternoon. Building whatever the answer implies can take as
long as it takes, on its own row.
Conflating the two is what would produce a twelve-month exception on the estate's largest block of unexecuted tests, and it would look reasonable the whole time.
Enumeration — deliberately NOT fixed in this pass
Adding test gates to builds that have never had them will break builds, and each needs its own decision about what "green" should mean. Recorded so the decisions can be made one at a time:
| Component | Build | Tests present | Blocking checks today |
|---|---|---|---|
forms-backend |
✅ | 208 | field-exposure-check; travel-instructions-check (1 of 7 suites) |
gpus-forms-routing-worker |
❌ none | 91 | nothing — no pipeline exists |
gpus-reports |
❌ none | 151 | nothing — deploys by gsutil |
gpus-forms-clamav-worker |
✅ | 0 | scan only, \|\| true |
forms-frontend |
✅ | 0 | none |
soc-backend |
✅ | 0 | coverage-check |
status-backend |
✅ | 0 | coverage-check |
soc-site |
✅ | 0 | coverage-check |
status-site |
✅ | 0 | coverage-check |
mkdocs-portal |
✅ | 0 | coverage-check |
security-backend |
✅ | 0 | none |
security-site |
✅ | 0 | none |
forms-backend/migrate |
✅ | 0 | none |
Two distinct problems, and they need different fixes. Six forms-backend
suites are ungated inside a build that now runs one — cheap to close, and
the reason the current gate is PARTIAL rather than green. Two components
with 242 tests between them have no pipeline at all — that is not a
cloudbuild edit, it is a decision about whether gpus-reports and the
routing worker should have builds, and it is the larger of the two.
The coverage-check steps are real blocking gates and are not criticised
here; they verify inventory coverage, which is a different question from
whether the code works.
VLN-039 — Pre-build check steps couple every deploy to PyPI availability¶
| Field | Value |
|---|---|
| id | VLN-039 |
| title | Cloud Build steps make network calls at build time with default retry settings, so a transient upstream timeout fails a build that has nothing wrong with it. Raised for pre-build pip install from PyPI; widened 2026-08-27 after the same shape appeared at the image push step, which no pip setting reaches |
| register | Security findings |
| status | in progress — pip half mitigated on forms-backend 2026-08-24, six other builds unchanged; the registry-push half is unmitigated everywhere (second instance 2026-08-27) |
| severity | Medium — argued below, not supplied |
| owner | R. Chhetry |
| evidence | Observed, not predicted. Build 8356f08a-6bf1-49eb-b38a-1e74ee6df218 (2026-08-24, gpus-forms-backend-trigger, commit 934c37f) failed with pip._vendor.urllib3.exceptions.ReadTimeoutError: HTTPSConnectionPool(host='files.pythonhosted.org', port=443): Read timed out. in step 0, field-exposure-check, during pip install --quiet pyyaml. Only step 0 started and finished — the enumerated step list contains that one entry and no other. A re-run of the same trigger at the same commit (a6637d70) passed with no code change. Across the estate, 7 of 11 cloudbuild.yaml files run pip install in a pre-build step — 11 install commands in total: forms-backend 3 (pyyaml; requirements.txt, 21 packages; bandit + pip-audit), soc-site 2, status-site 2, and mkdocs-portal, soc-backend, status-backend, gpus-forms-clamav-worker 1 each. |
| acceptance test | No Cloud Build step that gates a deploy fails on a transient third-party network error. Two halves, because the second instance proved one is not enough. (a) Dependency installs — every pre-build pip install in */cloudbuild.yaml carries explicit --retries and --timeout (grep, re-runnable), and no build in a 30-day window shows a ReadTimeoutError/ConnectionError from files.pythonhosted.org. (b) Registry I/O — no build in the same window fails on an i/o timeout or connection reset against *.pkg.dev, or a documented retry exists for the push/deploy steps. (b) cannot be satisfied by any pip flag and was not covered by the original test. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-24 |
| date_verified | — |
SECOND INSTANCE 2026-08-27 — different step, and pip settings do not reach it
Build fc88f7a8-1470-4f38-bc0b-9c66e7997f7e (gpus-deploy-mkdocs-portal,
commit 9c32983) failed at step 4, the image push — not at a pip
install, and not at a pre-build check:
Step #4: 3a98051e9797: Pushed
Step #4: 7da9fcd3e4f6: Pushed
Step #4: 746b97c19b23: Pushed
Step #4: Head "https://us-central1-docker.pkg.dev/v2/gpus-infra/gpus-images/
mkdocs/blobs/sha256:abf8772a…": dial tcp 142.250.125.82:443: i/o timeout
ERROR: build step 4 "gcr.io/cloud-builders/docker" failed: step exited with
non-zero status: 1
The image built. Every layer pushed. The final manifest HEAD against Artifact Registry timed out. Re-running the trigger at the same commit succeeded with no change — the same signature as the 2026-08-24 PyPI failure, one step later in the pipeline.
This is what widens the row rather than repeating it. The finding was
framed as "pre-build check steps couple every deploy to PyPI", and its
acceptance test was written entirely in those terms — --retries and
--timeout on every pip install. Every one of those mitigations was
already in place for this build and none of them could have helped, because
the call that failed was docker push to *.pkg.dev. Build-time network
fragility is not confined to dependency installs; the row's original scope
was the shape of the first instance, not the shape of the problem.
Note also which direction the exposure runs: this failure is safer than the pip one in that nothing shipped — but the deploy was blocked by an event with no relationship to the change, which is the row's actual harm.
The status is not the finding — reading it alone would have been wrong
Recorded because it nearly produced a false report.
The commit touched inventory.yaml, both portals, and five documents, and
fired five triggers. Four succeeded; gpus-deploy-mkdocs-portal showed
FAILURE. On a commit that had just added four inventory entries and a
supersession block to a published doc, a mkdocs build failure reads
exactly like a content regression — a malformed table, a broken link, a
bad YAML front-matter block.
It was none of those. The coverage check in step 0 had already passed
(Errors: 0, RESULT: PASS), MkDocs had built the site, the image was
tagged, and the failure was a TCP timeout two steps later. Only the log
distinguishes those two stories, and they call for opposite responses:
revert-and-fix versus re-run unchanged.
This is the same lesson as the frontend test gate's first run — a green status means least exactly when a pipeline has just changed — inverted. Here a red status meant least, and for the same reason: the status is a summary of the last thing that happened, not of what the build was for.
Why this is a SEPARATE row from VLN-038 and not an update to it
The two are adjacent and they are not the same finding. VLN-038 is "controls that never execute" — a silent absence. This is "a control that executes and blocks legitimate work" — a loud presence. They have different failure modes, different acceptance tests, and different fixes.
The decisive argument is that they move in opposite directions. Closing
VLN-038 means adding gates. Every gate added is another pre-build step,
another network round-trip, and another way for a deploy to be blocked —
so progress on VLN-038 makes VLN-039 worse. That is exactly what happened:
the unit-tests step that closed VLN-038's cheap half installs 21
packages where the step beside it installs one, and it was added hours
before this failure.
A row whose remediation worsens another row must not be that row. Folded together they would need one acceptance test spanning "every suite runs" and "no deploy is blocked by an outage", and no single check can say both.
Why Medium, and the argument against calling it High
The consequence is real and specific: a backend fix during an incident can be blocked by a read timeout in a test step. The gate has no knowledge of urgency, and the failure happens before any of the work being deployed is even evaluated.
But the honest read is Medium, not High, and the reason is that the mitigation is immediate and always available: a re-run at the same commit cleared it in under four minutes with no code change. Nothing is permanently blocked, no data is exposed, and an operator under pressure has a working action. High would put it beside VLN-010, VLN-011 and VLN-037 — findings where the estate was wrong about its own state for months without knowing. This is a known, visible, retryable delay.
It is not Low because the delay lands precisely when delay is most expensive, and because the number of round-trips is growing rather than shrinking.
Mitigated on forms-backend; what it does NOT solve
All three forms-backend pre-build installs now carry
--retries 10 --timeout 60.
--retries 5 is already pip's default, so raising the retry count
alone is close to a no-op — the lever that actually expired here is
--timeout, whose default is 15 seconds. Both are now explicit rather
than inherited from defaults that can change.
THIS DOES NOT DECOUPLE DEPLOYS FROM PyPI, and must not be read as
though it does. forms-backend/Dockerfile:10 runs
pip install --no-cache-dir -r requirements.txt to build the image that
actually ships. The coupling predates every test gate; what the gates
added was two earlier round-trips and a much larger package set. Lowering
the probability of a transient failure is all this change does.
The rejected option, and why — a pre-baked test image
The alternative was an Artifact Registry image carrying the test dependencies, used as the step image. It removes the round-trip outright, which is strictly better than tuning timeouts. It was rejected here for two reasons.
First, it does not deliver its headline benefit. It removes one of three round-trips and leaves the Dockerfile's, so the deploy remains coupled to PyPI either way — the incident scenario above is unchanged.
Second, and decisively: it creates a second copy of the dependency set
that must be kept in step with requirements.txt. This estate's
characteristic failure is two copies drifting — servers.py against
inventory.yaml, gpus-reports/ against soc-backend/ (99 lines), the
travel form's instructions against STEP_DESCRIPTIONS (GOV-021), and the
read-path argument that decided Option A in migration 022. A drifted test
image is worse than a flaky build, because tests would pass against
versions the shipped image does not contain, and nothing would report it.
It becomes the right answer if the flakiness persists after this change, or if the estate adopts a general staleness check for built images. Either would have to come first.
Adjacent finding — the security scan can no-op invisibly
forms-backend and gpus-forms-clamav-worker run pip install bandit
pip-audit without &&, and the block's last command ends || true.
So if that install fails, the step still reports success and scans
nothing. Unlike the gates above it fails open rather than closed.
Not fixed here — changing it to fail closed would make a scanning outage block deploys, which is the very problem this row is about, and that trade-off deserves its own decision rather than being made in passing.
VLN-040 — The dependency scan reports 24 advisories into || true¶
| Field | Value |
|---|---|
| id | VLN-040 |
| title | pip-audit runs on every forms-backend build, correctly reports 24 known vulnerabilities in 8 production dependencies, and its exit code is discarded by \|\| true. The build is green. The finding has been produced on every retained build and acted on by nothing. |
| register | Security findings |
| status | in progress — triaged 2026-08-24; step 1 of 6 done 2026-08-25 (flask-cors cleared, deployed and verified). 19 advisories remain across 7 packages and \|\| true is still in place. |
| severity | High |
| owner | R. Chhetry |
| evidence | Build a6637d70 (2026-08-24) step 2 security-scan: Found 24 known vulnerabilities in 8 packages, build status SUCCESS. The step is quoted below. Duration: the \|\| true was present in the FIRST forms-backend commit, 802b42c (2026-04-20) — 127 days. Cloud Build log bodies expire; the earliest retained log is 2026-07-27, reporting Found 22 known vulnerabilities in 8 packages. 41 builds since then, every one green, every one carrying the finding. |
| acceptance test | pip-audit --requirement forms-backend/requirements.txt --strict --ignore-vuln <each id in .pip-audit-exceptions with an unexpired date> exits 0. Deliberately NOT "the scan passes" — it already does that. The exception file must carry a dated entry per ignored id, per the .coverage-exceptions.yaml convention. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-24 |
| date_verified | — |
# forms-backend/cloudbuild.yaml, unchanged since 802b42c:
bandit -q -r forms-backend/ -x forms-backend/tests -ll || true
pip-audit --requirement forms-backend/requirements.txt --strict || true
# gpus-forms-clamav-worker/cloudbuild.yaml:
bandit -q -r gpus-forms-clamav-worker/ -ll || true
pip-audit --requirement gpus-forms-clamav-worker/requirements.txt --strict || true
VLN-010's shape, not VLN-039's — and the three must not be read as one
Three findings now concern this pipeline and they are genuinely different.
- VLN-038 — tests that never ran. A control absent. Fixed by adding it.
- VLN-039 — a gate that blocks legitimate work. A control too coupled. Fixed by loosening it.
- VLN-040 (this row) — a control that ran, produced correct output, and
was wired to nothing. This is VLN-010's shape: the Wazuh rules loaded,
validated, counted, and never matched. Here the scanner installed, ran,
correctly identified 24 real advisories, printed them — and
\|\| truethrew the exit code away, so the pipeline treated a true report as no report.
The distinction that matters operationally: VLN-038 was invisible, and this one has been on screen in every build log for at least 41 builds. Nobody was hiding it. It was reported, in full, and discarded by a two-character idiom.
Triage — 24 advisories, and most are NOT reachable. The count is not the finding.
Reachability assessed against actual estate usage, not against the package's full API. Where it could not be established it is marked unknown rather than guessed in either direction.
| Package | Advisories | Reachable? | Basis |
|---|---|---|---|
cryptography 42.0.7 |
8 | 5 NO, 3 UNKNOWN | The estate's ONLY import is AESGCM (forms-backend/crypto.py:16, worker mirror). No X.509, no key loading, no certificate verification anywhere. |
python-jose 3.3.0 |
5 | NO (all 5) | See the two notes below — one blocked empirically, four need JWE, which is never used. |
~~flask-cors 4.0.1~~ → 6.0.5 |
~~5~~ 0 | CLEARED 2026-08-25 | Was: CORS(app, origins="*") at app.py:70, in the request path. Bumped in 58567f0. See the step-1 note below — including a correction to this row's own "YES": the library is in the request path, but 3 of the 4 distinct defects were not. |
requests 2.32.2 |
2 | UNLIKELY | Used for JWKS + SOC API with fixed URLs. CVE-2024-47081 is a .netrc leak via malicious URLs; no .netrc, no caller-supplied URLs. |
flask 3.0.3 |
1 | NO | CVE-2026-27205 is a missing Vary: Cookie on session access. The portal uses no Flask session — it is a JWT API with no cookies. |
pg8000 1.31.2 |
1 | NO | CVE-2025-61385 is SQLi via pg8000.native.literal with a crafted list. The estate uses SQLAlchemy bound parameters; pg8000.native is never imported. |
protobuf 4.25.9 |
1 | UNKNOWN | Transitive via google-cloud-*. Not called directly. |
ecdsa 0.19.2 |
1 | NO | Minerva timing attack on P-256. Transitive via python-jose[cryptography]; the portal verifies RS256 against an RSA JWK, so no EC path is exercised. |
The three cryptography UNKNOWNs are the bundled-OpenSSL advisories
(GHSA-h4gh-qq45-vh27, GHSA-537c-gmf6-5ccf, PYSEC-2026-1284). Those are
flaws in the OpenSSL statically linked into the wheel, not in cryptography's
Python API. AESGCM does reach OpenSSL's EVP layer, so they cannot be
dismissed the way the X.509 ones can — but establishing whether the specific
OpenSSL defects touch AES-GCM requires reading the OpenSSL advisories of
2024-09-03 and 2026-06-09, which has not been done. Marked unknown, and it
is the gap this row owes.
A finding the scan did NOT report — algorithms is taken from the attacker's own token
Assessing CVE-2024-33663 (python-jose algorithm confusion) required
reading the calling code, and the calling code has a defect the scanner
cannot see. auth_v2.py:163-166:
The allowlist is derived from the unverified header of the token being
validated. An attacker controls alg, so they choose the algorithm their
own token is checked against. That is the precise shape algorithm-confusion
attacks exploit, and it is wrong independently of any CVE — an allowlist the
caller supplies is not an allowlist.
It is not currently exploitable, and that was verified rather than
assumed. _key_for_kid returns the Okta JWK dict (kty: "RSA"), and
python-jose refuses to build an HMAC key from it. Tested against the pinned
3.3.0:
jwk.construct(RSA-JWK, 'HS256') -> JWKError: Incorrect key type. Expected: 'oct', Received: RSA
jwk.construct(RSA-JWK, 'RS256') -> RSAKey
The defence is incidental and one refactor away. It lives in
python-jose's kty check, not in the portal's intent. Anyone
"simplifying" _key_for_kid to return a PEM string — a natural-looking
tidy-up — removes it silently, and the tests would not notice.
Fix regardless of patching: pin algorithms=["RS256"].
Cheap, independent of the version bump, and it makes the defence
deliberate.
Raised as its own row — VLN-041
— and pinned there on 2026-08-25 with its own tests. This row keeps the
provenance (the defect was found only because assessing CVE-2024-33663
forced a read of the calling code) and stops there: VLN-041 is a defect in
this estate's own source with a different fix, a different acceptance test
and a different lifecycle, and must not close or reopen with the scan-gating
work tracked here. One correction to the empirical note above, made
while writing VLN-041's tests: the kty check is not the only incidental
defence — see that row.
Why python-jose's other four are not reachable
CVE-2024-33664 and CVE-2024-29370 (each listed twice, making four of the
five) are both denial-of-service via jwe.decrypt on a high-compression
JWE — a "JWT bomb". The estate never decrypts JWE. Okta issues signed
(JWS) tokens; a grep for jwe across forms-backend/ and
gpus-forms-routing-worker/ returns nothing. Not reachable.
Two advisories have NO PUBLISHED FIX — this decides the remediation order
PYSEC-2026-1325—ecdsaMinerva timing attack. No fixed version.PYSEC-2025-185—python-josejwe.decryptDoS. No fixed version.
Both are assessed not reachable. But their existence means
pip-audit --strict can never exit 0 by patching alone, so removing
\|\| true without an exception mechanism would break the deploy pipeline
permanently rather than temporarily. That is why the acceptance test above
is written around a dated exception file and not around a clean scan.
Fixed versions, and where a bump will fight the pins
| Package | Now | Fix | Risk |
|---|---|---|---|
flask-cors |
4.0.1 | 6.0.0 | Major ×2. The only REACHABLE cluster — highest value, highest churn. |
cryptography |
42.0.7 | up to 49.0.0 | Seven majors. python-jose[cryptography] depends on it; bump them together. |
python-jose |
3.3.0 | 3.4.0 | Minor. Cheap. |
flask |
3.0.3 | 3.1.3 | Minor. |
requests |
2.32.2 | 2.33.0 | Minor. |
pg8000 |
1.31.2 | 1.31.5 | Patch. cloud-sql-python-connector[pg8000] pins a range — check it. |
protobuf |
4.25.9 | 5.29.6 / 6.33.5 | Major, and google-cloud-* pin protobuf ranges. Most likely to conflict. |
ecdsa |
0.19.2 | — | No fix. Exception. |
Recommended order: patch first, then remove || true — option (a)
Argued, and the argument is that (b) inverts the pressure.
Removing \|\| true now turns 24 open advisories into a broken deploy
pipeline immediately, and every one of them would have to be resolved
or excepted before anything ships — including the two with no published fix,
which can only ever be excepted. The estate would be patching
cryptography across seven major versions and protobuf against
google-cloud-* pins under a deploy freeze, which is the worst
condition for a change that touches the library performing AES-GCM on every
submission field. It also collides directly with VLN-039: a blocked
pipeline is exactly the incident-response problem that row is about.
Under (a) each bump is verified on its own, the 264-test suite (VLN-038's
gate) runs against it, and \|\| true comes off last — at which point the
green build is a true statement rather than a suppressed one.
Suggested sequence, cheapest and most reachable first:
1. ~~flask-cors → 6.0.0~~ — DONE 2026-08-25, at 6.0.5 (the fixed line's current patch, not 6.0.0). See the step-1 note below.
2. python-jose → 3.4.0. The algorithms=["RS256"] pin is no longer
part of this step — it shipped separately on 2026-08-25 under VLN-041,
deliberately decoupled so an auth-path change was not carried by a
dependency bump. test_auth_algorithm_pin.py is the regression gate for
the bump: run it against 3.4.0 before merging.
3. flask, requests, pg8000 — minors and a patch.
4. cryptography → current, with python-jose. Largest jump; AESGCM is the
only call site, so the blast radius is small and well-tested.
5. protobuf — last, because it is the one expected to fight the pins.
6. Except ecdsa and python-jose PYSEC-2025-185 with dates, then remove
\|\| true.
What (a) does not solve: nothing is patched until someone does step 1. Until then this row's status is the honest one — the finding is known, triaged, and open.
STEP 1 DONE 2026-08-25 — flask-cors 4.0.1 → 6.0.5, deployed and verified
Commit 58567f0, build f9bb3d10, serving revision
gpus-forms-backend-00108-sfw at digest sha256:1e21f227..., matching the
digest that build pushed. Built is not deployed; this was checked by digest,
not by tag.
The count, re-measured rather than subtracted. pip-audit --strict was
re-run against the requirements file as deployed (git show
58567f0:forms-backend/requirements.txt), and independently by the build's
own security-scan step on the file it shipped. Both report the same thing:
5 advisories were 4 defects. PYSEC-2024-71 and PYSEC-2024-260 are the
same CVE-2024-6221 listed twice — the same double-count already noted for
python-jose. Counting advisories overstates by one here, which is another
reason the headline number is not the finding.
A correction to this row's own triage table. It recorded flask-cors as
reachable — YES — on the grounds that CORS(app, origins="*") runs on
every request. The library is in the request path; 3 of the 4 defects were
not. CVE-2024-6866 (case-insensitive path match), CVE-2024-6839 (regex
pattern priority) and CVE-2024-6844 (unquote_plus turning + into a
space) are all path-matching defects, and each needs a per-path
resources= config to have anything to mismatch between. The portal has one
global policy and no per-path config. The fourth, CVE-2024-6221, sets
Access-Control-Allow-Private-Network: true; the backend is a public Cloud
Run service rather than a private-network target, so it was present but
inert. Reachable "library" and reachable "advisory" are different claims
and this table conflated them.
That does not make the bump optional — it makes its ordering load-bearing.
See VLN-042:
narrowing origins is normally written with a resources= dict, which is
exactly the construct that makes all three path defects reachable. Bumping
first is what keeps that narrowing from closing one row by opening three.
CVE-2024-6221 closed, observed on the live service rather than inferred
from a version number:
The standard this bump set, and what the next one owes
None of the nine test suites exercises a browser preflight. Had the
evidence been "282 tests pass", it would have covered none of the behaviour
a CORS library bump can break. flask-cors sits in the request path on
every request; the suite says nothing about that path.
What was done instead, and what the next dependency bump should be held to:
the exact call the portal makes — CORS(app, origins="*") — was driven
through Flask's test client on both versions, and every
Access-Control-* and Vary header compared across five request shapes:
the SPA's real preflight (Authorization + Content-Type), a
Private-Network preflight, a simple cross-origin GET on the
unauthenticated /metrics, a request with no Origin header (the nginx
same-origin path), and a case-variant path.
Exactly one header differed, in one case — Access-Control-Allow-Private-
Network: true → false, which is precisely CVE-2024-6221's fix. Everything
else byte-identical. The deployed service was then probed directly and
returned the same preflight headers the test client predicted.
That measurement is why this shipped without a staged rollout, and it is
recorded here so the next bump is argued from a measurement of the changed
behaviour rather than from a green suite. cryptography (step 4) is the
obvious case: AESGCM is the estate's only call site, so the equivalent
evidence is an encrypt/decrypt round-trip across both versions on the same
ciphertext — not the suite, which stubs it.
What step 1 does NOT do, stated so the status is not misread: nothing
else is patched, 19 advisories remain in 7 packages, and \|\| true is
still discarding the scanner's exit code on every build. This row closes at
step 6, not step 1.
VLN-041 — The JWT algorithm allowlist was taken from the token being validated¶
| Field | Value |
|---|---|
| id | VLN-041 |
| title | forms-backend/auth_v2.py derived jwt.decode(algorithms=[...]) from the unverified header of the token being validated, so a caller chose which algorithm their own token was checked against. Classic JWT algorithm confusion, in the authentication path in front of every authz check and every decrypt endpoint in the forms portal. |
| register | Security findings |
| status | done — 2026-08-25. Both acceptance conditions met and observed; see the closure note below. |
| severity | High (argued) — not exploitable on the pinned dependency set (see below), but the defect sits on the sole authentication path for every Phase 2 route, and the only thing preventing exploitation is third-party type checking that a routine refactor removes. Argued rather than supplied; no CVSS calculated. |
| owner | R. Chhetry |
| evidence | forms-backend/auth_v2.py:163-166 at commit 86c9eee, read 2026-08-25: algorithms=[unverified_header.get("alg", "RS256")], where unverified_header = jwt.get_unverified_header(token) at :150. Found while triaging CVE-2024-33663 for VLN-040 — pip-audit does not and cannot report it. Estate sweep, same date: one other decode site, forms-backend/auth.py:59-72, which takes algorithms=[key.get("alg", "RS256")] from the JWKS entry — server-controlled, not attacker-controlled, a materially different shape and not this defect. Nothing else in the estate decodes a JWT: gpus-forms-routing-worker uses google.auth only for signBlob; gpus-forms-clamav-worker leaves Pub/Sub push OIDC to Cloud Run's IAM layer; gpus-reports, scripts/, soc-backend, status-backend, security-backend have no jose import and no jwt requirement; the frontends do not parse tokens. Okta signing algorithm read live, not assumed: GET https://greenpeaceeu.okta.com/oauth2/v1/keys returns a single key {kty: RSA, alg: RS256, use: sig, kid: RWZkwN0Cn0ZK31J9hs2-sdu3YiKyUINyH2KJtD7RTdo}, and the discovery document advertises id_token_signing_alg_values_supported: ["RS256"]. |
| acceptance test | Two conditions, both required. (1) Live: the deployed gpus-forms-backend revision serving /api/forms is built from a commit in which python3 -m unittest test_auth_algorithm_pin passes, and a real Okta ID token still authenticates against it end-to-end — a pin that authenticates nobody is a worse defect, so the closing observation is a 200 on a real user's token, not a 401 on a forged one. (2) Structural: test_auth_algorithm_pin.TestSourceNeverReadsAlgFromTheToken is executed by a pipeline that can fail the build. Condition 2 cannot be met until VLN-038 lands; until then the guard protects a developer running the suite, not a deploy, and this row stays open on that basis rather than being closed on condition 1 alone. |
| blocks | — |
| blocked_by | — (was VLN-038; cleared 2026-08-25 when that row's remediation put the suite behind a hard gate) |
| date_raised | 2026-08-25 |
| date_verified | 2026-08-25 |
# forms-backend/auth_v2.py, before:
unverified_header = jwt.get_unverified_header(token) # :150 — attacker-supplied
...
claims = jwt.decode(
token,
key,
algorithms=[unverified_header.get("alg", "RS256")], # :166 — and so is this
audience=OKTA_CLIENT_ID,
issuer=OKTA_ISSUER,
)
# after:
algorithms=["RS256"],
CLOSURE — both conditions observed 2026-08-25, on the deployed revision
Condition 1 — a real Okta token returns 200 against the deployed
revision. MET. Observed by R. Chhetry at 15:41:12 UTC: signed in at
forms.greenpeace.us in a private window (empty sessionStorage, so a
genuine Okta round-trip and a freshly issued token, not a rendered shell),
and /my-requests listed 4 requests. Correlated in the Cloud Run request
log:
TIMESTAMP REVISION_NAME STATUS REQUEST_URL
15:41:12 gpus-forms-backend-00107-n8k 200 .../api/my-submissions
/api/my-submissions is @require_auth, so that 200 is
auth_v2._validate_token returning claims from a token verified against
algorithms=["RS256"]. /api/me (also @require_auth) returned 200 twice
in the same session, and /api/forms — the legacy auth.py path — returned
200 as well, so neither authentication path regressed.
The revision is the right one, by digest and not by tag. Serving
revision gpus-forms-backend-00107-n8k, Ready=True, 100% of traffic,
resolved image
...gpus-forms-backend@sha256:471a0c98dba6098b9ab7a58984b39b0e3f727a07e53c6845fac0208080876ed6
— byte-identical to the digest build df2e414c pushed for tag 1285a6b.
Cloud Run pins the digest at deploy, so this cannot be a tag that moved
afterwards.
Nobody was locked out — checked across all callers, not just one. Every 401 on the service since the revision went live at 14:37:46Z:
Both are the deliberate pre-flight probes described below. Zero 401s from any real client. That is the observation this condition was written to require: a pin that authenticates nobody is a worse defect than the one it fixes, so the closing evidence is a success, not a rejection.
Pre-flight probes against the live revision, run before the login so a failure would not be mistaken for a lockout — and worth keeping, because they exercise the rejection path end-to-end through the real deployment:
GET /health -> 200
GET /api/my-submissions (garbage bearer) -> 401 {"error":"invalid_token",
"detail":"Malformed token header: ..."}
GET /api/my-submissions (alg:HS256, fake kid) -> 401 {"error":"invalid_token",
"detail":"Signing key not found in JWKS"}
The second is the attack's shape, refused at the kid lookup before the
algorithm check is even reached — and refused with the portal's own error
contract. Under the old code the incidental defence raised JWKError, which
is not a JWTError and escaped uncaught: a 500. The clean 401 is itself
evidence the pinned path is what is running.
Condition 2 — the structural guard is executed by a pipeline that can fail
the build. MET, and evidenced rather than argued. The unit-tests step in
forms-backend/cloudbuild.yaml is a hard gate with no \|\| true, runs
unittest discover, and demonstrably failed a build over this exact
suite: build 34fe8dcb (8b7800c) died at step 1 with
ModuleNotFoundError: No module named
'cryptography.hazmat.primitives.asymmetric', Ran 265 tests,
FAILED (errors=1), build step 1 ... exited with non-zero status: 1. No
build, push or deploy step ran. Build df2e414c (1285a6b) then
reported Ran 282 tests ... OK with all nine suites named in the step
output, test_auth_algorithm_pin among them and all 18 of its tests listed
individually.
blocked_by: VLN-038 is cleared. That row's remediation is what makes
this condition true, and its first real catch was a defect in this row's own
commit.
The finding is that the defence was incidental — not that the attack failed
This was not exploitable on the pinned dependency set, and that was
established empirically rather than argued. It is recorded as High anyway,
because nothing in this repository was doing the defending. A forged HS256
token was built by hand and put through _validate_token against each shape
_key_for_kid could return:
_key_for_kid returns |
old — algorithms=[header alg] |
new — algorithms=["RS256"] |
|---|---|---|
| JWK dict (today) | JWKError: Incorrect key type. Expected: 'oct', Received: RSA |
rejected by the pin |
| PEM string | JWKError: ...asymmetric key... should not be used as an HMAC secret |
rejected by the pin |
| plain secret | ACCEPTED — token minted a submitter |
rejected by the pin |
Every cell that says "rejected" in the old column is python-jose's, and
none of them is the portal's. The VLN-040 note called out one such check
(kty); writing the tests found a second (HMACKey's asymmetric-key
guard), which is why the PEM refactor that note warned about would in fact
still have been caught on 3.3.0. That does not weaken the finding — it
means the estate was relying on two library behaviours it never asked for,
across a dependency the same triage is scheduled to bump. The third row is
the one with no library defence in it at all.
Two further observations from the same test run. JWKError is not a
subclass of JWTError, which _validate_token catches — so under the old
code the incidental defence escaped uncaught and would have produced a
500, not a 401. The portal's own error contract was never reached. And
once pinned, python-jose raises JWTError: The specified alg value is not
allowed, which the existing handler turns into a clean 401.
Why this is its own row and not an update to VLN-040
It was found inside VLN-040's triage, and the provenance is recorded there. It does not belong there.
- VLN-040 is a pipeline defect — a scanner's exit code discarded. Its
fix is
cloudbuild.yamland a dated exception file; it closes whenpip-audit --strictexits 0. - VLN-041 is a source defect in this estate's own code. No scanner reports it, no version bump fixes it, and no exception file could ever cover it. It closes when a deployed revision carries the pin.
Different fix, different acceptance test, different owner surface, different lifecycle. Nested inside VLN-040 it would have no id to cite from a commit, no line in the Index, and it would reopen every time the dependency work regressed. The distinction is not bookkeeping: a dependency advisory is someone else's defect that the estate inherits and can defer with a dated exception, whereas this is the estate's own, and "not currently exploitable" is not a disposition that can be written against it.
What a legitimate Okta change would do to the pin, and what it would not
A pin that breaks on a legitimate change would be a different problem, not a better one, so this was checked rather than assumed.
Key rotation does not touch it. Okta rotating to a new kid with the
same RS256 is the routine event, and _key_for_kid already force-refreshes
the JWKS cache and retries once (auth_v2.py:129-134). The algorithm is
unchanged; the pin is unaffected.
The org authorization server cannot change the algorithm.
0oazce9jd5inTjtrn417 is an app on https://greenpeaceeu.okta.com — the
org AS, no /oauth2/<asId> path segment. It signs ID tokens RS256 only
and exposes no per-app signing-algorithm setting.
Only a migration to a CUSTOM authorization server could, since those
permit RS384/RS512/ES256 and others. That is a deliberate configuration
project, not something Okta does to the estate unannounced. Handling:
widen the list from the new server's own
id_token_signing_alg_values_supported, in the same change as the cutover.
test_there_is_exactly_one_place_to_change_on_an_okta_migration asserts
there is exactly one algorithms= in the module, so that stays a one-line
edit in a known place.
auth.py:68 — a known second instance. DECIDED 2026-08-25: it stays.
The Phase 1 module takes its allowlist from key.get("alg", "RS256") where
key is the JWKS entry fetched from Okta over TLS. The attacker does not
supply it, so this is not the defect this row is about and pinning it
would close nothing.
DECIDED 2026-08-25 (R. Chhetry):
auth.py:68is not changed. Not attacker-controlled, andauth.pyis scheduled for retirement under the legacy-authz work — incremental hardening of code with a removal date gets deleted along with the code.
It is recorded here so the retirement does not ship the idiom forward.
That is the whole point of the entry: /api/admin/reload, /api/forms and
/api/pulldowns still authenticate through auth.verify_token, and when
that work migrates them to auth_v2 or to whatever replaces it, the
replacement must carry algorithms=["RS256"] as a literal. Copying
key.get("alg", ...) across would reintroduce a caller-supplied allowlist
into new code, where the argument that it is server-controlled would have to
be re-established from scratch rather than inherited. Acceptance
condition on the retirement work, not on this row: when auth.py is
deleted, no module that absorbs its routes reads alg from anything —
neither a token header nor a JWKS entry. test_auth_algorithm_pin's AST
assertion is the pattern to copy; it is written against auth_v2 and would
need pointing at the replacement module.
Test coverage as shipped
forms-backend/test_auth_algorithm_pin.py — 18 tests, all passing.
Real crypto: unlike the sibling suites it does not stub jose, because
what python-jose actually does is the subject. No network:
_key_for_kid is monkeypatched, so _get_jwks never runs.
- Behavioural — a forged HS256 token rejected across all three key
shapes in the table above; a valid RS256 token still verifies against
both a JWK and a PEM and still mints a
submitter; a token signed by the wrong RSA key still rejected (the pin is not a skipped signature check); an unknownkidstill fails closed. - Discriminating —
TestOnlyThePinRejectsItasserts that the same token and the same key python-jose accepts under["HS256"]is rejected under["RS256"]. This is what makes the pin load-bearing rather than decorative, and it is the assertion the other two key shapes cannot make. - Structural — same shape as
TestNoDecryptInQueueModule: noalgread from the unverified header in comment-stripped source, everyunverified_headerline is the assignment or.get("kid"), exactly onealgorithms=. Two AST assertions go tighter than a grep can — noast.Constantin the module has the value"alg"(immune to spelling, and docstring prose cannot hide in an AST), and everyalgorithms=kwarg must be a list literal of string constants equal to["RS256"], soalgorithms=[anything_computed]fails whatever the computation is.
Mutation-proved, not assumed. The old line was reintroduced, confirmed
to compile (py_compile clean) and confirmed live in the file, and the
suite re-run: 9 of 18 failed — 5 structural, 4 behavioural, the latter
including the two JWKError escapes described above. Restored; 18/18 pass.
The eight pre-existing forms-backend suites are unaffected.
The suite passed standalone and did not run at all under the build's own command
Recorded because it is the same failure spine as VLN-010, VLN-017 and VLN-040 — a control that looks present and does nothing — and this row nearly shipped an instance of it.
forms-backend/cloudbuild.yaml runs python3 -m unittest discover -p
'test_*.py'. The pin suite was written and verified with python3 -m
unittest test_auth_algorithm_pin. Under discover it failed to import:
test_approval_access and test_approval_persist register in-memory
placeholder modules for jose and cryptography at import time so they can
run where those packages are absent, test_approval_* sorts before
test_auth_*, and this is the one suite that must not be stubbed. The
unit-tests step is a hard gate with no \|\| true, so the build for
8b7800c fails there and the pin does not deploy — the fix is 26021a3.
Two things worth keeping from it. The gate worked: VLN-038's remediation
caught a broken test on its first real exercise, which is exactly what it
was added for. And "the tests pass" is not a claim about the pipeline
unless the tests were run the way the pipeline runs them; standalone and
discover are different commands and only one of them gates a deploy.
STANDING HAZARD for every test added to forms-backend — not a one-off
The paragraph above describes how one suite broke. This is the general form, recorded here because there is nowhere else it currently lives and the next person to add a test file will hit it without warning.
1. A suite verified by running it directly is not verified for the
pipeline. python3 -m unittest test_x and python3 -m unittest discover -p
'test_*.py' differ in more than convenience: under discover, every test
module in the directory is imported into one interpreter, in
alphabetical order, before any test runs. A module's import-time side
effects therefore land on every module sorting after it. The only run that
tells you anything about the deploy is the one the deploy performs.
2. Two suites mutate sys.modules at import time, and the mutation is
undetectable by inspection. test_approval_access.py:26 and
test_approval_persist.py:16 define _stub(), which registers in-memory
placeholders for google.cloud.*, cryptography.* and jose.* so those
suites can run where the packages are absent. Each placeholder carries
so it fabricates any attribute asked of it — __file__ and __spec__
included. A placeholder cannot be told from the real library by looking at
it; hasattr, __file__ and __spec__ all lie. The only reliable
discriminator is the module's name.
3. The stub is guarded on if mod not in sys.modules, not on whether the
package is installed. It therefore fires in Cloud Build too, where the
whole requirements.txt is installed — replacing real, present
libraries with fakes for every module imported afterwards. That is not a
local-dev-only device, and reading it as one is the trap.
What to do when adding a test to this directory. If the new suite needs
a real jose, cryptography or google.cloud — that is, if the library's
actual behaviour is the subject rather than an obstacle — evict the
placeholders by name at the top of the module, as
test_auth_algorithm_pin.py does, and do not restore them afterwards
(restoring fails: both libraries import submodules lazily, so a restored
placeholder is what those late imports find). If it does not need them,
nothing is required. Either way, run python3 -m unittest discover -p
'test_*.py' from forms-backend/ before pushing. That is the command
that gates the deploy, and it is the only one whose result means anything.
VLN-042 — origins="*" on an unauthenticated backend reachable from any origin¶
| Field | Value |
|---|---|
| id | VLN-042 |
| title | forms-backend/app.py:70 configures CORS(app, origins="*"). flask-cors reflects the caller's origin rather than sending a literal *, the service is deployed --allow-unauthenticated with ingress: all, and /health, /health/deep and /metrics carry no auth decorator. Any website a staff member visits can therefore read the portal's Prometheus metrics and its DB/KMS health from that person's browser. |
| register | Security findings |
| status | done — 2026-08-25. Narrowed in 158f932, deployed as revision gpus-forms-backend-00110-hvk. Both conditions observed on that revision. |
| severity | MEDIUM (argued) — bounded exposure today (see the scope note below, which is as much of this row as the exposure claim), raised above LOW for the supports_credentials edge and because the wildcard has outlived the condition its own TODO made it conditional on. Argued, not supplied; no CVSS calculated. |
| owner | R. Chhetry |
| evidence | forms-backend/app.py:61-70, read 2026-08-25 — the CORS(app, origins="*") call and the TODO(phase-2-followup) above it. Deployment posture: forms-backend/cloudbuild.yaml:107 --allow-unauthenticated; ingress: all per VLN-018. Unauthenticated endpoints: forms-backend/routes/health.py:31-47 — /health, /health/deep and /metrics carry no @require_auth or @require_role, unlike /api/admin/reload at :51. Header behaviour measured, not assumed: the exact call was driven through Flask's test client on flask-cors 4.0.1 and 6.0.5, and a simple cross-origin GET /metrics carrying Origin: https://evil.example returns Access-Control-Allow-Origin: https://evil.example with Vary: Origin on both versions — the wildcard is origin-reflection, not a literal *. Metrics exposed: forms_decrypt_total, forms_happyfox_total, forms_up and submission counters (routes/health.py:20-28). |
| acceptance test | CORRECTED 2026-08-25 — the original could not fail. See the note below. Two conditions, BOTH required. (a) The deployed SPA at forms.greenpeace.us still works end-to-end, observed in a browser — sign in, list forms, open a submission, submit. Necessary, and not sufficient: it guards against a config error that throws at init, and nothing more. (b) A cross-origin request from a NON-LISTED origin is refused. curl -H 'Origin: https://evil.example' <backend>/metrics must come back without an Access-Control-Allow-Origin header reflecting that origin. curl, not a browser — browsers will not generate this request, which is precisely why the condition needs a tool that will. (b) is the only condition that distinguishes a working narrowing from a no-op. |
| blocks | — |
| blocked_by | — (was VLN-040 step 1; cleared 2026-08-25 when flask-cors 6.0.5 deployed. The ordering note below is kept: the reason it blocked was technical, not procedural.) |
| date_raised | 2026-08-25 |
| date_verified | 2026-08-25 |
CLOSED 2026-08-25 — condition (a) observed on the narrowed revision
Condition (a) — the deployed SPA still works, observed in a browser.
MET. All four named flows exercised against revision 00110-hvk, from
the Cloud Run request log:
17:35:33 200 GET /api/forms ← list forms
17:35:38 200 GET /api/forms/53f29068c37e7067 ← open a form
17:35:46 201 POST /api/submissions ← the write
17:35:47 200 POST /api/submissions/8d9578b0-…/submit ← submit
18:06:53 200 GET /api/my-submissions
18:06:58 200 GET /api/submissions/c3e322ab-…/status ← open a submission
13 browser requests on this revision, 12× 200 and 1× 201, zero errors;
the /status read rendered the full page — chain, three CANCELLED steps.
Every non-2xx on the revision is a curl probe of ours. Condition (b) is
evidenced by the verbatim live response headers recorded below.
Why the fourth flow was nearly missed, and it is worth one line: it was
blocked by a different finding. The first pass exercised three of the
four. "Open a submission" did not run because the only submission to hand
was a Facilities Support Request — and a non-chain form is filtered out of
/my-requests (approval_queue.py:46), so there was nothing in the list
to click into. VLN-044
prevented VLN-042's own acceptance test from completing, and it did so
silently: the flow simply had no route in, and a green sweep over the
remaining three would have read as a complete pass.
Two findings interacting is how a test silently goes unrun. Neither row
predicts it. Nothing in VLN-042 mentions /my-requests scoping, and
nothing in VLN-044 mentions CORS. The only thing that caught it was reading
the log for the specific request the acceptance test named, rather than
accepting "all four flows passed" as reported.
NARROWED 2026-08-25 — condition (b) verified on the live service
158f932, build 11bb3cba, serving revision gpus-forms-backend-00110-hvk
at digest sha256:5f23ab44…, matching the digest that build pushed.
CORS(app, origins="*") is replaced by a flat three-entry allowlist:
the forms.greenpeace.us domain mapping and the frontend's two Cloud Run
URLs (hash form and project-number form — Cloud Run serves both and the
domain mapping replaces neither).
Condition (b) — an unlisted origin is refused. MET 2026-08-25.
Evidence is the live service's own response headers, captured verbatim
against revision 00110-hvk. Not a test client: test_client performs no
preflight, and the whole point of this condition is what the deployment
actually returns.
$ curl -D - -o /dev/null -H 'Origin: https://evil.example' $BACKEND/metrics
HTTP/2 200
$ curl -D - -o /dev/null -H 'Origin: https://forms.greenpeace.us' $BACKEND/metrics
HTTP/2 200
access-control-allow-origin: https://forms.greenpeace.us
vary: Origin
$ curl -D - -o /dev/null $BACKEND/metrics # NO Origin — the nginx-proxied path
HTTP/2 200
access-control-allow-origin: https://forms.greenpeace.us
vary: Origin
$ curl -X OPTIONS -H 'Origin: https://evil.example' \
-H 'Access-Control-Request-Method: POST' \
-H 'Access-Control-Request-Headers: authorization, content-type' \
-D - -o /dev/null $BACKEND/api/submissions
HTTP/2 200
$ curl -X OPTIONS -H 'Origin: https://forms.greenpeace.us' \
-H 'Access-Control-Request-Method: POST' \
-H 'Access-Control-Request-Headers: authorization, content-type' \
-D - -o /dev/null $BACKEND/api/submissions
HTTP/2 200
access-control-allow-origin: https://forms.greenpeace.us
access-control-allow-headers: authorization, content-type
access-control-allow-methods: DELETE, GET, HEAD, OPTIONS, PATCH, POST, PUT
vary: Origin
The unlisted cases return no access-control-allow-origin at all. On
revision 00108-sfw, the same GET /metrics returned
access-control-allow-origin: https://evil.example.
THE POSITIVE CASES ARE PART OF THE EVIDENCE, NOT DECORATION. A negative
result on its own cannot tell "the narrowing worked" apart from "CORS broke
entirely" — both produce a missing header. The listed origin still being
reflected, with its full allow-headers / allow-methods / Vary set, is
what makes the negative result mean what it is claimed to mean.
The no-Origin path — the one every real request takes — was re-checked
deliberately. It changed from access-control-allow-origin: * to
: https://forms.greenpeace.us (flask-cors emits the first listed origin
when no Origin is present). Non-browser callers ignore the header
entirely, and the response is still 200. Confirmed end-to-end through the
real proxy rather than only against the backend's own hostname:
$ curl -H 'Authorization: Bearer not.a.jwt' https://forms.greenpeace.us/api/my-submissions
HTTP 401 {"detail":"Malformed token header: ...","error":"invalid_token"}
$ curl https://forms.greenpeace.us/ HTTP 200 # the SPA shell
A 401 from the backend's own error contract — rather than a 502/504 —
is the success signal here: it proves nginx reached the backend and the
backend answered. /api/forms (legacy auth.py) and /api/me return the
same. /health and /health/deep are 200.
Condition (a) — the SPA still works, observed in a browser — is NOT yet met. No browser has touched this revision. Per this row's own corrected acceptance test, (a) is necessary and insufficient, and (b) alone does not close the row either. The row stays open until both are observed.
Flat list, not resources= — and the code says so at length
forms-backend/app.py carries the reasoning inline, because the dangerous
change here is a plausible-looking improvement rather than a mistake:
resources={r"/api/*": {...}}would reopen three CVEs.CVE-2024-6866,-6839and-6844are path-matching defects that need per-path config to have anything to mismatch between. A flat list has one policy for every path and cannot exercise that class at all. The tempting next step — exempting/healthand/metricsfrom the policy — is exactly the change that would do it, and they need no exemption: they are unauthenticated and answer any caller regardless.supports_credentials=Truemust not be added. flask-cors will not pair a literal*with credentials, but it will pair a reflected origin with them — this configuration plus credentials, on an API that returns decrypted submission fields from/approval-view.- This is not a load-bearing auth control.
@require_auth/@require_roleare the gate. CORS stops another origin reading a response and stops nothing else; non-browser callers are unaffected by any value here. - No
localhostentry and no env var. Dev proxies/apithrough Vite, so dev is same-origin too. An allowlist that can be widened by untracked config is not an allowlist — a developer pointing a dev server straight at a deployed backend adds their origin in a commit.
Scope — what this does NOT expose. An overstated row is as bad as a missed one.
No PII, no submission content, no token. The API is Bearer-only: the
SPA attaches Authorization: Bearer <id_token> from JavaScript
(forms-frontend/src/lib/api.ts:47). The token lives in sessionStorage
(src/lib/auth.ts:45), which is scoped to the SPA's own origin, and no
browser attaches it to a cross-origin request — a script on another site
cannot read it and cannot cause it to be sent. supports_credentials is not
set, so cookies are not in play either (and the portal uses no Flask
session).
The exposure is exactly and only the unauthenticated endpoints, which
would answer any caller anyway — curl reaches them today. What
origin-reflection adds is that the read can be performed by a staff
member's browser, from a page they merely visited, which curl from the
open internet cannot do: it attributes the request to an internal-looking
client and needs no attacker infrastructure. That is a real difference and
it is the whole of the difference.
This row is a missing defence-in-depth layer, not an open door. Recorded at that weight deliberately.
Why the original acceptance test was wrong — correct about the library, wrong about the system
The test this row was raised with read: "the deployed SPA still works, observed in a browser." It could not fail. Any narrowing at all — including one with an empty origin list, or a typo'd domain — would have passed it.
The reasoning was drawn from the code and never checked against the
deployment topology. From forms-frontend/src/lib/api.ts:47, every SPA
request carries an Authorization header; a request with that header is
not a CORS-simple request; therefore every request is preflighted;
therefore a broken origin list breaks everything and a working SPA proves
the list is right. Every step of that is true of the library, and the
conclusion is false of this deployment, because preflight applies only
to cross-origin requests. forms-frontend/nginx.conf reverse-proxies
/api/ under forms.greenpeace.us, so the browser sees same-origin
requests, sends no Origin header, and never asks permission. Narrowing
the list cannot break something that never consults it.
Proved, not argued. Cloud Run request logs for the browser session that
verified the flask-cors bump on 2026-08-25 (revision
gpus-forms-backend-00108-sfw) show two API calls from a real browser and
zero OPTIONS requests:
16:02:27 200 GET /api/my-submissions Mozilla/5.0 (Macintosh…)
16:02:31 200 GET /api/submissions/…/status Mozilla/5.0 (Macintosh…)
The only OPTIONS in the window were curl probes. nginx.conf's own
comment says it plainly — the proxy exists "so that the frontend and
backend appear at the same origin … and CORS becomes a non-issue."
Name the mistake, because it is a different one from the register's others. VLN-010, VLN-017 and VLN-040 are all controls that ran and were wired to nothing. This is not that. This is a correct inference about a component, applied to a system whose topology makes it irrelevant — code read accurately, deployment not read at all. It produced an acceptance test that would have been signed off, on evidence that was real but measured the wrong thing.
The generalisation worth carrying: an acceptance test must be checked for whether it can fail before it is checked for whether it passes. A condition no realistic defect could violate is not a weak test, it is not a test. Ask what would have to be true for it to go red; if the answer is "nothing that could plausibly happen", the condition is decoration.
The sharp edge — one flag from wildcard to origin-reflection WITH credentials
flask-cors will not send a literal * together with credentials. Given
supports_credentials=True it echoes the request origin instead, which
is the genuinely dangerous configuration: any origin, credentialed, on an
API that returns decrypted submission fields from
/api/submissions/<id>/approval-view.
Nothing in the repository sets that flag today. The finding is that one
keyword argument separates the current posture from that one, on a line
whose own comment (app.py:61-69) says
"Currently wildcard so local dev + the about-to-be-deployed Cloud Run frontend both work; tighten to the specific frontend origin in a follow-up commit once that URL is known"
TODO(phase-2-followup): replace "*" with explicit origins=[...]
Both URLs have been known for months. forms.greenpeace.us is live and
is recorded in inventory.yaml:1044. The condition the TODO made itself
conditional on was met long ago and the TODO did not fire — which is why
this is a register row and not a code comment.
Why this is NOT part of VLN-040
VLN-040 is dependency advisories — someone else's defects, inherited, dispositioned by version bumps and dated exceptions. This is the estate's own configuration. No scanner reports it, no bump fixes it, and no exception file could cover it. Same distinction that separated VLN-041 from VLN-040, applied to a config defect rather than a source one.
Ordering: narrow AFTER the flask-cors bump deploys — and the reason is technical
This looks like ordinary change hygiene. It is not.
Three of the four distinct advisories against flask-cors 4.0.1 —
CVE-2024-6866 (case-insensitive path matching), CVE-2024-6839 (regex
pattern priority), CVE-2024-6844 (unquote_plus turning + into a space)
— are path-matching defects. Every one of them needs a per-path
resources= configuration to have something to mismatch between. The
portal has a single global policy and no per-path config, so none of the
three is reachable today.
Narrowing origins is normally written as CORS(app, resources={r"/api/*":
{"origins": [...]}}). That is precisely the construct that makes all
three reachable. Narrowing on 4.0.1 would close this row by opening three
others. The bump (e56c646, VLN-040 step 1) must deploy first.
If the narrowing is written without a resources= dict — a flat
CORS(app, origins=[...]) — the three do not become reachable. Do not rely
on that: the natural next step after narrowing origins is exempting
/health and /metrics from the policy, which requires per-path config.
VLN-043 — The forms SPA reads none of the three variables its env template defines¶
| Field | Value |
|---|---|
| id | VLN-043 |
| title | forms-frontend/.env.local.example defines VITE_OKTA_ISSUER, VITE_OKTA_CLIENT_ID and VITE_API_BASE_URL. The SPA reads none of them — two values are hardcoded in src/lib/auth.ts and the third is read under a different name (VITE_API_BASE). The template's VITE_OKTA_CLIENT_ID is also the pre-cutover shared-portal app id, wrong since 2026-08-05. |
| register | Security findings |
| status | open |
| severity | LOW (argued) — nothing is currently misrouted, because the variables are inert. The severity is about what the template would do if the inertness were ever "fixed", and about an operator trusting a value that has no effect. Argued, not supplied. |
| owner | unassigned |
| evidence | Read 2026-08-25. forms-frontend/.env.local.example and the working .env.local both set VITE_OKTA_ISSUER=https://greenpeaceeu.okta.com, VITE_OKTA_CLIENT_ID=0oavvg1y33wTWFsmP417, VITE_API_BASE_URL=https://forms.greenpeace.us. Against that: src/lib/auth.ts:31-32 hardcodes const ISSUER = 'https://greenpeaceeu.okta.com' and const CLIENT_ID = '0oazce9jd5inTjtrn417', and src/lib/api.ts:30 reads import.meta.env.VITE_API_BASE — not VITE_API_BASE_URL. A repo-wide grep for import.meta.env returns exactly one hit, api.ts:30. auth.ts:14-18 already documents its own half: "The frontend Cloud Run service carries no env vars, and auth.ts does not read import.meta.env … (.env.local.example defines VITE_OKTA_CLIENT_ID but nothing reads it.)" — the VITE_API_BASE_URL / VITE_API_BASE name mismatch is not documented anywhere. The clientId in the template is the shared "GPUS Internal Portals" app, which auth_v2.py:35-39 records Forms as having cut over FROM on 2026-08-05 and which status-site, soc-site and mkdocs-portal still use. |
| acceptance test | Either the template defines only variables the code actually reads (names matching, stale values corrected or removed), or it is deleted and the hardcoded values documented as the single source. A grep of forms-frontend/ finds no VITE_ name in an env file that is absent from src/, and none present in src/ that is absent from the template. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-25 |
| date_verified | — |
Why a dead variable is worth a row
Same class as the env-template drift that would have rerouted every forms notification on a DR restore: a configuration value that looks authoritative, is read by nothing, and is wrong. All three failure surfaces are here at once.
- It misleads. An operator changing
VITE_API_BASE_URLto point at a staging backend observes no change and starts looking for the fault somewhere else. Nothing warns them; the variable is spelled plausibly and sits under a comment block that describes it as the backend API setting. - It is a loaded gun for whoever "fixes" it. The obvious tidy-up —
wiring
auth.tstoimport.meta.envso the clientId is configurable, or renamingVITE_API_BASEto match the template — would immediately point Forms at0oavvg1y33wTWFsmP417. The backend checksaud == OKTA_CLIENT_ID, deployed as0oazce9jd5inTjtrn417(cloudbuild.yaml:126), so every token would fail the audience check and every user would get a 401. The template does not merely do nothing; it holds a value that would break authentication for everyone the moment it started doing something. - It rots silently. The clientId went stale at the 2026-08-05 cutover and nothing noticed, because nothing reads it. A value nothing reads has no feedback path and cannot self-correct.
Not fixed in this pass — deliberately. The correct fix is a decision (wire the variables up, or delete the template and document the hardcoded values), not an edit, and choosing wrongly is how the second bullet happens.
VLN-044 — The confirmation page promised an approval chain that 28 of 29 forms do not have¶
| Field | Value |
|---|---|
| id | VLN-044 |
| title | forms-frontend/src/routes/Submitted.tsx rendered a "Following this request" panel — "You can check which approval step it has reached at any time — and you will be emailed when it is approved, returned or declined" — on every submission, with links to /status/<id> and /my-requests. Only travel-request-001 has an approval chain. For every other form there is no step, no outcome email, and no row in the list the second button leads to. |
| register | Security findings |
| status | in progress — fixed in the working tree 2026-08-25 with a both-directions test; not deployed, not verified against live state |
| severity | MEDIUM (argued) — no data exposure and no authz failure; this is an integrity-of-record defect. It is the portal telling its user, at the moment they decide what to do next, something specific and false about what will happen to their request. Argued, not supplied; no CVSS. |
| owner | R. Chhetry |
| evidence | Found by submitting, 2026-08-25 17:35 UTC. A real Facilities Support Request, 8d9578b0-046a-4604-a36d-2006c6e37a69, submitted through the production SPA during VLN-042's condition (a); the confirmation page rendered the panel. Cloud Run request log: 17:35:47 200 POST /api/submissions/8d9578b0-…/submit. The gate: Submitted.tsx:179 read {result.id && ( — id is the submission's own UUID, returned by every successful finalize, so the condition was a tautology. The field it was supposed to read did not exist: the comment directly above claimed "routing.approval is present only when finalize ran one", but contract.ts:246-249 declared routing as {email, happyfox} and routes_phase2.py:723-735 returned exactly those two. The chain gate: routes_phase2.py:691 — if submission.form_id == "travel-request-001":. The chain set: approval_resolver.py:66-72, FORM_CHAINS has exactly one key, extracted from source rather than read by eye. The list scope: approval_queue.py:46-48 — if sub.form_id not in chain_form_ids: return False, so a non-chain submission is silently absent from /my-requests. 31 form YAML files in the repo (27 in the live catalogue); 1 has a chain. |
| acceptance test | cd forms-backend && python3 -m unittest test_submitted_panel_gate passes, and the deployed SPA shows the panel on a travel-request confirmation and no panel on a non-travel one, observed in a browser. BOTH DIRECTIONS ARE REQUIRED IN BOTH HALVES. Deleting the panel outright removes the false statement and would pass any test that only checks "absent for a facilities request" — so the travel direction is what makes this an acceptance test rather than a regression waiting to be signed off. Re-runnable: the suite is in the hard unit-tests gate (VLN-038's remediation), so it runs on every build. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-25 |
| date_verified | — |
The guard was designed, documented, and implemented against a field nobody built
This is not a missing check. Submitted.tsx carried, verbatim:
/*
Only for a submission that HAS an approval chain: linking a status page
for a form with no chain would promise a progress view that has no
progress to show. `routing.approval` is present only when finalize ran
one.
*/}
{result.id && (
The comment states the correct rule. The code implements a tautology.
And routing.approval existed on neither side of the wire. So the author
reasoned it through, wrote the reasoning down, and then wrote a condition
against a contract that was never built — on either the frontend or the
backend.
That is worse than an absent guard, and the reason is worth stating. An ungated block invites the question "should this be conditional?". A block whose comment says it is conditional answers that question wrongly, to every future reader, including a reviewer looking specifically for this class of defect. The comment was evidence the case had been handled. It was the opposite.
Rule taken from it: a comment asserting a guard is a claim about the code and must be checked like one. Where the two disagree, the comment is not documentation — it is a false statement in the place people look to avoid reading the code.
Third instance of the same pattern — and the first found by USING the portal
UI or copy that describes a workflow the system does not perform:
| Finding | How it was found | |
|---|---|---|
| 1 | GOV-021 — travel-request-001's own instructions column described an approval order the resolver does not perform ("…then to Finance"), served live to every submitter for the life of the chain |
Reading — comparing the YAML to STEP_DESCRIPTIONS |
| 2 | The WITHDRAWN fallback — SubmissionStatus.tsx has no branch for it, so the submitter reads "This request is recorded with this state. Contact IT if you need more detail." (recorded in schema/021_withdraw_test_chains.sql) |
Reading — while verifying an unrelated deploy |
| 3 | This row | Submitting a real request |
The first two were found by someone reading code against other code. This
one required a human to fill in a form and look at the page — and it was
found on the first non-travel submission anyone has made since the panel
shipped. That is the finding about the finding: the estate's verification
has been strong on reading and weak on using, and this class of defect is
invisible to reading alone precisely because each artifact is internally
coherent. Nothing in Submitted.tsx is wrong on its own terms; it is wrong
only against a fact that lives in approval_resolver.py.
The unexercised remainder is named, not implied: the revise path
(/revision-source, prefill, resubmit) has never been driven through a
browser end to end. It is the only piece of what shipped to staff still in
that state.
Population — 6 days, 4 submissions, 3 affected, ONE PERSON (revised twice)
Stated precisely rather than as a shape, because the shape and the blast radius are very different numbers here.
The panel shipped in 1daeea0, deployed by build 3ddb9338 at 2026-08-19
18:51 UTC. Every finalize since, from the Cloud Run request log:
08-25 18:15 — /api/submissions/b5e998f7-…/submit ← Change of Address (AFFECTED)
08-25 18:08:24 200 /api/submissions/f8f89f9e-…/submit ← Facilities Support (AFFECTED)
08-25 17:35:47 200 /api/submissions/8d9578b0-…/submit ← Facilities Support (AFFECTED)
08-20 16:41:24 200 /api/submissions/e5ff2b59-…/submit ← travel; panel truthful
e5ff2b59 persisted a chain — travel_approval.persisted submission=e5ff2b59-…
outcome=PERSISTED state=IN_APPROVAL current_step=3 — so the panel was
correct for it and for it alone.
Shape: 28 of 29 forms would have shown a false statement. Realized: three submissions — and, importantly, ONE PERSON.
The blast radius is one person finding this three times, not three people misled. All three affected submissions are R. Chhetry's, in a single session on 2026-08-25, made while deliberately exercising the portal. No staff member outside IT has encountered this panel on a non-chain form. That distinction is the difference between a defect with a user population and a defect with a discovery narrative, and the row would misrepresent itself by reporting "3 affected" without it.
The third instance confirms the diagnosis rather than extending it. The
gate is a tautology ({result.id && (), so the panel shows on every form
by construction — Change of Address is not a new class of case, it is the
same case a third time. No further instances are needed or should be
collected; another submission would add a row to this table and nothing
to the finding.
A SECOND COUNT WAS ALSO WRONG, in the other direction. This row first
said "30 of 31 forms". That counted forms/*.yaml minus pulldowns.yaml,
which swept in templates.yaml and _schema.yaml — neither is a form, and
yaml_loader.py:174 does not load either. The live figure is 29 forms,
of which 1 has a chain: 28 affected. Corrected 2026-08-26. It changes
nothing about the finding and is fixed because a register that rounds its
own denominator invites the next reader to round theirs.
Still carrying the old number: forms-frontend/src/lib/contract.ts, in
the RoutingApproval doc comment ("30 of 31 forms"). Comment only, no
behaviour, deliberately not corrected tonight — it would trigger a frontend
deploy for a comment. Fix it with the next change to that file.
THE COUNT WAS EVIDENCE AND IT CHANGED — TWICE. Recorded as revisions rather than edits. This note first read "1 affected", written at 17:50 when that was true; revised to 2 at 18:08; revised to 3 at 18:15. The population of a live defect is a measurement with a timestamp, not a property of the finding: it is correct only as of when it was taken and it grows until the fix deploys. Anything quoting an earlier figure is quoting an earlier clock, which is why each revision is dated in place rather than overwritten.
The fix — the backend reports what happened; the SPA is told, not required to know
routing.approval now exists, which is what the original comment assumed.
It carries the outcome for this submission, not the form's
configuration:
status |
meaning | panel? |
|---|---|---|
persisted |
a round exists with steps to follow and an approver to email | yes |
none |
this form has no chain — 28 of 29 forms | no |
blocked |
a chain was configured but resolved to no steps | no |
failed |
persistence raised; the submission is finalized and routed, the chain is not | no |
Outcome, not configuration, and that is the load-bearing choice. A form
lookup would answer "travel-request-001 has a chain" for a submission whose
chain failed to persist — and _run_travel_approval swallows that failure
by design, because a submitted request must not 500 over an approval
problem. The return value is the only place that fact can reach the page. So
blocked and failed are travel-request outcomes that leave nothing to
follow, and the UI treats them exactly like none.
Absent means false. A cached SPA bundle talking to an older backend gets
no panel rather than the old always-on one. Guessing wrong in that direction
costs a travel submitter one link that the /my-requests nav item already
provides; guessing wrong in the other direction is this row.
No hedge replaces it. For a form with no chain there is genuinely nothing to follow, and the honest render of that is silence. "You may be contacted if anything is needed" would be the same defect with better manners. The headline prose already states what did happen — logged, and queued for delivery to the relevant team.
No form id in the frontend. The backend owns FORM_CHAINS; a hardcoded
travel-request-001 in the SPA would be a second copy that silently stops
matching the day a second chain is configured. Asserted by a test, against
comment-stripped source — the code comment does name the form id when
explaining where the backend gate lives, and a naive grep would have fired
on the explanation.
Mutation-proved in BOTH directions, on mutants that compile
19 tests in test_submitted_panel_gate.py; the forms-backend suite is now
301 tests across 10 suites, green under unittest discover.
| Mutant | Compiles? | Caught by |
|---|---|---|
Gate back to {result.id && (, unused const removed |
yes, tsc exit 0 |
6 tests — …is_not_gated_on_something_always_true, …gate_reads_the_approval_outcome, …absent_status_is_treated_as_no_chain, and 3 more |
| Panel deleted outright ("fixed" by removal) | yes | test_the_panel_still_exists_for_a_chain_bearing_submission |
The two mutants are caught by disjoint tests, which is the property that
makes this bidirectional rather than merely thorough. The first mutant was
initially written without removing the now-unused hasApprovalChain, which
made tsc fail for an incidental reason — re-run with the const removed so
the mutant genuinely compiles and the test is demonstrably what catches it.
VLN-045 — Every submission's status page is headed "Travel request."¶
| Field | Value |
|---|---|
| id | VLN-045 |
| title | forms-frontend/src/routes/SubmissionStatus.tsx:90 renders a hardcoded <h1>Travel request.</h1>. The page serves any submission the caller owns, so a Facilities Support Request's own status page is titled with a different form's name. Two sibling views carry the same hardcoded string and are correct only by accident of scoping. |
| register | Security findings |
| status | in progress — fixed 2026-08-26 in ee2da8a; not yet verified in a browser |
| severity | LOW (argued) — no disclosure and no authz failure. It is the fourth instance of the estate's recurring pattern: an artifact that is internally coherent and false against a fact living elsewhere. Argued, not supplied. |
| owner | unassigned |
| evidence | Read 2026-08-25 while investigating a report of "the approver in the other forms". SubmissionStatus.tsx:90 — <h1 style={{ fontSize: 32, marginBottom: 6 }}>Travel request.</h1>, outside every conditional. The route is /status/<submission_id> and its endpoint authorizes per submission with no form filter (routes_phase2.py:1223, resolve_submitter_access), so it renders for any owned submission. Two siblings carry the same literal and are currently correct: MyRequests.tsx:148 (the list is scoped to chain-bearing forms, approval_queue.py:46) and ApprovalDetail.tsx:117 (the approver queue is chain-only). Both become wrong the day a second chain is configured — they are latent, not safe. |
| acceptance test | /status/<id> for a non-travel submission shows that form's own name, observed in a browser; and /status/<id> for a travel request still shows "Travel request." Both directions, for the same reason as VLN-044: replacing the heading with a generic "Request." would fix the false statement and pass a one-directional test while making the travel page worse. The form name is already available — GET /api/submissions/<id>/status returns form_id, and the SPA has the catalogue from /api/forms. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-25 |
| date_verified | — |
What this row is NOT — the investigation that produced it
Raised out of a report that non-travel forms were "showing the approver", which would have been a far larger finding: chain data rendered for a submission that has none, or read from another submission. Neither is happening. Three checks, all negative:
- No non-travel form claims an approval chain. All 31 form YAMLs were
grepped for approver / approval / SMT / sign-off / supervisor / manager
wording. Seven HR forms carry fields asking for a manager or timesheet
approver (
Manager,TimeApproval,CurrentManager— "Current time card approver"). Those are input questions the submitter answers, not statements about routing, and they are correct.travel-request.yamlis the only form whose text describes an approval chain. - Nothing approver-shaped renders for a chainless submission.
SubmissionStatus.tsx:98gates the whole chain block on{approval.steps.length > 0 && <Chain steps={approval.steps} />}, and the endpoint returnssteps: []when there is no round.total_stepsis returned as the travel constant3even with no chain, but it is read only inside theSTATE_IN_APPROVALbranch, whichstate === nullcannot reach.stateLabel(null)correctly returns "Submitted". No approver name, no step, no roster reaches a chainless page. - No path reads another submission's chain. Every chain read is keyed
on
submission_id:load_current_roundandload_steps(approval_access.py:376,414) both filter on it,load_submission_refis a primary-keyget, and_decrypt_submission_fields(routes_phase2.py:852) filtersSubmissionField.submission_id. The one query that is not per-submission isload_current_rounds()(approval_queue.py:317), which loads every round to build the queue — and it attributes steps to submissions by a(submission_id, version)dict key, never positionally. A chainless submission has no row insubmission_approvals, so it cannot appear in that result set at all, with or without steps.
What the reporter most likely saw is either the HR forms' legitimate manager fields, or the Approvals queue they opened at 17:36:55 — which shows approvers because it is the approver queue and is chain-only. The one thing genuinely wrong on a non-travel page is this heading.
FIXED 2026-08-26 — the name is data, in all three views
ee2da8a. QueueSubmission carries form_title (one batch lookup, not one
per row); both row builders emit it; /status and approval-view emit it.
A shared submissionHeading() renders "<name>." and falls back to
"Request." — deliberately generic, because a form name as the default is
the defect whichever form it names.
Two explicit key-allowlist tests fired on the new row key and were
extended deliberately rather than loosened. That is exactly what they exist
for: TestQueueRowShape.ROW_KEYS and TestMySubmissionsRowShape.ROW_KEYS
pin the row's ENTIRE key set so that no key can be added by accident, and
form_title is not field content — it is the public name of a form whose id
the caller already holds.
And a test fake was wrong in a way worth recording. _FakeSession.get
ignored its model argument and returned the submission for every lookup.
Harmless while the handler fetched one model; approval-view now also reads
the Form row, so the fake handed back a Submission, .name raised, and
eight unrelated tests failed at once as a 500. A stub that is loose
about what it is asked for does not fail where the looseness is — it fails
somewhere else, later, in bulk.
Fourth instance of the pattern, and the one that dates it
UI or copy describing a workflow the system does not perform: GOV-021's
served instructions text, the WITHDRAWN fallback, VLN-044's panel, and
this. All four share a shape — an artifact that is correct on its own
terms and false against a fact that lives in another file. Nothing about
<h1>Travel request.</h1> is wrong in SubmissionStatus.tsx; it is wrong
only because resolve_submitter_access does not filter by form.
This one dates the pattern rather than just repeating it. The heading was
written when /status served travel requests and nothing else, and it was
true then. It became false when the page began serving every form — a
change made elsewhere, which no test and no reviewer connected back to a
literal string three files away. The defect was introduced by a change
that did not touch it.
What a test pinning the copy would NOT have caught — and what would
Asked for explicitly, because this row is a different class from its three siblings and the difference decides what to build.
GOV-021, the WITHDRAWN copy and VLN-044 were all wrong when written. A test asserting the copy at authoring time — "the panel says X", "the instructions say Y" — would have been written against the same wrong understanding and would have passed, then defended the error. Snapshot tests on copy are worth little against them.
This heading was TRUE when written. /status served travel requests and
nothing else. It was made false by a change three files away that did not
touch it: the page began serving every form the caller owns, because
resolve_submitter_access is per-submission and has no form filter. A
snapshot test would have passed before the change and passed after it,
since the literal never moved.
What catches this class is asserting the relationship, not the value:
- "No view hardcodes a form name" — a structural assertion over source, which fails the moment a literal is reintroduced anywhere, and which cannot be satisfied by a literal that happens to be right today.
- "The heading is a function of the payload" — the view must read
form_titlefrom its own data. A view that renders a constant fails whatever the constant says. - Both directions, so replacing one literal with a generic literal fails too — which is mutant M2 below.
That is the general form: when a view's correctness depends on a fact owned by another module, pin the dependency, not the output. A value can be right for the wrong reason and stay right until the reason changes.
VLN-046 — "Ticket: NOT CREATED" asserts a negative the portal cannot verify¶
| Field | Value |
|---|---|
| id | VLN-046 |
| title | Submitted.tsx:47 renders an amber Not created pill against a Ticket row whenever a form declares any happyfox_template action and no ticket id came back. It states as fact that no ticket exists. The portal cannot know that: the same submission's email_template legs deliver to gpus-*@greenpeace.org addresses, at least one of which is proven to auto-open a HappyFox ticket by email ingest. |
| register | Security findings |
| status | in progress — copy fixed 2026-08-26 in ee2da8a; not yet verified in a browser |
| severity | LOW (argued) — no disclosure, no data loss. It is the inverse of VLN-044: that one is reassuring and false, this one is alarming and unverifiable. Argued, not supplied. |
| owner | unassigned |
| evidence | Observed 2026-08-25 on the Change of Address Notification confirmation (b5e998f7), and absent from both Facilities Support confirmations. The trigger is form shape, not submission state: Submitted.tsx:47 — const ticketNotCreated = happyfoxCount > 0 && !result.happyfox_ticket_id; where happyfoxCount is routing.happyfox.count, computed in routes_phase2.py by counting the form's happyfox_template actions. change-of-address-notification.yaml declares 2 (destination: 92, destination: 96); facilities-support-request.yaml declares 0 — which is the whole of why one showed the row and the other did not. No ticket is ever created by the API path: gpus-forms-routing-worker/routing_worker.py:683-694 writes status='deferred' for every happyfox_template action and dispatches nothing, pending 2.5(e). But the email path is a different story, and it is documented: architecture/forms-phase2.5c-design.md:157 records the G4.4 finding of 2026-06-12 — "some email_template destinations are HappyFox email-ingest addresses — the Gate-4 close-out email to gpus-it-support@greenpeace.org auto-opened ticket #USITS00357259 via email-to-ticket. The worker dispatched no HappyFox action (verified)." Change of Address's two email legs go to gpus-finance-support@ and gpus-people@, from the same seven-address gpus-*@greenpeace.org family. |
| acceptance test | For each of the seven email destinations, it is recorded whether that inbox opens a HappyFox ticket on receipt — the T3 per-form deconfliction answer. Then the confirmation page says only what follows from it. A page that cannot determine ticket state says nothing about tickets; it must not assert a negative. Re-runnable: a submission on a form whose email leg IS an ingest inbox must not render "Not created". |
| blocks | — |
| blocked_by | — CLEARED 2026-08-26. The row was blocked on T3 while the copy tried to report the OUTCOME (does a ticket exist?), which needs a per-mailbox answer. It is unblocked by no longer making that claim: the copy now describes the MECHANISM and the portal's own blindness, both of which are knowable today. T3 still gates 2.5(e) and still gates ever showing a ticket NUMBER. |
| date_raised | 2026-08-26 |
| date_verified | — |
CONFIRMED, then FIXED — 2026-08-26
The finding is no longer inferential. Ticket #USFIN00367708 opened
in US - Finance Support for change-of-address submission b5e998f7,
delivered by email, Requestor: Alerts - alerts@greenpeace.us. A ticket
existed while the confirmation page said one had not been created. The
routing is correct — the body concerns reimbursement addresses and Finance
is the right queue. Only the copy was wrong.
Note the requestor: alerts@greenpeace.us, the relay-rewritten sender
(the send-as decision). The ticket is attributed to the alerts mailbox
rather than to the submitter — a separate consequence of that decision, not
tracked here, and worth knowing before anyone tries to correlate tickets to
people.
Fixed in ee2da8a. The row now reads:
Ticket —
Not shown here· the team's system assigns it when the email arrives
It states what the portal did and what it cannot see, and asserts nothing about whether a ticket exists — the portal cannot observe that either way. Neutral, not amber: the warning colour did as much of the work as the word "not", and there is no fault here to warn about.
Also gated on an email leg existing, which is not cosmetic. Of the 12
live forms declaring a happyfox_template action,
new-employee-notification-contractor-intern declares no email action at
all — nothing is dispatched and nothing is emailed, so no mailbox can open
anything, and wording about email ingest would be false for exactly that
form. The headline prose already tells that submitter their request has not
been sent to a team, which is the honest place for it.
Not overcorrected. "A ticket has been created" would have been the mirror-image defect: the portal cannot see that either. The copy describes the mechanism and its own blindness, both of which are knowable today, and neither of which is a claim about this particular ticket.
The assertion the portal is not entitled to make
Three separate things are true at once, and the row collapses them into one false sentence:
| Fact | Confidence |
|---|---|
| No ticket was created by API dispatch | Certain. The worker defers every happyfox_template and dispatches nothing. |
| A ticket may have been created by email ingest | Unknown per address, and demonstrated for at least one. gpus-it-support@ provably opens tickets (#USITS00357259). |
| The email was handed off for delivery | True, and it is what the small print beside the pill already says. |
"Not created" reports the first as though it were the whole picture. The honest statement is "the portal cannot tell you the ticket number" — which is a statement about the portal, not about the ticket.
This is not pedantry about wording. The submitter's decision is whether to chase. Told a ticket was not created, the reasonable action is to contact IT and open one — and if the ingest path did its job, they have now caused a duplicate for the queue to close, which is precisely the double-ticket failure T3 exists to prevent. The copy can generate the outcome it is warning about.
RESIDUAL — 17 forms are told nothing about a ticket that probably exists
Recorded because the 2026-08-26 fix corrected the copy and did not change the rule that decides who sees it. Anyone touching that gate next should know what it actually keys on.
The gate is happyfoxCount > 0 && !happyfox_ticket_id && hasEmailLeg.
In words: this form declares a HappyFox action. It is NOT a ticket can
result from this submission. Those are different questions, and the second
one is the one a submitter is asking.
Counted from the 29 live form definitions on 2026-08-26 (forms/*.yaml
less pulldowns.yaml, templates.yaml and _schema.yaml, which
yaml_loader.py:174 does not load):
| forms | ticket row | a ticket can result? | |
|---|---|---|---|
| HF action + email leg | 11 | shown | yes — by ingest |
| HF action, no email leg | 1 | hidden (correctly) | no — nothing is sent |
| email leg, no HF action | 17 | hidden | yes — by ingest, identically |
| 29 |
The 17 are the residual. Their email goes to the same
gpus-*@greenpeace.org family — gpus-facilities@, gpus-it-accounts@,
gpus-data-request@ and the rest — and ingest is what opens a ticket, not
the happyfox_template action, which dispatches nothing at all. A
Facilities Support Request very likely opens a ticket by exactly the
mechanism that opened #USFIN00367708, and its submitter is told nothing
about tickets whatsoever.
This is the same asymmetry as this row, pointing the other way. The original defect was a form being told something false about a ticket. This is a form being told nothing about one that probably exists. Both come from the same root: the row is gated on a declared action rather than on whether a ticket can result.
Not a new finding and not wrong. Silence is not a false statement, and saying nothing is strictly better than the "Not created" it replaced. It is logged here rather than raised because the fix is the same fix — and it is the same blocked answer: which of the seven destination mailboxes are HappyFox-watched. Once T3 answers that, the honest gate is "this submission's email goes to a watched mailbox", which would show the row on the 17 and could remove it from any of the 11 whose mailbox is not watched.
Do not tighten the gate before T3. Showing the row on all 28 email-bearing forms would assert the ingest mechanism for mailboxes nobody has confirmed are watched — trading a silence for a guess.
Coverage is inverted — the form most likely to have opened a ticket says nothing
Worth stating plainly because it is the opposite of what the row's presence implies.
- Change of Address — 2 email legs + 2
happyfox_templatelegs → shows "Not created". It has the higher chance of a ticket existing, because it has two email legs into thegpus-*family. - Facilities Support — 1 email leg, 0
happyfox_templatelegs → shows nothing about tickets at all. Its email togpus-facilities@greenpeace.orgmay equally have opened one, and the submitter is told nothing either way.
The row keys off the unbuilt API path while the thing that actually creates tickets today — email ingest — is invisible to it. A submitter reading both confirmations would conclude the facilities request is the cleaner outcome. Nothing supports that.
Why its OWN row and not part of VLN-044
Both are on the same page, in the same component, found in the same session. They are still different findings.
- VLN-044 is a promise about a workflow that does not exist, gated by a tautology. Fix: gate it on a fact the backend already has. Owned entirely within this repo; shipped the same day it was found.
- VLN-046 is an assertion about a downstream system's state that this repo cannot determine from anything it holds. No gate fixes it. The fix requires an answer from outside the code — which of seven mailboxes are HappyFox-watched — and that answer is T3, which also blocks 2.5(e).
Different fix, different acceptance test, different blocking dependency, and one closes while the other waits on a mailbox audit. Folding this into VLN-044 would have closed it by association when the panel shipped — exactly the reason VLN-041 was separated from VLN-040.
The T3 block is cleared, and how is the point. This row was
blocked_by T3 while the copy tried to report the outcome — does a
ticket exist? — which genuinely needs a per-mailbox answer nobody has. It is
unblocked not by getting that answer but by no longer making that claim.
The mechanism ("the team's system assigns it when the email arrives") and
the portal's blindness ("not shown here") are both knowable today.
T3 still gates 2.5(e), and still gates ever showing a ticket number. What it does not gate is telling the truth about what the portal can see — which was available the whole time.
VLN-047 — The revise path shipped complete and unreachable, and the page it should start from says it does not exist¶
| Field | Value |
|---|---|
| id | VLN-047 |
| title | Migration 022's revise path — /revision-source, prefill, parent-linked submit, chain rendering — is deployed and working, and no UI anywhere constructs the URL that starts it. /status/<id>, the one page a returned request lands on, asserts in bold that the request "cannot be edited" and offers a button that produces an unlinked resubmission. Two defects: an absent entry point, and a surface that actively contradicts the feature while silently generating the wrong data shape. |
| register | Security findings |
| status | done — CLOSED 2026-08-26. Both halves of the acceptance test met and observed in the database; see the closure note. Frontend revision gpus-forms-frontend-00021-dlw, worker bc0c84b, both content-verified |
| severity | MEDIUM (argued) — no disclosure and no authz failure; resolve_revision_access is sound and was never bypassed. Argued up from LOW on a ground the copy findings lack: a blocked staff member who reported it rather than working around it — a live, staff-facing feature was unreachable to every user for a week while the page a returned traveller lands on denied it existed. Stated precisely: no real trip is known to have been blocked. The submission that surfaced it, 81c74ae9-…, was returned by IT with the comment "I am returning this as a test"; what was real was the user, who could not find the control and asked. The defect is the unreachability, which was total; the discovery was a test return that a real user then walked into. A second ground, "one orphan resubmission already in submissions", was asserted here on 2026-08-26 and WITHDRAWN the same day — see the population note; it was an inference from timing that the field data contradicts. Argued, not supplied. |
| owner | unassigned |
| evidence | Raised 2026-08-26 from Tanu Garg's report on 81c74ae9-22b0-451e-be8c-6123b39c0a18 (RETURNED_FOR_REVISION, returned at step 2, 14:33:24 UTC), verbatim: "how do I edit the same response? I have an option to submit a new request, but do not see re-submit the request." Both halves of that sentence are the two defects. (a) grep -rn "revise" forms-frontend/src returns seven hits: six consume the parameter (FormFill.tsx:29,48,67,69,85,90), one is a stale comment (SubmissionStatus.tsx:30), zero construct it. No template literal, no searchParams.set, no navigate() carrying it. The URL is reachable only by typing it with a UUID already in hand. FormFill.tsx:29 documents a caller that was never written: "?revise=<parent_id> is how the SPA carries the parent, set by the link on /my-requests." There is no such link — MyRequests.tsx contains no RETURNED branch at all, and its whole card is one <Link to={/status/…}>. (b) SubmissionStatus.tsx:152-154 renders <strong>This request cannot be edited and will not move any further.</strong> followed by the only control on the panel, <Link to={TRAVEL_FORM_PATH}> — '/forms/travel-request', bare, no ?revise=. Commit 65350f8 ("wire the revise path", 2026-08-25) touched four frontend files — api.ts, contract.ts, FormFill.tsx, MyRequests.tsx. SubmissionStatus.tsx is not in its file list. The +45 lines to MyRequests.tsx are ChainLine and its call site: rendering of an existing chain, not creation of one. |
| acceptance test | Two halves, both required. (1) A submitter reaches the revise path from the UI alone — from /status/<id> on a returned request, without typing a URL — and the prefilled form opens carrying the returning approver's comment. (2) The submission that results carries a non-NULL parent_submission_id, observed in the database. The second half is what distinguishes a fix from a link that produces orphans: a revise button pointing at the bare form would pass half 1 and reproduce the exact defect the row exists for. Half 2 additionally proves the chain restarts at step 1 and ChainLine renders on both ends. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-26 |
| date_verified | 2026-08-26 |
CLOSED 2026-08-26 — the column was written by the real path, by the person who reported it
Both halves, observed in gpus_forms rather than inferred from a green
deploy. The verification was Tanu Garg's, which is correct: she found it.
17:08:58 submission_viewed + decrypt_success on 81c74ae9… actor tgarg
17:09:03 submission_viewed + decrypt_success on 81c74ae9… actor tgarg
17:09:33 35750caf-80b7-4879-adf7-ecf174ce1bcf submitted
parent_submission_id = 81c74ae9-22b0-451e-be8c-6123b39c0a18
Half 2 — the parent is recorded. SELECT count(*) FROM submissions WHERE
parent_submission_id IS NOT NULL went from 0 to 1. That column has
existed since migration 022 on 2026-08-25 and had never been written by any
path; this is the first time, and it is what separates a fix from a link
that produces orphans.
Half 1 — reached from the UI. The audit pair at 17:08:58 and 17:09:03 is
GET /revision-source, the portal's fourth decrypt path, which is called
only when ?revise=<parent_id> is present — and /status/<id> is the only
thing in the SPA that constructs that URL. Stated precisely: the audit log
proves the endpoint was exercised, not which control produced the URL,
since a hand-typed URL is indistinguishable here. Typing it is not a
plausible reading — she had reported two and a half hours earlier that she
could not find any way to do this — but the distinction is recorded rather
than glossed.
The chain restarted and completed. The revision's steps are
1 SKIPPED_DUPLICATE (Kevin Toruno) · 2 APPROVED (Rajesh Chhetry) ·
3 APPROVED (Kevin Toruno), round state APPROVED. A fresh round from step
1, exactly as designed — no special case, because a revision routes,
finalizes and dispatches through the same code as any other submission.
What this does not close. VLN-049:
the form's own instructions went on telling every submitter a returned
request "cannot be edited" for a further day after this row closed. The
control existed and worked; the page introducing the form still said it did
not.
Population — ZERO. This note asserted one orphan on 2026-08-26 and was WRONG.
Corrected 2026-08-26, hours after it was written, before anyone acted on it. The retraction is kept in place rather than edited away, because the row's whole subject is a claim that outlived its evidence.
What the row originally said. That
1b4dfd50-fe3c-4cfc-b3e0-c8dada16c713 was an orphan resubmission of
6418f160-…: Tanu Garg was returned, followed the "Submit a new request"
button, and filed a request the database could not connect to the one it
replaced. The reasoning was a timeline:
08-19 16:29:00 6418f160-… tgarg submitted
08-19 16:30:49 6418f160-… RETURNED at step 1
08-19 17:44:55 1b4dfd50-… tgarg submitted ← 74 min later, parent NULL
08-19 17:46:16 1b4dfd50-… step 1 APPROVED
Same person, same form, 74 minutes after a return. It reads as obvious.
The field data says otherwise, and it was never checked. The return
comment on 6418f160-… step 1 asks for exactly one thing:
Lodging is too much
6418f160-… (returned) |
1b4dfd50-… (the supposed revision) |
|
|---|---|---|
| Lodging | 1000 | 1000 — unchanged |
| Airfare | 400 | 150 |
| Depart | 2026-10-10 | 2026-09-10 |
| Return | 2026-12-10 | 2026-10-20 |
The one field the approver asked to change is identical. Every field they did not mention changed. A revision answering "lodging is too much" that leaves lodging untouched and moves both dates by a month is not a revision. These are two different trips.
So the population is ZERO — no orphan, and no known instance of anyone
following the button into an unlinked resubmission. All 11
travel-request-001 rows have parent_submission_id IS NULL, and
SELECT count(*) … WHERE parent_submission_id IS NOT NULL returns 0
across the whole table. That is fully explained by "no revision has ever
been created" — the entry point's absence stated as data — and needs no
second explanation.
R5's August trip count is therefore NOT off by one. 1b4dfd50-… is its
own trip and counting it as one is correct.
What this cost, and why it is recorded here rather than quietly fixed.
The claim shipped in two commit messages (01957b4, bc0c84b), a test
docstring, and this row's severity argument, which cited "a realized data
defect" as a ground for MEDIUM. It was an inference from timing, written
with the confidence of a query result, in a row whose own thesis is that a
well-argued statement defends a stale conclusion better than a bare one.
The register reproduced the defect it was describing, in the paragraph
describing it. The evidence that settles it — one comment column and
four searchable_values keys — was one query away the entire time and was
not run, because the timeline already looked like an answer.
Severity is unchanged at MEDIUM, re-argued. The ground that survives is the realized one: a live, staff-facing feature was unreachable for a week and the page a returned traveller lands on actively denied it existed, with a confirmed blocked user who reported it. The orphan ground is withdrawn.
Nothing to link retroactively. The Part-4 data decision this row was
carrying — whether to backfill parent_submission_id on 1b4dfd50-… — is
moot, and backfilling it would now be the error: it would assert a
supersession between two unrelated trips, inside a live approval record
currently sitting at step 2. Closed as "no action, premise false" rather
than as "decided against".
TWO defects in one row, and why they are not the same defect
Distinguished because the fixes are different and half a fix is worse than none.
- (a) No entry point. The feature is complete and has no front door.
Everything from
/revision-sourcethrough parent-linked submit toChainLineshipped and works; nothing calls it. Fix: construct the URL somewhere a submitter will find it. - (b) An active contradiction.
/status/<id>does not merely omit the control — it denies the control exists, in bold, and offers a button that produces an orphan. Fix: delete the false claim and remove the orphan-producing button.
(b) is the more serious half and it is not a copy defect. A page that silently says nothing about revising leaves the submitter stuck. A page that says "cannot be edited" and hands them a button routes them into producing the wrong data shape — which is what happened on 08-19. The false sentence is not the harm; the button beside it is.
Which is why the fix must not leave both controls standing. Two buttons, one of which produces an orphan, is worse than the current state, because today's single wrong button is at least unambiguous.
The mechanism — TRUE when written, made false by a change that never touched it
This is the part worth the row, and it is the same class as VLN-045 rather than a repeat of the copy findings.
SubmissionStatus.tsx's comment was correct on the day it was written,
1daeea0, 2026-08-20:
* READ ONLY. There is no edit and no resubmit-in-place, because the revision
* path is blocked on an unsettled schema question about how a revised
* submission chains to the original (ASVS §2.4). The page therefore never
* offers a control that does not exist — on a returned request it says a NEW
* request is required and links the form, which is the same thing the outcome
* email says, so the two cannot tell the traveller different stories.
Every clause was true. ASVS §2.4 did record the revise endpoint as "NOT BUILT, PROVISIONAL". The schema question was genuinely unsettled. And the reasoning is good — it refuses to name a control that does not exist, precisely so as not to send a traveller hunting for a missing button. The inline comment at the panel says it again:
NO EDIT PATH EXISTS. Saying "update your request" would send the
traveller looking for a button that is not there, and someone who
believes they have edited a request that never changed is worse off
than someone told plainly to start again. Same wording as the outcome
email, deliberately.
Migration 022 settled the schema question on 2026-08-25 — Option A, a
revision is a new row with parent_submission_id. 65350f8 built
everything on top of it the same day. The refusal outlived its reason by
one day and now says the opposite of the truth in the same words.
The care is what makes it dangerous. Both comments explain why the control
is absent so convincingly that a reader arriving later — including the
author of 65350f8 — is told, in the file, that this page is correct as it
stands. A well-argued comment defends a stale conclusion better than a bare
one does.
Three rows in this family now split two ways
Asked for explicitly, because the split decides what test is worth writing.
Wrong when written — GOV-021's served instructions, the WITHDRAWN
fallback, and VLN-044's
panel. Each was authored against a mistaken understanding of the system. A
copy snapshot test written at authoring time would have been written against
the same mistake, passed, and then defended the error.
True when written, made false elsewhere — VLN-045 and this row.
VLN-045's <h1>Travel request.</h1> was true while /status served travel
requests only, and was falsified by resolve_submitter_access gaining no
form filter three files away. This row's "cannot be edited" was true while
no revise path existed, and was falsified by migration 022 in a different
directory. In both, a snapshot test passes before the change and passes
after it, because the literal never moves.
The difference between the two halves of the family: for the first, the fix is to correct a statement. For the second, the fix is to pin the relationship rather than the value — assert that the view's claim is a function of the fact it depends on, so it cannot survive the fact changing. That is what the structural test below does, and it is the only test here that would have caught this before Tanu Garg did.
Register mechanics — VLN-044 predicted this exactly, one day early, and nothing acted on it
Recorded as a finding about the register itself, not as a footnote.
VLN-044 already contained this, written 2026-08-26, the day before the report:
The unexercised remainder is named, not implied: the revise path (
/revision-source, prefill, resubmit) has never been driven through a browser end to end. It is the only piece of what shipped to staff still in that state.
It named the exact path, the exact reason, and the exact risk — three lines below VLN-044's own lesson that "the estate's verification has been strong on reading and weak on using." It was right. It changed nothing.
Because it was prose inside another row. It had no id, so nothing could reference it; no owner, so nobody held it; no acceptance test, so no build and no review could fail on it; and no status, so it could not be open. It could not appear in the summary table, could not be triaged, and closed silently when VLN-044 closed — which is the failure mode VLN-046's own "Why its OWN row" note warns about, reached from the other direction.
A prediction without an id is a note. The register's rule should be read as: if a sentence in a row identifies unexercised behaviour in shipped code, it is a row. Naming a risk inside another finding feels like diligence and produces none of a row's mechanics.
This is also the second time the estate has been told and not heard: ASVS
§2.4 recorded the endpoint as NOT BUILT, PROVISIONAL, and that line went
stale on 2026-08-25 without anyone revisiting the surfaces that cite it.
Why tsc and eslint both passed, and what does catch it
65350f8's verification line reads "tsc and eslint clean on the four
frontend files touched." True, and structurally incapable of finding this.
A query parameter that is consumed but never produced is perfectly typed.
searchParams.get('revise') returns string | null whether or not anything
in the universe ever sets it. null is a legal, handled value — FormFill
branches on it correctly and renders a normal blank form. There is no unused
symbol for eslint, no type error for tsc, no dead code: every line of the
revise path is reachable, just never reached. The feature's absence has
the same type signature as its presence.
Compounding it: forms-frontend has no test runner at all. No vitest, no
jest, zero test files, and forms-frontend/cloudbuild.yaml invokes no test
step. check-test-gates.sh cannot see this, because it enumerates
directories containing test_*.py — a component with no Python tests is
invisible to the gate that exists to find ungated components. VLN-038 closed
the backend half; the SPA was never in scope.
What catches it is an assertion over the relationship between two files: every query parameter read by a route must be constructed somewhere in the SPA. It is cheap, it needs no DOM, and it fails on exactly this shape — a reader with no writer — without knowing anything about revisions.
VLN-048 — The portal-coverage gate passes every cloud_services entity automatically, forever¶
| Field | Value |
|---|---|
| id | VLN-048 |
| title | validate_portal_presence() computes present = (sentinel in blob) or (eid.lower() in blob), and both portals carry the cloud-services-render sentinel in a comment. The or is therefore satisfied for every cloud_services entity by a string that has nothing to do with that entity, so the gate passes with zero portal edits and cannot distinguish "this entity is rendered" from "a render path exists somewhere on this page". It is the only automated enforcement of the three-portal rule for cloud services, and it enforces nothing. |
| register | Security findings |
| status | open |
| severity | MEDIUM (argued) — no disclosure and no runtime impact. Argued at MEDIUM rather than LOW because this is a control that reports success without testing anything, and the estate has already acted on its green: the Component Coverage Standard tells contributors the pipeline enforces portal presence, CLAUDE.md repeats it, and two live internet-facing Cloud Run services sat unrendered on both portals for three weeks with the gate green throughout. A check whose pass carries no information is worse than an absent check, because an absent check does not get cited. Argued, not supplied. |
| owner | unassigned |
| evidence | scripts/check-component-coverage.py:557,564,584-608. The sentinel constant is PORTAL_RENDER_SENTINEL = "cloud-services-render" and the presence expression is present = (sentinel in blob) or (eid.lower() in blob). _portal_source_blob() concatenates every *.html/*.js in the portal directory including comments, so the sentinel is matched inside an HTML comment. Both portals carry one — status-site/index.html:318 and soc-site/index.html:1215, each reading <!-- cloud-services-render \| T5: replace this hardcoded block with an inventory.json render … -->. Demonstrated, not theorised: gpus_forms_frontend and gpus_forms_backend have been registered in inventory.yaml since 2026-08-05 and appear as a literal data-svc-id on neither portal; they passed every coverage run for three weeks purely on the sentinel. Confirmed by scanning both portals for each ID on 2026-08-26 — 7 of 9 entries rendered, 2 absent, gate green. |
| acceptance test | Re-runnable, and it must fail on a sentinel match — that is the whole point, so the test is written against data-svc-id attributes rather than raw substrings:python3 - <<'EOF'import re, sys, yamlcs = yaml.safe_load(open("inventory.yaml"))["inventory"]["cloud_services"]gaps = [f"{d}: {e}" for d in ("status-site","soc-site")for e, v in cs.items()if v.get("monitoring_status") != "decommissioned"and e not in set(re.findall(r'data-svc-id="([^"]+)"', open(f"{d}/index.html", errors="replace").read()))]print("\n".join(gaps) or "PASS"); sys.exit(1 if gaps else 0)EOFBoth directions required. (1) It must exit 0 on the current tree. (2) It must exit non-zero on the tree as it stood before 2026-08-26, where it reports gpus_forms_frontend and gpus_forms_backend missing from both portals — the two entries validate_portal_presence() passes today. A replacement gate that cannot fail on that input has reproduced this finding. Verified in both directions on 2026-08-26. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-26 |
| date_verified | — |
The docstring overclaims in a second, independent way
PORTAL_DIRS = ("status-site", "soc-site")
def validate_portal_presence(...):
"""Every non-decommissioned cloud_services entity must be present in
BOTH status-site and soc-site render sources (literal ID or generic
render sentinel). Makes the three-portal coverage requirement an
enforced gate for cloud_services entities."""
PORTAL_DIRS has two entries. The "three-portal coverage requirement"
is enforced across two portals. The third, mkdocs-portal, is not checked
here at all — it is covered by validate_documented_in(), which is a
materially weaker check:
| Condition | Severity |
|---|---|
documented_in is empty |
error — fails the build |
| a listed path does not exist | warning |
| a listed path exists but never mentions the entity | warning |
So for the third portal, listing one path that exists and says nothing is a pass. There are 122 such warnings in the current run and they have never failed anything.
Two separate overclaims, then, in eight lines: the sentinel makes the
two-portal check unconditional, and the docstring calls a two-portal check
a three-portal one. Neither is visible to anyone reading the summary line
RESULT: PASS.
Why this outranks the task that found it
Raised while propagating the travel approval workflow into inventory.yaml
and both portals. That work is four entries and eight portal rows. This is
the reason such work can be skipped and still look done.
The propagation rule is real and is stated in three places — the Component
Coverage Standard, CLAUDE.md's "Adding Infrastructure" section, and the
quick-start. All three tell a contributor that the pre-build check enforces
it and that drift fails Cloud Build. For cloud_services and portal
presence specifically, that is not true and has not been true since the
sentinel was introduced. A contributor who adds an entry and no portal row
gets a green build and a documented assurance that green means covered.
The estate's recurring failure mode is an artifact that is internally
coherent and false against a fact living elsewhere. This is that shape in a
control: validate_portal_presence is correct on its own terms — it
genuinely tests what its expression says — and false against the claim the
standard makes about it.
The sentinel is not a mistake, and it should not simply be deleted
Recorded because the obvious fix is wrong.
The sentinel exists for a real reason, stated at its definition: the T5
refactor will replace the hardcoded portal blocks with an inventory.json
render, and after that refactor no literal entity ID will appear in the
portal source at all — the rows will be generated at runtime from
window._inventory. A gate that demanded literal IDs would fail on the day
the render path became correct, which is exactly backwards. The or was put
there so the check would survive that transition.
So the fix is not "drop the sentinel". It is to make each branch prove something:
- Pre-T5 (today). Require a per-entity marker —
data-svc-id="<id>"— rather than any substring. That is already the convention every rendered row follows, so the check costs nothing and the acceptance test above passes on the current tree. - Post-T5. The sentinel branch is the right idea but must be paired with
evidence that the render path actually enumerates the inventory — e.g.
asserting the block iterates
cloud_servicesand that the bakedinventory.jsoncontains the entity. Presence of a comment is not that evidence.
Two more sentinel occurrences were added on 2026-08-26 by the very
commit that fixed the underlying coverage gap (status-site/index.html:335,
soc-site/index.html:1234), marking the new approval-workflow tables as
T5-replaceable alongside the existing ones. Correct on its own terms, and it
also means the free pass is now four comments wide. That is how a check like
this erodes — not by anyone disabling it, but by ordinary correct work
adding more of the thing it accepts.
Scope — what this does and does not cover
validate_portal_presencealso checkscompute_instancesand any asset taggedproject: gpus-it-infrastructure, againstCOMPUTE_RENDER_SENTINEL=compute-instances-render. Same defect, same shape — both portals carry that sentinel too, so all nine compute instances pass unconditionally. They happen to be genuinely rendered today; nothing checks that they stay so.- The IAR sub-check inside the same function (
eid.lower() not in iar_blob) is a real check — it matches the entity ID against the file's text and can fail. It applies only togpus-it-infrastructureassets, so nogpus-infracloud service is covered by it. network_devices,hypervisors,power_devices,storageandlinux_hostshave no portal-presence check at all, by design. Not part of this finding; noted so a reader does not infer coverage from its absence.
VLN-049 — The form's own instructions still deny the revise path, and a test keeps them that way¶
| Field | Value |
|---|---|
| id | VLN-049 |
| title | VLN-047 fixed /status/<id> and the outcome email. It did not reach the third surface that carries the same false sentence — the forms.instructions column served at the top of the travel form itself, which reads "If your request is returned, it cannot be edited. Read the comment, then submit a new request." This is the same column GOV-021 was raised and closed about, wrong again for a different reason. A passing test requires it to say this, so correcting it fails the suite. |
| register | Security findings |
| status | done — CLOSED 2026-08-27. All four surfaces fixed in the version: 3 bump and the served column verified against gpus_forms, which is this row's acceptance test and GOV-021's standard |
| severity | MEDIUM (argued) — same class and the same realized-harm ground as VLN-047, argued at the same level rather than lower because this surface is read earlier. /status/<id> is seen after a return; the instructions are seen by every submitter before they fill the form in, so this is where the expectation is set. A traveller told at submission time that returns cannot be edited will not go looking for the control when one is returned — the fix shipped, and the form still teaches people it does not exist. |
| owner | unassigned |
| evidence | Live, not just in the repo. SELECT current_version, md5(instructions), length(instructions), (instructions LIKE '%cannot be edited%') FROM forms WHERE id='travel-request-001' returned 2 \| e9b6e1acc6bd0b0bf3ce8f2aacc1bee9 \| 570 \| t on 2026-08-27 — byte-identical to forms/travel-request.yaml, so the served text is exactly the wrong text. Four places carry it: (1) the YAML instructions block; (2) forms-backend/test_travel_instructions.py::test_returned_requests_are_described_as_a_new_submission, which asserts the text must say a returned request needs a new submission — a green test pinning the false copy, the same shape as the outcome-mail guard evolved in bc0c84b; (3) forms-backend/approval_decide.py:128-130, whose comment says "What is still absent is a submitter STATUS view and an edit path" — the status view shipped 2026-08-19 and the revise path 2026-08-25, so both clauses are false and the test above cites this comment as its contract; (4) approval_decide.py:138, which still describes a request moving "from Admin Ops to Finance" — step 3 is Final approval, settled in 07d5766 and recorded in the platform doc. |
| acceptance test | Two halves. (1) The live forms.instructions column for travel-request-001 describes the revise path — verified by query against gpus_forms, not by reading the YAML and not by a green build, per GOV-021's own closing standard and VLN-014's worked example of why. (2) test_travel_instructions passes with an assertion that the text describes revising, and fails if the text says a returned request cannot be edited — the guard must be inverted, not deleted, or the copy is left unprotected in both directions. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-27 |
| date_verified | 2026-08-27 |
CLOSED 2026-08-27 — verified in the database, not in the YAML
Both halves of the acceptance test.
(1) The served column. Read back from gpus_forms after the v3 reload:
current_version | md5(instructions) | length | cannot_be_edited | says_revise
3 | 721984482ad6de97b0fa8465561add09 | 780 | f | t
The md5 and length were recomputed from forms/travel-request.yaml locally
and match exactly, so the text staff are served is the text that was
written — the same check GOV-021 closed on, for the same reason: reading
the YAML proves what was intended, not what is being served.
(2) The guard is inverted, not deleted.
test_returned_requests_are_described_as_revisable now asserts "revise"
present and "cannot be edited" absent, and
test_a_revision_is_still_never_described_as_an_in_place_edit keeps the
half that survives. Mutation-proved both ways: restoring the old sentence
fails the first, and phrasing revise as "just edit your request and
resubmit" fails the second.
All four surfaces landed in the same version: 3 bump as the
funding-source fields, so the correction cost one line and no additional
deploy — the sequencing note below, acted on rather than recorded and
deferred.
THE DURABLE LESSON — VLN-047's sweep grepped SOURCE; this lives in DATA
The reason this survived, stated as the class rather than the instance.
VLN-047 enumerated every surface that denied the revise path and found two —
SubmissionStatus.tsx and approval_render.py. Both are code, and the
search that found them was a grep over source files for the sentence and for
revise. It was a thorough search of the wrong space.
The instructions are data. They live in a YAML block scalar
(forms/travel-request.yaml) and, after yaml_loader._upsert_form runs on
the next container boot, in a database column (forms.instructions).
Neither is a source file in the sense the sweep meant, and a grep of
*.py / *.ts / *.tsx cannot see either.
This is the same blind spot that made GOV-021 possible, on the same
column. That row's whole finding was that the form's served instructions
described a workflow the resolver does not perform, and it survived because
"the instructions and STEP_DESCRIPTIONS are both machine-readable and
nothing compared them". GOV-021 was closed by fixing the sentence and
adding test_travel_instructions.py. What was not done was widening the
definition of "a surface" for the next sweep — so the next time a claim
went stale, the same column was missed again, by a search that was
otherwise careful.
What a sweep for this class would have to cover. Not "grep harder" — a different and enumerable set of places prose about system behaviour is served from:
| Class | Where | Reachable by a source grep? |
|---|---|---|
| Component source | *.tsx, *.ts, *.py |
yes — this is what VLN-047 searched |
| Assembled mail bodies | approval_render.py, approval_dispatch.py |
yes, the literals are in source |
| Form definitions | forms/*.yaml — instructions, every field description, every pulldowns.yaml value |
no — data, not code |
The live forms table |
forms.instructions, fields.description |
no — and can drift from the YAML independently |
| Email templates | forms/templates.yaml → the templates table |
no |
| Published docs | mkdocs-portal/docs/** |
yes |
| The registers themselves | this file, gap-register.md |
yes, and they go stale too — see VLN-047's own withdrawn population note |
The two rows in bold are the ones both sweeps missed, and they are precisely
the surfaces a submitter reads. The estate's own guides
(docs/guides/travel-*.md) are a further copy of the same claims, correct
today only because someone remembered them.
So the standard a claim-staleness sweep has to meet: enumerate by
audience, not by file extension. Ask who is told this, and then find every
place they are told it — which for a submitter means the form's instructions
and field descriptions before it means any .py file. A grep over source is
a search of the developer-facing surfaces only.
A green test is holding the false copy in place
test_returned_requests_are_described_as_a_new_submission asserts
"new request" in text or "submit a new" in text. It was written against a
correct understanding — no edit path existed — and it is now the reason the
copy cannot be corrected without a red suite. Anyone fixing the sentence
sees a failing test and has to decide whether they have broken something.
Third instance of this exact pattern in four days. bc0c84b evolved
the outcome-mail guard (test_returned_says_a_NEW_request_is_required),
01957b4 evolved two more in the frontend, and this is the same shape
again: a guard written against a true fact, outliving the fact, and
defending the error it was meant to prevent.
The fix follows the precedent already set — evolve, do not delete. The
half that survives is real: a revision is a NEW submission
(parent_submission_id, migration 022), so the instructions must still not
imply this request is edited in place. "Revise" is now correct; "edit your
request" is still wrong.
FIXED 2026-08-27 — four surfaces, one version bump
Folded into the same version: 3 bump as the funding-source fields, on the
sequencing argument below: the reload happens regardless, so the correction
cost one line and no extra deploy.
| # | Surface | Change |
|---|---|---|
| 1 | forms/travel-request.yaml instructions |
"it cannot be edited … submit a new request" → tells the submitter they can revise, from /my-requests, with answers prefilled, and that approval restarts at the first approver |
| 2 | test_travel_instructions.py |
guard evolved, not deleted — now asserts "revise" present and "cannot be edited" absent, plus a second test keeping the half that survives |
| 3 | approval_decide.py:128-130 |
"what is still absent is a submitter STATUS view and an edit path" — both shipped (2026-08-19, 2026-08-25); corrected in place with what changed and what did not |
| 4 | approval_decide.py:138 |
"from Admin Ops to Finance" → "to final approval", matching STEP_DESCRIPTIONS[3] and 07d5766 |
Mutation-proved, each mutant parsing first. Restoring the old sentence
fails test_returned_requests_are_described_as_revisable; rephrasing
"revise" as "just edit your request and resubmit" fails
test_a_revision_is_still_never_described_as_an_in_place_edit. A guard that
only fires in one direction would have accepted the second, which is the
in-place edit the original comment was right to forbid.
Cheap now, dearer later — a sequencing note, not a severity argument
travel-request-001 is being bumped to version: 3 for the funding-source
fields. That bump reloads forms.instructions from the YAML on the next
container boot regardless. Correcting the sentence in the same version
costs one line and no additional deploy; deferring it costs a second
version bump and a second reload, and leaves the wrong text served in the
interval. Recorded as a fact about sequencing so the decision is made
deliberately rather than by default.
Lynis scan findings¶
Daily Lynis scans run at 03:00 on all 4 servers. See Lynis Scan Results for the latest hardening index and warnings per server.
Current hardening indices (2026-03-16):
| Server | Hardening Index | Warnings | Suggestions |
|---|---|---|---|
| SKY | 79/100 | 1 (kernel reboot) | 23 |
| RAIN | 79/100 | 1 (kernel reboot) | 24 |
| SUN | 76/100 | 1 (kernel reboot) | 27 |
| WIND | 76/100 | 1 (kernel reboot) | 28 |
Remediation guidance¶
VLN-001 — SSH MFA¶
# Install Google Authenticator PAM module
dnf install -y google-authenticator pam
# Configure for each service account
su - dnsadmin -c "google-authenticator -t -d -f -r 3 -R 30 -w 3"
# Add to /etc/pam.d/sshd
echo "auth required pam_google_authenticator.so" >> /etc/pam.d/sshd
# Enable ChallengeResponseAuthentication in sshd_config
sed -i 's/ChallengeResponseAuthentication no/ChallengeResponseAuthentication yes/' /etc/ssh/sshd_config
systemctl restart sshd
VLN-006 — DNS recursive query restriction¶
# Add ACL to /etc/named.conf
# acl "internal" { 192.168.120.0/23; 192.168.124.0/24; 172.16.0.0/24; };
# options { allow-recursion { internal; }; };
sudo rndc reload
VLN-050 — The HappyFox ticket body is a third render surface, and the v3 bump left it enumerating a 13-field form¶
| Field | Value |
|---|---|
| id | VLN-050 |
| title | forms/templates.yaml → travelrequest001 is the body of the only ticket a travel request raises (queue 45). It enumerates its fields as literal {{ Placeholder }} tokens rather than iterating them. The version: 3 bump of 2026-08-27 added FundingSource, GrantName and CostCenter to travel-request.yaml — a different file — and the template carries no version of its own, so nothing updated it and nothing failed. The portal and the approval mail show the funding fields; the ticket does not. |
| register | Security findings |
| status | CLOSED 2026-09-01 — both halves met. The template + guard closed 2026-08-31; the second half closed on af7ea350 (Jack Sundius), submitted 13:28 UTC on 2026-09-01, the first submission ever to reach the corrected template. Ticket #USITS00368946 read in queue 45 and confirmed line by line. c84bfb45's short ticket is not repaired and is not a condition of closure — see the closing note |
| severity | LOW → MEDIUM (argued), raised 2026-08-31 — the original LOW rested on realized harm being zero. It is no longer zero: it is one real travel request, c84bfb45, submitted by Madison Carter, whose queue-45 ticket omits the funding attribution she supplied. Still not HIGH — no disclosure, no authz failure, and the approval mail and portal both carried the fields correctly, so the approver was never misinformed. But a durable IT record of a real person's travel is missing two values she entered, and "argued on zero harm" is no longer available as the argument. No disclosure, no authz failure, no data loss: the ticket under-reports fields the triager can see in the portal. Deliberately not argued at MEDIUM like VLN-049, whose surface was read by every submitter before filling the form in. This one is read by one triager, after the fact, beside a portal page that is correct. What raises it above trivia is not this instance — it is that the surface was uncounted, and would have taken a fourth cost field down with it silently. |
| owner | unassigned |
| evidence | Live, not just the repo — the VLN-049 standard. Queried gpus_forms as maple-agent@gpus-infra.iam on 2026-08-31: SELECT id, length(body), md5(body), body LIKE '%FundingSource%' … FROM templates WHERE id='travelrequest001' returned travelrequest001 \| 531 \| 5016c0c4b24e1cb1672b3fcbdef230e1 \| f \| f \| f \| t — no funding fields, CostAirfare present. The md5 and byte length were recomputed from forms/templates.yaml locally and matched exactly, so the served body is the written body. Meanwhile SELECT id, current_version FROM forms WHERE id='travel-request-001' returned 3. Parsed diff: the form defines 16 field keys, the template rendered 13 placeholders; IN FORM, NOT IN TEMPLATE = ['FundingSource','GrantName','CostCenter'], IN TEMPLATE, NOT IN FORM = []. One action only: SELECT form_id, action_order, action_type, destination, template_id FROM actions WHERE template_id='travelrequest001' → a single email_template to gpus-it-support@greenpeace.org, so there is no per-queue fan-out to make partiality correct. travelrequest001 has been edited exactly once in its life — 78cdd23, 2026-08-18, the commit that created the travel form. It predates version: 2 and version: 3 entirely. |
| acceptance test | forms-backend/test_travel_ticket_template.py — asserts the template's placeholder list equals the form's field-key list, in order. A relationship, not a literal, per this family's own diagnosis. Mutation-proved both directions on 2026-08-31: run against the pre-fix templates.yaml, 4 failures; against the corrected file, 4 passed. Second half, and the one that actually closes it: a real submission's ticket body in queue 45 carries the three funding lines — the YAML is what was intended, the ticket is what a triager reads. NOT MET. The first real submission — c84bfb45 — was rendered by the old template 6m06s before the fix reached the templates row, so its ticket is evidence the defect was real, not evidence it is fixed. The next travel submission is the first that can meet this half. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-31 |
| date_verified | Template + guard 2026-08-31. Ticket-body half MET 2026-09-01 on af7ea350 / #USITS00368946 — submitted 13:28:23 UTC, routed 13:28:26, against a templates row reloaded 12:31:36 |
CLOSED 2026-09-01 — on a real ticket, read line by line, not on the YAML
af7ea350, Jack Sundius, submitted 13:28:23 UTC, routed 13:28:26 —
the first submission ever to reach the corrected template. The
templates row was reloaded at 12:31:36 by a container boot after the
12:50 report deploy window, so the render is 57 minutes downstream of the
fix rather than 6m06s upstream of it, which is what closed c84bfb45's
ticket short.
Ticket #USITS00368946, queue 45, read by R. Chhetry. The four lines
appear consecutively and in form order immediately after Department: —
How is this travel funded?: Cost Center
Grant name:
Cost centre: 20100 — INFORMATION TECHNOLOGY
SMT member you report to: Kevin Toruno
— and Event registration fee (USD): 0 is present, the v4 field
rendering in a real ticket for the first time, two days after it shipped.
WHICH RENDERING SHIPPED, and why the question was asked. GrantName
is legitimately blank here: funding is Cost Center, and the field's own
description says to leave it blank in that case. Two renderings were
possible and they mean different things:
| ships when | reads as | |
|---|---|---|
Grant name: + nothing ← THIS ONE |
the SPA stored an empty-string row | a question asked and answered "none" |
Grant name: [GrantName: —] |
no submission_fields row was stored |
template_render.VisibleUndefined — "the template names a key the form did not supply" |
The empty-string rendering shipped. Confirmed in the ticket, and it
matches the database: GrantName appears exactly once in the whole
submission_fields table — this submission — at ciphertext length 16,
the global minimum across every row. All 17 of 17 declared fields have
stored rows. So the SPA writes a row for a touched-but-empty optional text
field, and {{ GrantName }} is defined and empty rather than undefined.
That distinction is recorded because the marker would also have been a
pass under this row's acceptance test — the label survives either way —
while meaning something quite different about the SPA. A future reader
seeing [GrantName: —] on some later ticket should know it is a change in
behaviour, not the normal blank.
What could NOT have failed, stated so the test is not over-credited.
The labels are literal static text in the template (Grant name:\t{{ GrantName }}).
A label cannot go missing unless the whole line is absent from the
template — which is the original defect, and is now held by
test_travel_ticket_template asserting the placeholder list equals the
form's field-key list in order. What this ticket proves is the part the
test cannot reach: that the deployed templates row, the live renderer and
the queue-45 delivery path all agree with the YAML.
c84bfb45's ticket is still short, and that is deliberately not a condition of closure
Madison Carter's ticket remains missing the funding attribution she
supplied. The correct action is unchanged — a note on the existing
queue-45 ticket carrying the two values, not a re-route: c84bfb45 is
IN_APPROVAL with a real approver holding step 2, there is no re-route
endpoint, and inventing one to re-render a ticket would be a state change
on a live travel authorisation.
This row closes without it because the row is about a render surface, and the surface is now proven correct end to end on a real submission. The outstanding item is a single historical record repair, owned by Rajesh, tracked in the workstream close-out rather than by holding a verified finding open. Holding it open would misreport the estate's state: the defect is fixed and cannot recur.
SUPERSEDED 2026-08-31 — TRUE WHEN WRITTEN, falsified six minutes later. Read the reopening note below it.
CORRECTION to the raising brief — the affected population was ZERO at the time of writing
This row was raised on the understanding that "every HappyFox ticket raised from a travel request since 2026-08-27 is missing the funding fields" and that queue 45's triager "has been reading an incomplete record for four days." Checked, and it did not happen.
Twelve travel submissions all-time, none since the v3 bump, the most
recent one predating it by a day. So no incomplete ticket was ever raised
and nobody has read one. That is consistent with M-01 — the same reason
count(*) WHERE searchable_values ? 'FundingSource' is zero.
The defect was real and the harm was not. Recorded as a correction rather than quietly downgrading the severity, because the difference between "this happened twelve times" and "this has never happened" is the difference between a remediation and a near-miss, and the register should be able to tell them apart afterwards.
It also inverts the sequencing argument in the useful direction. The first travel submission to hit this template would have been the v3 exercise itself — the one submission the estate is running specifically to prove the funding fields reach the surfaces that consume them. Fixing the template before that submission is not tidiness; it is the difference between the exercise proving three surfaces and proving two.
REOPENED — the population became ONE, 6m06s before the fix landed, and it is a real person
c84bfb45-27e2-4250-aa67-62b42b215947, Madison Carter (mcarter@),
submitted 2026-08-31 15:46:12.683Z. The first genuine end-user travel
request since the v3 bump, from someone who did not know any of this had
changed.
| Time (UTC, 2026-08-31) | Event |
|---|---|
| 15:46:12.683 | Madison submits. form_version = 3 — served all 16 fields, answers 15 |
| 15:46:15.668 | Step 1 dispatched to Kalina Francis |
| 15:46:30.259 | Approval mail sent — carries Funded by: and Cost centre: correctly |
| 15:46:32.088 | Ticket email rendered and sent to queue 45 — OLD 13-placeholder template |
| 15:46:32.412 | submission_routed |
| 15:48:38 | The fix commit is pushed; Cloud Build triggers fire |
| 15:52:38.981 | templates.travelrequest001 upserted with the corrected 640-byte body |
Her ticket missed the fix by 6 minutes and 6 seconds.
What queue 45 is holding. Two values she actually entered are absent — not blank, absent:
FundingSource = Cost CenterCostCenter = 23051 — COMMUNICATIONS CORE
GrantName is not among them: she chose the Cost Center branch and left it
blank, so it has no row in submission_fields at all (15 rows, not 16).
A correct rendering would have shown it as a present label with an empty
value.
The blank-vs-missing distinction — the whole risk — resolves cleanly
here, by luck of this template's shape. A missing placeholder removes the
entire line, label included. So the ticket contains no How is this travel
funded? line and no Cost centre line anywhere. A legitimately blank value
would have shown the label followed by nothing. A triager cannot confuse
the two in this instance because the labels themselves are gone — but that
is a property of how this template is written, not a guarantee. A template
that emitted labels unconditionally would have produced the ambiguous case,
and the next one might.
What was NOT affected, so the blast radius is not overstated: the
approval mail Kalina Francis received carries both funding lines and a
correct Total: 500.00 (lodging 0 rendered as 0.00 and kept, so the
total is complete, not partial); the portal renders all 15 answered
fields in field order; and searchable_values stored the cost centre
byte-identical to the pulldown option — md5 dd61e655b30d38f828c126cccfd9f445
on both sides, 29 bytes over 27 characters. Nothing was lost from the
database and no approver was misinformed. Only the ticket is short.
Repair is by hand, and must not touch the request. c84bfb45 is
IN_APPROVAL with a real approver holding it. There is no re-route
endpoint, and inventing one to re-render a ticket would be a state change on
a live travel authorisation for a real trip. The correct action is a note on
the existing queue-45 ticket carrying the two values.
M-14, applied to this row by events rather than by a sweep
The superseded note above said the affected population was zero. It was true when written — 12 submissions, none since the v3 bump — and a user falsified it six minutes later by filling in a form.
This is M-14 with a new operand, and it is worth recording because the previous three instances were all spatial: a contradicting claim elsewhere on the page, in another file, in the database. This one is temporal. No sweep could have found it; at the moment of writing there was nothing to find.
"Realized harm is zero" is a measurement, not a property. It carries a timestamp whether or not one is written, and it decays. A row argued down to LOW on the basis that nothing has happened yet must record when that was measured and must be re-measured before the row closes — because the interval between raising a defect and fixing it is precisely the interval in which the thing is still broken.
The superseded note argued that fixing before the exercise was "the difference between the exercise proving three surfaces and proving two." The reasoning was right and the timing beat it. The estate did not get to choose which submission was the first to hit this, and the one that arrived was better evidence than the planned test — and a live problem for someone's actual trip.
THE CLASS — 'true when written, made false elsewhere', third instance
Filed against the split this tracker already drew for VLN-045 and VLN-049.
travelrequest001's body was complete and correct on 2026-08-18 against
a 13-field form. It was falsified on 2026-08-27 by an edit to
travel-request.yaml, in a different file, that never touched it. As with
the other two, a snapshot test on the template passes before the change
and after it, because the literal never moves.
What is new here, and why this is worth its own row rather than a footnote
on VLN-049: the other two were surfaces someone had at least thought about
and got wrong. This one was not counted at all. The v3 work identified
"both renderers" — approval_dispatch.py and approval_render.py — and
updated both. There were three.
Why the version column could not catch it. version: is a property of
the form (forms.current_version), and it advanced 2 → 3 correctly. The
template is a separate row in a separate table with no version of its
own, and yaml_loader._load_templates upserts on the template YAML's
content — which had not changed. So the bump was well-formed, the loader was
correct, and the drift is invisible to both.
What a v-bump checklist would have to cover to catch the class
The trigger is not "a YAML changed". It is "the field set changed" — and the file that changes is never the file that goes stale. So a checklist keyed on the edited file cannot work. It has to be keyed on the field set, and it has to enumerate the consumers.
The discriminator is ENUMERATES vs ITERATES. A surface that walks the fields it is given adapts for free. A surface that names them in a literal list must be edited by hand on every field-set change, and nothing tells anyone. Only the second kind belongs on the checklist:
| Surface | Shape | v3 outcome |
|---|---|---|
forms/templates.yaml → the ticket body |
enumerates | ❌ missed — this row |
mkdocs-portal/docs/guides/travel-request-submitting.md |
enumerates (numbered table) | ❌ missed — see below |
approval_dispatch.py — COST_FIELDS + literal sv.get() |
enumerates | ✅ updated |
approval_render.py — READ_KEYS, OUTCOME_READ_KEYS, COST_FIELDS |
enumerates | ✅ updated, and the only one with a structural guard |
gpus-reports/summary_report.py — R1's four key constants |
enumerates | ✅ updated |
forms-frontend/.../ApprovalDetail.tsx — fields.map(...) |
iterates | ✅ safe by construction |
So: four enumerating surfaces were in scope for v3 and two were missed —
both of them outside forms-backend/, which is where the attention was.
The durable form of the checklist is not a checklist. It is the property
approval_render.py already has and the other three did not:
TestReadKeysIsLoadBearing compares the declared key set against the keys
the source actually reads, and fails in both directions. That is why
approval_render.py was updated and the template was not — one of them had
a test that could notice. test_travel_ticket_template.py gives the template
the same property. R1's constants and the docs guide still do not have it.
What cannot be generalised, stated so nobody tries. "Every form's
template must cover every field" is false for this estate and would fail
61 of 67 form/template pairs. Templates fan out per queue: a
multi-action form renders a different template per destination, each scoped
to what that queue needs (89ebf11). Partial is correct there. The strict
rule only holds for a form whose sole action is one email_template,
which is why the new test asserts that premise first and fails if travel
ever gains a second destination.
A FOURTH surface, found while writing this row, and fixed with it
mkdocs-portal/docs/guides/travel-request-submitting.md — the staff-facing
guide, published on infra.greenpeace.us — carried three claims the same
v3 bump falsified:
- "Thirteen questions. All of them are required." Sixteen, and two are not.
- A numbered table of 13 questions, correct through question 2 and wrong from
question 3 on, because the funding fields sit between
DepartmentandSMTApproverin the served order. - "Question 3 matters more than the others" and "the SMT member you named in question 3" — SMT approver is now question 6. Two co-located cross-references to a number the same edit moved.
Corrected here rather than deferred, because deferring it would repeat the exact mechanism this row is about. Also added: an explicit note that nothing validates questions 3–5 against each other, which is the documented consequence of the platform having no conditional display and no server-side cross-field validation — a submitter should be told that before they fill the form in, not discover it.
Two more stale surfaces the sweep found, and what was done with each
Applying M-14 to this row's own fix — grep the estate for the claim you just corrected — turned up two more artifacts the v3 bump falsified, neither of them a render surface:
forms-frontend/src/routes/ApprovalDetail.tsx:295,347— two comments asserting "All thirteen fields" and "scrolled past thirteen fields". Fixed. The componentmaps over what the API gives it, so it was never wrong in behaviour; the number in the prose was a claim nothing checked. It now states no count, and says why.forms-backend/test_approval_access.py:495— a module constantTHIRTEENlisting the pre-v3 field set, consumed bytest_returns_all_thirteen_fields_decryptedand bytest_submission_status.py. Not changed. It is a fixture, not an assertion about the live form: the test proves the endpoint returns what it was handed, which is still true and still worth proving. But the name and the list now describe a form that no longer exists, and a reader would reasonably take it for the travel field set. Renaming it and extending the fixture to sixteen is correct and is deliberately left as its own change, because it touches two test modules and this row is closed on a template.
Not in scope, and named so it is not mistaken for clean
16 of the 17 single-action forms in forms/ show the same drift today.
The travel form is where this class was caught, not the only place it lives.
Widening the guard is real work with a real backlog behind it and it is not
this row; raising it here so the next person does not read
test_travel_ticket_template.py's narrow scope as evidence the rest are fine.
VLN-051 — The approval decision audit row was built for non-repudiation and reaches no consumer¶
| Field | Value |
|---|---|
| id | VLN-051 |
| title | approval_decision_recorded and approval_notification_sent are values 32 and 33 of models.py AUDIT_ACTIONS, and appear in neither ship_actions nor excluded_actions in soc/forms-authz-detection/soc_shipper_allowlist.json — 31 of 33 declared, and the two undeclared are exactly the approval ones. soc-log-shipper builds its filter from ship_actions, so no approval decision in this estate has ever reached the SOC stream. |
| register | Security findings |
| status | open — raised 2026-08-31. No change made to ship_actions or the Wazuh ruleset in this pass, deliberately |
| severity | MEDIUM (argued) — argued on what the row was built to be, not on a hypothetical attack. approval_decide.py writes the audit row inside the decision's transaction, and its own comment states the reasoning: audit_log has UPDATE/DELETE revoked from PUBLIC and from every application role, while approval_steps is UPDATE-able by forms_app and carries no immutability constraint (ASVS fact G). So the immutable audit row is the thing the non-repudiation guarantee actually rests on — §2.2 asserts non-repudiation as the endpoint's product. The row exists, it is correct, it is written atomically with the decision, and nothing consumes it. A control that is implemented and unread is this estate's characteristic failure, and it is the same shape as VLN-011 and VLN-040: the mechanism runs and the result goes nowhere. |
| owner | unassigned |
| evidence | From source, 2026-08-31. forms-backend/models.py:204-222 — AUDIT_ACTIONS is a 33-tuple; the comment at :201-203 records it reaching 33 "after 020 added approval_notification_sent" and requires any enum-altering migration to update it. soc/forms-authz-detection/soc_shipper_allowlist.json — ship_actions has 17 entries, excluded_actions 14; 17 + 14 = 31. Set difference computed directly: AUDIT_ACTIONS − ship_actions − excluded_actions = ['approval_decision_recorded', 'approval_notification_sent'], and (ship_actions ∪ excluded_actions) − AUDIT_ACTIONS = []. The file's own header is the standard it fails: "Do NOT hand-edit… Regenerate via gen_authz_detection.py; a new enum value forces a ship decision." Two enum values were added — migrations 018 and 020 — and neither forced one. The write site is forms-backend/routes_phase2.py:1189-1204 (audit_row, success: True, inside the decision transaction) and the durability reasoning is approval_decide.py:21-30. |
| acceptance test | Two halves, both able to fail. (1) Structural, and it fails today: a test asserting set(AUDIT_ACTIONS) == set(ship_actions) | set(excluded_actions) — i.e. every enum value carries an explicit ship decision, in one direction or the other. This is the file's stated contract and nothing currently enforces it; it belongs beside gen_authz_detection.py so a future enum value cannot be added without one. Note it must assert declaredness, not shipping — "excluded, with a reason" is a valid answer for a high-volume action and the test must accept it. (2) Live, and it is the one that closes the row: after a ship decision is taken and applied, a real approval decision produces a matching event in Wazuh — observed in the SOC stream by submission_id, not inferred from config. Per VLN-010's worked example, a rule present in a file is not a rule that fires. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-31 |
| date_verified | — |
What this row does NOT claim
It does not claim the audit row is missing, wrong, or losable. It is
written in the same transaction as the decision, so a decision without its
audit row cannot arise from the application path — which is precisely why
FP-IR-06 can use audit_id IS NULL
on a decided step as a tampering signal. The durable record is sound.
The finding is that the record is unread. Detection and forensics are
different things: the row is excellent forensics and there is no detection,
because nothing is watching. Whether the right answer is ship_actions, a
scheduled reconciliation query, or an explicit excluded_actions entry with
a reason is a ship decision, and taking it inside a playbook edit or a
register row would be exactly the hand-edit the allowlist header forbids.
Deliberately not fixed in this pass
ship_actions feeds a live log shipper and the Wazuh ruleset feeds live
alerting. Changing either is a detection-config change with a blast radius
on the SOC's alert volume, and it should be made as its own change with its
own verification — not as a side effect of writing up the finding.
Recorded, not actioned.
VLN-052 — A denied approval attempt is logged at level 3 and pages nobody¶
| Field | Value |
|---|---|
| id | VLN-052 |
| title | approval_access's denial reasons are absent from soc/forms-authz-detection/authz_reasons.json, so a denied approval-decision attempt matches none of rules 100031–100036 and falls through to the parent rule 100030 at level 3 — below the level ≥ 10 threshold that auto-creates a SOC ticket. The denial is recorded and is shipped; it simply arrives beneath the floor at which anyone is told. An attempt to forge an approval decision is logged and nobody is paged. |
| register | Security findings |
| status | open — raised 2026-08-31. No change made to authz_reasons.json or the ruleset in this pass, deliberately |
| severity | MEDIUM (argued) — argued at the same level as VLN-051 and for the complementary reason. VLN-051 is the successful decision that is never seen; this is the failed one. Together they mean the approval endpoint — the control approval_access.py exists to enforce, on a live financial authorisation — produces no SOC signal in either outcome. Not argued higher because the denial is genuinely recorded and queryable, 403 is uniform and the endpoint is not an enumeration oracle, and the authorisation itself holds: this is a monitoring gap, not an access-control failure. |
| owner | unassigned |
| evidence | From source, 2026-08-31. The write site is forms-backend/routes_phase2.py:1154-1166: on step is None or not authorizes_decision(reason) it logs approval_decision.denied and writes _audit_v2(session, "auth_failure", …, details={"reason": reason, "endpoint": "approval_decision"}, success=False). auth_failure is in ship_actions, so the event reaches Wazuh. soc/forms-authz-detection/forms_authz_rules.xml:6-11 — rule 100030, <action>auth_failure</action>, level 3, no reason field. Rules 100031–100036 each add <if_sid>100030</if_sid> plus <field name="details.reason">^…$</field>, anchored exact-match. authz_reasons.json lists six reasons — auto_provision_denied, inactive_user, insufficient_role, not_owner, unknown_required_role, user_resolution_failed — all sourced from models.py AUTHZ_DENIAL_REASONS, which is auth_v2's vocabulary. approval_access has a different and disjoint vocabulary (approval_access.py:91 — the reason constants behind authorizes_read / authorizes_decision), none of which appears in that file. With ^…$ anchoring, a non-matching reason cannot fall into 100031–100036; it stops at the parent. The ≥ 10 threshold is the SOC ticketing rule already cited in FP-IR-02. |
| acceptance test | Two halves. (1) Structural, fails today: approval_access's denial-reason constants are the single source for a generated tier in authz_reasons.json, exactly as AUTHZ_DENIAL_REASONS already is for auth_v2 — with a test asserting the generated file covers every constant, so a new reason cannot be added without a level decision. This is the same relationship-not-literal shape as TestReadKeysIsLoadBearing. (2) Live, and it closes the row: a deliberate denied approval attempt against a non-approver identity produces a Wazuh alert at the chosen level, observed via wazuh-logtest-legacy -v and in the live archive. Per VLN-010, a rule that loads is not a rule that fires — that row's whole mechanism was rules that parsed cleanly and were unreachable. |
| blocks | — |
| blocked_by | — |
| date_raised | 2026-08-31 |
| date_verified | — |
The level is a decision, not an oversight to be corrected upward
A denied approval attempt is not automatically level 10. not_owner earns
10 because cross-user ID-walking is an IDOR probe; insufficient_role earns
6 because a role probe is common and mostly benign. The approval_access
reasons are not uniform either — round_not_in_approval fires whenever an
approver opens a request someone else has already decided, which is normal
behaviour, while no_matching_step on a DISPATCHED round is the one that
looks like a forgery attempt.
So the fix is a per-reason tier, not a blanket promotion, and tiering them is the ship decision this row asks for. Setting them all to 10 would produce exactly the alert-fatigue that makes a SOC stop reading, which is a worse outcome than level 3.
Read with VLN-051 — together they are the whole endpoint
Neither row is interesting alone. Together: the approval endpoint records a
successful decision in a table nothing ships, and a denied one at a
level nothing tickets. approval_access.py exists because unauthorised
state transitions on a travel authorisation are the threat, and the SOC
currently cannot see either side of it.
This is why FP-IR-06 states its detection as "a human noticing" and labels its queries forensics rather than monitors. That playbook is honest about the gap; these two rows are the gap.