VLN-010 — Loaded, valid, counted, and structurally unreachable: the 100015 baseline rules can never be evaluated¶
Classification: CONFIDENTIAL — Internal Use Only Document:
security/finding-2026-08-05-wazuh-lolbin-inert.md· v2.1 · 2026-08-12 · GPUS-IT Status: CLOSED 2026-09-08. Deployed 2026-08-11 18:29:50 UTC; D1–D5 passed at deploy; D6(a), D6(b) and D6(c) all measured 2026-09-08.Mail dropped ~95% across the fix boundary, indexed volume held flat at 182–184/day — which is the pass, not a partial one — and the three baselines fire non-zero and exclusively at level 3. BT-003 is proven in production. The rules were loaded, valid and counted, and sorted behind a level-12 catch-all that matched first — so they could never be evaluated. Six hypotheses were eliminated getting here; every one of them assumed the rules were being rejected, and none was.
The fix was verified before deployment, not after. Re-parenting the three baselines as children of
100015via<if_sid>100015</if_sid>was demonstrated on 4.14.4-rc2 — compiled tree inspected, both directions traced — using a zero-touch method that never wrote to/var/ossec(§3.11). §3.9 required this be verified and not asserted; it was.It is now installed and running (§7.2).
wazuh-analysisdPID 1491909 → 2361899. Acceptance D1–D5 passed, with D3/D4/D5 measured against the running daemon rather than the file on disk.The finding is CLOSED. All three parts of D6 reported. (a) mail volume fell ~95% measured at the inbox on the real recipient; (b) indexed volume is flat at 182–184/day against a 183–184/day baseline, which is the intended outcome because the fix removes mail and not documents; (c)
100016/100017/100018are non-zero and exclusively level 3. The compound failure — silence mistaken for suppression — was excluded by measurement rather than assumed away.The result the acceptance table did not ask for is the one worth keeping:
100015has fired 15 times since 2026-08-12, every one at level 12. The daemon noise was removed without touching the detection. That is what the fix existed to do, and it is now measured rather than argued.Filename and title diverge deliberately
The file is still
finding-2026-08-05-wazuh-lolbin-inert.mdand the nav entry still reads "Wazuh LOLBin Baseline Rules Inert". Both are kept so the published URL and every existing cross-reference keep working. The title is what changed: v1.0 and v1.1 said absent from the running ruleset, which is now known to be wrong. The rules are present. They are unreachable. "Inert" in the filename remains accurate as to effect.
1. Summary¶
| Finding ID | VLN-010 |
| Severity | High — detection-quality defect, not an exposure |
| Affected | Wazuh manager MAPLE; rules 100016, 100017, 100018 |
| Discovered | 2026-08-05, read-only against live manager + CEDAR indexer |
| Behaviour confirmed | 2026-08-06 (Step 3 verbose trace) |
| Published | 2026-08-10 |
| Revised | 2026-08-10 (v1.2) — mechanism identified; v1.0's named cause disproved |
| Revised | 2026-08-11 (v2.0) — fix designed, verified and committed. Drafted but never published; superseded by v2.1 the next day. |
| Revised | 2026-08-12 (v2.1) — DEPLOYED; D1–D5 passed; D6 open |
| Revised | 2026-09-08 (v2.2) — D6 measured; (b) and (c) PASS; (a) open, blocked on root |
| Revised | 2026-09-08 (v2.3) — CLOSED. D6(a) measured at the inbox; all three parts pass |
| Mechanism | IDENTIFIED. Loaded but unreachable — sorted behind a level-12 catch-all (§3.5). |
| Fix | VERIFIED then DEPLOYED. <if_sid>100015</if_sid> re-parenting, demonstrated on 4.14.4-rc2 (§3.10), installed 2026-08-11 (§7.2). Committed at git-blob 8ce43cf0. |
| Deploy | DONE 2026-08-11 18:29:50 UTC. D1–D5 passed (§7.2). D6(a)+(b)+(c) all passed 2026-09-08. |
| Closed? | YES — closed 2026-09-08. All of D1–D6 measured and passed. Reproduction method retained (§7.3). |
Rules 100016/100017/100018 were written to suppress ~184 daemon-generated
level-12 alerts per day. The on-disk file is byte-identical to the repo copy,
it sits in a configured rule_dir, and wazuh-analysisd restarted after the
file was written. Every artifact says the change shipped.
The rules have never matched a single event. Zero hits, all-time. 100015
— the rule they were written to get in front of — still fires at its full
pre-tuning rate.
They are not being evaluated and failing to match. They are never evaluated at all.
They are, however, in the ruleset. v1.0 and v1.1 said the opposite — that
the rules were absent from the compiled ruleset — and that was wrong. The
standalone loader counts them (Total rules enabled: '8521', the same number
the daemon logged on Jul 30) and lists all three by id and level with no error,
no warning and nothing ignored.
They are loaded, valid, counted — and structurally unreachable. Wazuh
evaluates sibling rules by level, not in file order, and stops at the first
match. 100015 is level 12 and matches every credential_access event by
design; the baseline rules are level 3 and sort behind it. They can never be
reached for any event they were written to catch (§3.5).
The sharpest variant of the spine yet
The previous findings in this series each had a visible tell: a file absent from the manager, a rule that existed only in prose, a control whose scope stopped at the cloud half of the estate. This one has none. The file is present. Its SHA matches git. The restart happened after the write. Config validation passes. Every check an operator would think to run comes back green — and the control does nothing at all.
And there is no check that would have caught it. The rules load cleanly, the count is correct, validation passes — there is nothing anywhere for a diagnostic to report (§4). Absence of the artifact was never the real failure mode; absence of evidence that the artifact works is.
2. Evidence — the control is inert¶
Measured on the CEDAR indexer (wazuh-alerts-*) on 2026-08-05, the same
corpus soc-backend reads.
| Rule | Purpose | Hits, all-time |
|---|---|---|
100016 |
modulesd / syscheckd / aide, daemon context | 0 |
100017 |
stat, Wazuh SCA context |
0 |
100018 |
tar, estate backup context |
0 |
100015 |
the level-12 rule they sit ahead of | 2,576 in 14d |
100015 daily rate, spanning the deploy:
2026-07-31 74 (partial — window edge)
2026-08-01 182
2026-08-02 184
2026-08-03 184
2026-08-04 184
2026-08-05 108 (partial — day in progress)
Flat through and after the deploy, evenly across all four WDC hosts (sky 649, rain 647, sun 643, wind 636). No inflection. The tuning had no effect of any size.
3. Mechanism — loaded, valid, counted, unreachable¶
Four facts about deployment, then the mechanism itself in §3.5.
3.1 The file is deployed and current. The on-disk copy on MAPLE is
byte-identical to the repo copy at HEAD:
Confirmed from this workstation against git show HEAD:soc/wazuh-rules/gpus-lolbin-rules.xml
on 2026-08-10. It is loaded from a configured rule_dir — not an orphan file
in a directory the manager never reads.
3.2 The restart happened after the write. The file was written
2026-07-30 20:00. wazuh-analysisd started 2026-07-30 20:08:21, with
wazuh-remoted at 20:08:24 and wazuh-modulesd at 20:08:27. Verified live
2026-08-10: PID 1491909, uptime 10 days 19 hours, not restarted since. The
ruleset running in production today is the one loaded on Jul 30, eight minutes
after the file landed.
3.3 The rules never appear in the trace. wazuh-logtest -v on a real
daemon-context event walks the parent chain under 80700 and tries:
100016, 100017 and 100018 are never reached and never named.
Re-confirmed 2026-08-10 with wazuh-logtest-legacy, which loads the ruleset
from disk in its own process rather than querying the running daemon (§3.8).
Same sequence against a freshly parsed ruleset.
v1.0 and v1.1 drew the wrong conclusion from this
Both versions read "absent from the trace" as "absent from the compiled ruleset", on the reasoning that an evaluated-and-rejected rule still shows up in a verbose trace as a considered candidate.
That reasoning omits the case that turned out to be true: a rule that is in the ruleset but is never reached, because evaluation stopped at an earlier sibling. Such a rule is equally invisible in the trace, and the trace alone cannot distinguish it from a rule that was never loaded. That distinction needed the load-time rule listing (§3.5), which was not gathered until 2026-08-10.
3.4 The positive direction still works. An interactive AIDE event with
auid=1000 fires 100015 at level 12, correctly. Detection was never
weakened by any of this — the estate is noisy, not blind.
3.5 Mechanism — IDENTIFIED. The rules are unreachable, not rejected.¶
Established 2026-08-10 with wazuh-logtest-legacy -d -d, which parses the
ruleset from disk in its own process and prints both the rule count and the
compiled rule tree. No restart, no daemon involvement.
The rules load. All nine of them.¶
8521 — byte-identical to the count the daemon logged on 2026-07-30.
100016, 100017 and 100018 were inside that number the whole time.
printRuleinfo lists them explicitly:
No error, no warning, nothing ignored. The only note against any rule in the
file is deprecated Mitre format, on the five rules carrying <mitre> blocks
— which does not include these three.
Evaluation order is by LEVEL, not by file order¶
Every printRuleinfo block emits the nine rules in the same sequence:
| Position | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
|---|---|---|---|---|---|---|---|---|---|
| Rule | 100012 |
100015 |
100013 |
100014 |
100010 |
100011 |
100016 |
100017 |
100018 |
| Level | 0 | 12 | 12 | 12 | 10 | 10 | 3 | 3 | 3 |
Level 0 first, then strictly descending. File order — 100010, 100011,
100012, 100016, 100017, 100018, 100015, 100013, 100014 — has no
bearing on it.
Therefore¶
100015 is level 12 and matches every credential_access event by design;
it is the catch-all the baseline rules were written to get in front of. It
sorts ahead of them. Evaluation stops at the first match.
The baseline rules are structurally unreachable. They can never be evaluated for any event they were written to catch. Not rejected, not malformed, not missing — present, correct, and behind a door that always closes first.
Two independent traces confirm the order end to end:
credential_access event : 100012 → [80740] → 100015 ✱match, stop
lolbin_download event : 100012 → 100015 → 100013 → 100014 → 100010 ✱match, stop
Six hypotheses eliminated — and what they had in common¶
| # | Hypothesis | How it died |
|---|---|---|
| 1 | type="pcre2" on <field> is rejected |
Attribute removed from all eight conditions, artifact installed and verified 92885b62…, gate re-run. No change. |
| 2 | 100016's grouped alternation cannot match |
Flattened to per-alternative anchored literals. No change. |
| 3 | A shadow copy in a load path overrides the file | Two .bak files were in /var/ossec/etc/rules/, a configured rule_dir; one was the pre-tuning version of this file. Both moved out. No change. |
| 4 | Duplicate rule id | 100016–100018 are defined in exactly one file across both load paths. No duplicates, and no 7612 warning in any archive, ever. |
| 5 | Unresolvable if_sid parent |
All four of 80700, 80780, 80784, 80790 exist in 0365-auditd_rules.xml. |
| 6 | Ordering | Excluded in v1.0 on file-line grounds — and it was the answer. See below. |
Every one of the first five assumed the rules were being REJECTED
That is the single sentence worth carrying out of this investigation. Hypotheses 1–5 differ in mechanism but share one premise: that something was refusing to load these rules. Nothing was. The premise came from §3.3 — absent from the trace — and was never itself questioned, so five tests were spent inside a frame that could not contain the answer.
The load-time rule listing that settled it costs one command and no restart. It was available on day one.
§3.1 was the answer, and it was excluded first¶
v1.0's §3.1 excluded ordering with this reasoning:
100016/100017/100018are at lines 60, 69 and 88 ofgpus-lolbin-rules.xml;100015is at line 98. The baseline rules precede the catch-all, which is the required order.
That used file order as a proxy for evaluation order, and it is not one.
Wazuh sorts siblings by level. The baseline rules do precede 100015 in the
file and are evaluated after it, and nothing in the file's appearance suggests
that.
This is the sharpest correction in the document, for two reasons.
It was first. §3.1 is the opening exclusion. Everything after it was searching a space that had already had the answer removed from it — five hypotheses, two workstation sessions, one install-and-revert cycle on the manager, and a published finding naming the wrong cause, all downstream of one plausible-looking sentence that nobody went back to.
It looked conclusive. Line numbers are checkable, objective, and easy to verify — which is exactly what made the exclusion feel safe enough to build on. A weak-looking exclusion invites revisiting. A strong-looking one does not.
3.6 Two reasoning failures, and the one they share¶
This is the most transferable material in the document, and it is worth more than the fix would have been.
Instance 1 — a real correlation that was not causal¶
The evidence for type="pcre2" was not weak. It was an exact estate-wide split
with no exceptions:
| Rule | <field> syntax |
Anchored | Fires? |
|---|---|---|---|
100034 |
<field name="details.reason">^not_owner$</field> |
yes | ✅ L10, confirmed live |
100012 |
<field name="audit.command">unix_chkpwd\|sshd\|…</field> |
no | ✅ proven suppressing |
100015 |
<field name="audit.key">credential_access</field> |
no | ✅ 184/day |
100016 |
<field name="audit.auid" type="pcre2">^4294967295$</field> |
yes | ❌ 0 all-time |
100017 |
<field name="audit.exe" type="pcre2">^/usr/bin/stat$</field> |
yes | ❌ 0 all-time |
100018 |
<field name="audit.exe" type="pcre2">^/usr/bin/tar$</field> |
yes | ❌ 0 all-time |
Every rule in the estate carrying the attribute was inert. Every rule without
it fired. 100034 even controlled for the obvious alternative — same ^…$
anchors, no type attribute, fires fine — so anchoring was excluded rather
than assumed. The reasoning was disciplined, the confound was stated openly
rather than buried, and the conclusion was still wrong.
What made it wrong is that the split was drawn over three rules that share
far more than the attribute. They are the same three rules, in the same
commit, in the same region of the same file, added on the same day, all at
level 3, all matching the same audit.key. Any one of those properties
produces the identical estate-wide split. With n=3 on one side, "no exceptions"
is a much weaker statement than it reads as.
An estate-wide correlation with no exceptions is evidence about where to test, never a substitute for the test. The strength of a correlation scales with how many independent ways the two groups differ — and here there was only ever one group, described several different ways.
Instance 2 — an exclusion that was wrong, and was never revisited¶
§3.1 excluded ordering using file line numbers as a proxy for evaluation order. The proxy was invalid, the exclusion was wrong, and it was the answer.
What makes this worse than instance 1 is position. The pcre2 conclusion was reached late, published as a conclusion, and therefore attracted the scrutiny that killed it within days. The ordering exclusion was made first, in passing, as throat-clearing before the real analysis — and nothing ever came back to it. A wrong conclusion gets tested. A wrong exclusion just quietly shrinks the search space and is never seen again.
The rule this yields
An exclusion made early and never revisited is more dangerous than a hypothesis held late. A hypothesis is under scrutiny by construction; an exclusion removes its subject from scrutiny by construction. The earlier it sits in the chain, the more work inherits it, and the more confident the surrounding reasoning looks precisely because the space has been narrowed.
Practically: when a line of investigation stalls after several failed hypotheses, re-open the exclusions before generating a sixth hypothesis — and check what each one actually measured, not what it claimed to. §3.1 measured file order. It asserted evaluation order. Nobody read it closely enough to notice, for five days, because it was excluded early and looked objective.
Elimination arguments inherit both failures. "Everything else is excluded, therefore X" is only as sound as the enumeration, and the enumeration is written by whoever already suspects X. v1.0's §3.5 was titled cause, by elimination; the elimination was real and the cause was not.
Instance 3 — a count read as a proxy for a thing it does not measure¶
Added 2026-08-11, during the fix verification.
The verified fix moves Total rules enabled from 8521 to 8506. Fifteen
fewer. The reflex reading — the one this document's own §7 had written into its
acceptance test — is fifteen rules went missing, and it is wrong.
Total rules enabled is a count of tree NODES, not of rules. A rule with
<if_sid>a,b,c,d</if_sid> is instantiated at every position those parents
occupy in the compiled tree. The three baselines listed four parents and were
materialised at six tree positions each — 18 nodes. Re-parenting them under the
single <if_sid>100015</if_sid> collapses that to one position each — 3 nodes.
18 − 3 = 15. The ruleset still contains nine rules in this file, before and
after; printRuleinfo lists all three in both loads.
This is the THIRD time a count has been misread in this investigation
8521→8524(§7). The primary acceptance test was inverted — it would have rejected a working fix, because the count already contained the rules.- The count as evidence (§4).
8521being correct was read as confirmation the rules were live. It confirmed only that they were loaded, which was never in question. - Node count vs rule count (here).
8506read as fifteen rules lost.
The same error three times, in three different directions, over six days —
by people who had already written the previous two down. A number that is
easy to obtain gets used as a proxy for the thing that is hard to obtain.
Total rules enabled is cheap, precise, and answers a question nobody in
this investigation was actually asking.
The correction is not "be careful with counts." It is: before using a metric as acceptance, state what it physically counts. All three instances dissolve on contact with the sentence "this counts nodes in the compiled tree." Nobody wrote that sentence for five days because the number looked self-explanatory.
3.7 The gate worked — this is the finding's own thesis, in the affirmative¶
The failed fix cost nothing: no restart was burned, no production change was made, and no false result was published. The manager ran continuously on the Jul 30 ruleset throughout, and the one-shot load-time capture (§4) is still unspent.
That is not luck. It is because acceptance required a trace, not just a
count. wazuh-logtest-legacy loads the ruleset from disk in its own process,
so a candidate fix can be walked through the real parser without restarting the
manager — and the gate caught the failure before the restart.
The count-based check was worse than useless — it was INVERTED
v1.1 congratulated itself here for keeping the rule count as a non-substitutable check alongside the trace. That was right for the wrong reason, and the reasoning has to be corrected now that the count is known.
The runbook's primary acceptance test was "the rules-enabled count moves
8521 → 8524". 8521 already included all three rules. The count
could not have moved for any fix, correct or incorrect, because nothing was
ever missing from it.
So A1 was not merely weak. It would have FAILED against a correct fix
— a real repair that made the rules reachable would leave the count at
8521 and be rejected as unsuccessful. A test built to catch a false
positive was quietly producing false negatives instead.
The general form is worth keeping: a check derived from a hypothesis
inherits that hypothesis's errors. 8521 → 8524 only makes sense if the
rules are being rejected at load. They never were, so the test measured
nothing about the thing it was gating.
The count tests loading. Only a trace tests matching. That sentence survives v1.1 intact; what changes is that the count was not just insufficient here, it was actively misleading.
One gate defect, found the hard way
The first gate attempt used /var/ossec/bin/wazuh-logtest, which in
Wazuh 4.x is a ~1 KB wrapper around a client that talks to the running
wazuh-analysisd over a socket. It answered from the ruleset loaded on
Jul 30 and reported a stale result that looked like a fix failure.
The standalone loader is the separate 1.2 MB binary,
wazuh-logtest-legacy. Any pre-restart gate in this estate must use it.
Testing an on-disk file with a tool served by the running daemon measures
nothing.
3.8 The tooling lesson¶
One gate defect, found the hard way
The first gate attempt used /var/ossec/bin/wazuh-logtest, which in
Wazuh 4.x is a ~1 KB wrapper around a client that talks to the running
wazuh-analysisd over a socket. It answered from the ruleset loaded on
Jul 30 and reported a stale result that looked like a fix failure.
The standalone loader is the separate 1.2 MB binary,
wazuh-logtest-legacy. Any pre-restart gate in this estate must use it.
Testing an on-disk file with a tool served by the running daemon measures
nothing.
The same binary, at -d -d, is what finally answered the question: it prints
Total rules enabled and the full compiled rule tree with every rule's id and
level. That is the diagnostic that separates never loaded from loaded and
never reached, and it needs no restart and no privileged change — only
sudo to read /var/ossec (§8.5).
It was available from the first day of this investigation. Nothing prevented running it except not knowing it was the question.
Second gate defect — -t exits 0 and prints nothing
Found 2026-08-11. wazuh-logtest-legacy -t (config test) was run against
both the baseline and the verified-fix candidate. Both exited 0. Both
printed nothing at all.
The two rulesets differ in compiled structure and in node count — 8521
versus 8506 — and produce different alerts for the same input event. -t
distinguished them not at all. It did not flag the count change, and it
would not have flagged a regression.
-t answers does this parse. It is routinely read as is this correct.
The general form, and the reason this belongs in the finding rather than in a runbook: the validation tools on this platform verify substantially less than their names suggest, and each one fails in a way that presents as success.
| Tool | Reads as | Actually verifies |
|---|---|---|
wazuh-logtest |
tests my rules | the ruleset the daemon loaded, over a socket — not the file on disk |
analysisd -t / -t |
configuration is valid | the XML parses |
Total rules enabled |
my rules are live | some number of tree nodes exist (§3.6) |
| a clean restart | the change took effect | the process started |
Every one of these was green for the entire period the control did nothing.
A gate that cannot fail is not a gate. The only instrument in this estate
that answered a real question was wazuh-logtest-legacy -d -d plus a -v
trace on a real event — because it is the only one that reports what the engine
did, rather than what it accepted.
3.9 Fix direction — recorded, NOT designed (resolved — see §3.10)¶
The idiomatic Wazuh answer is to stop making the baseline rules siblings of
100015 and make them children of it, via <if_sid>100015</if_sid>, so
they are evaluated after 100015 matches and override its level.
That shape preserves the BT-003 path by construction: an interactive read with
a real auid matches 100015 at level 12 and no child overrides it; a
daemon-context read matches 100015, then the child matches and takes the
alert down to level 3.
This is a direction, not a design, and it must be VERIFIED not asserted
Six hypotheses have been eliminated in this investigation, several of
them stated with more confidence than this one deserves. The child-rule
reordering behaviour must be demonstrated in 4.14.4-rc2 specifically,
with wazuh-logtest-legacy -d -d showing the new tree and a trace showing
the override firing, before anything is installed.
An RC build is running (§8.3) and rule-tree construction is exactly the kind of thing an RC can differ on. Asserting how Wazuh orders child rules, on the strength of how it ought to work, is the identical error to §3.1 asserting evaluation order from file order.
Also unresolved by this direction: whether level 3 or level 0 is wanted (§6.1), and whether the same reordering issue affects any other custom rule in the estate that sits at a lower level than a sibling catch-all. Neither is in scope here.
The working-tree artifact is irrelevant to reachability. Removing
type="pcre2" and flattening 100016's pattern do nothing about which rules
get evaluated. The flatten remains a genuine precondition — 100016 cannot
match under the default matcher until it is corrected (§5) — but it is a
precondition for a rule that must first be made reachable.
RESOLVED 2026-08-11 — verified, and the demand above was the right one
The direction held. It was demonstrated on 4.14.4-rc2, not asserted,
exactly as this section required — see §3.10. The file is now committed at
blob 8ce43cf0 (sha1sum 8588a9ce), superseding the uncommitted
working-tree state this section described.
Worth recording that the demand was not ceremonial: six wrong hypotheses preceded this one, and the verification itself surfaced two defects in its own harness plus one wrong expectation in this document (§3.6, instance 3) before it produced a usable answer.
3.10 Fix VERIFIED — the mechanism demonstrated, both directions traced¶
Verified 2026-08-11 on MAPLE, Wazuh 4.14.4-rc2, against the committed
artifact. No install, no restart, no write to /var/ossec (§3.11).
The compiled tree, inspected via -d -d and not inferred. printRuleinfo
emits one line per tree position, as <depth> : rule:<id>, level <n>:
100015 |
100016 / 100017 / 100018 |
|
|---|---|---|
| Before | depths 2 3 4 4 5 5 |
depths 2 3 4 4 5 5 — identical to the parent |
| After | depths 2 3 4 4 5 5 (unchanged) |
depth 3 — one below 100015's reachable copy |
The "before" row is the mechanism stated exactly: the baselines sat at the
same depths as the rule they were meant to get in front of, because they
were siblings of it. After re-parenting they sit one level below it. Nothing
else in the 8521-node tree moved — 100010–100014 hold six positions in both
loads, 100020–100039 hold one in both.
Q2 — which id and level does a matching child emit? From the -v trace on
a real modulesd event, not from documentation:
*Rule 80700 matched.
*Trying child rules.
*Rule 100015 matched.
*Trying child rules.
*Rule 100016 matched.
**Phase 3: Rule id: '100016' Level: '3'
The child's id and level are emitted, not the parent's. And with no matching child the parent's alert stands unchanged — which is precisely what preserves BT-003:
Both directions, on real events harvested from alerts.json, each vector
run independently so a dropped alternative in 100016's flattened pattern
could not hide behind a passing sibling:
| Vector | Before | After |
|---|---|---|
wazuh-modulesd |
100015 L12 |
100016 L3 |
wazuh-syscheckd |
100015 L12 |
100016 L3 |
aide (cron) |
100015 L12 |
100016 L3 |
stat |
100015 L12 |
100017 L3 |
tar |
100015 L12 |
100018 L3 |
aide, auid=1000 |
100015 L12 |
100015 L12 — unchanged |
The interactive vector is byte-identical to the cron one but for auid, ses
and AUID — verified by diff, not by construction.
The inbox/index split, observed rather than reasoned. -a reports the
output channel per alert. At L12 the alert is tagged : mail; at L3 the alert
is still generated and carries no mail tag. That is the §6.1 threshold
behaviour confirmed directly.
One attachment point is correct, and this was checked rather than assumed
The children occupy one tree position where the parent occupies six, which
reads as under-coverage. It is not. 80700 is <decoded_as>auditd</decoded_as>
— the single root for every auditd event — and 80780/80784/80790 are
its own descendants (depths 2, 3, and 4/4/3), not alternate entry points.
100015's six positions are it hanging off 80700 plus three of 80700's
descendants, and only the depth-2 copy directly under 80700 is ever
reached: it matches there and evaluation descends into its subtree rather
than continuing along depth 2. Across all six vectors, 80780/80784/80790
were never tried at all — not tried and unmatched. The children sit
below the one reachable copy.
A forced-record-type sweep found baseline and candidate agreeing on every
auditd record type except SYSCALL, the intended change. It also surfaced
a property worth keeping: an event carrying key=credential_access whose
decoder does not populate audit.exe fails every child condition and
alerts at 100015 L12. An unattributable credential access stays loud.
3.11 The zero-touch verification method — reusable, and the most useful output here¶
This is the operational takeaway. Anyone changing a Wazuh rule in this estate should start here.
wazuh-logtest-legacy accepts -D <dir>, which relocates the entire ruleset
root. A candidate ruleset can therefore be compiled, inspected and traced
without writing anything to /var/ossec and without restarting the manager:
# build a scratch root — ruleset material only, no client.keys / authd.pass
mkdir -p /tmp/probe/home/{etc,ruleset}
cp -a /var/ossec/ruleset/{rules,decoders} /tmp/probe/home/ruleset/
cp -a /var/ossec/etc/{rules,decoders,lists} /tmp/probe/home/etc/
cp -a /var/ossec/etc/ossec.conf /tmp/probe/home/etc/
cp -a /var/ossec/etc/internal_options.conf /tmp/probe/home/etc/
cp -a /var/ossec/queue/fts /tmp/probe/home/queue/fts
# drop the candidate rule file in, then:
/var/ossec/bin/wazuh-logtest-legacy -D /tmp/probe/home -d # node count + tree
/var/ossec/bin/wazuh-logtest-legacy -D /tmp/probe/home -v < event.log # trace
/var/ossec/bin/wazuh-logtest-legacy -D /tmp/probe/home -a < event.log # alert + mail tag
Build a second root holding the unmodified file and run both. The baseline root is the control: if it does not reproduce the current production behaviour, the harness is wrong and no result from the candidate root means anything.
Three failure modes that cost real time — all of them present as a dead run
- Do not
chmodthe scratch tree.wazuh-testrulechroots and thensetuid()s to thewazuhuser. Tightening permissions as root makesetc/internal_options.confunreadable after the drop, which is a CRITICAL that fires after rule loading — so the node count prints and every event run dies silently.cp -apreserves modes that already work; leave them alone. queue/fts/must exist. FTS init runs after config load; without it every event run aborts withError initiating FTS listwhile-tstill passes.- Dump raw output before grepping it. A format mismatch and a dead process are indistinguishable through a filter. Both of the above were initially invisible for exactly this reason.
Backups of the live rule file go to /var/ossec/backup/rules/ — never into
etc/rules/, which is a configured rule_dir and will load anything left
there.
Cost of the whole method: a few minutes, sudo to read /var/ossec (§8.5),
and no production change. It was available throughout this investigation.
4. Nothing was dropped — which makes this HARDER to detect, not easier¶
The manager start sequence for 2026-07-30, from the archived log, records:
That is the whole of it. No warning naming the file. No warning naming the rule ids. No parse error. No count discrepancy.
v1.0 understated this. The real version is worse.
v1.0 read the silence as "a ruleset that silently dropped three rules starts identically to one that loaded them" — and offered the rule count as the signal nobody was watching.
Nothing was dropped. The rules loaded. The count is correct — 8521
includes them. wazuh-analysisd -t passes because the ruleset is valid.
Every artifact is not merely unrevealing but genuinely, accurately
green, because there is no fault at load time to report.
That is a strictly harder failure mode than the one v1.0 described:
| A rule silently dropped | A rule loaded and unreachable | |
|---|---|---|
| Parse error | none | none |
| Startup warning | none | none |
| Rule count | discrepancy — a real signal, if watched | correct. No signal exists at all |
analysisd -t |
passes | passes |
| Detectable by | counting rules | only by tracing an event |
A dropped rule leaves a number that is wrong. Somebody could watch that number. An unreachable rule leaves nothing wrong anywhere — every observable is correct, and the only way to see the defect is to walk a representative event through the ruleset and notice which rule fires.
So the v1.0 remediation suggestion — watch the rule count — would not have
caught this. It would have shown 8521 on Jul 30, 8521 today, no change,
no alarm, for a control that has never once run.
This is the part that generalises past Wazuh. Any gate that stops at "did it load" passes a rule that can never execute. Loading is not reachability, and reachability is not matching; three different properties, and only the third is what anybody actually wanted.
Corroborated by a second, independent loader (2026-08-10)
wazuh-logtest-legacy parses the ruleset from disk at invocation, entirely
outside the daemon, and at default verbosity says nothing beyond its
startup line and a deprecation notice. Two independent code paths, no
complaint from either — correctly, since the ruleset is valid.
At -d -d the same binary prints the count and the full rule tree. The
information exists; it is simply not surfaced at any normal verbosity, by
either loader, and there is no fault for it to surface.
Note also that this section's evidence itself nearly went unrecoverable.
The live ossec.log was 3916 bytes and no longer contained the Jul 30
load line at all; the 8521 baseline survives only in the rotated archive
(§9). A finding about undetectable failure was one log rotation away from
losing its own baseline.
5. Second defect on the same rules — real, latent, and NOT the blocker¶
Status, as of v1.2
v1.0 recorded this as a gap in the staged fix. It was corrected in the
working tree, and flattening it changed nothing (§3.5, hypothesis 2) —
because 100016 is never reached in the first place.
The defect is still real and still has to be carried. It is a
precondition, not a cause: once the reachability fix (§3.9) puts
100016 in a position to be evaluated, it will match nothing until this
pattern is corrected. Two independent defects on the same rule, and the
reachability one masks the matching one completely.
While staging the fix, a second, independent failure was found in 100016.
Its exe condition is:
This is PCRE2 grouping syntax, and the default Wazuh matcher cannot
evaluate it. Under the default matcher the parentheses are literal characters
and | splits at the top level, so the condition degrades to a set of
alternatives that no real value can equal. Removing the type="pcre2"
attribute alone would leave 100016 inert by a different mechanism — the
rule would load, be evaluated, and still never match.
100017 and 100018 are not affected: their patterns are single anchored
literals (^/usr/bin/stat$, ^/usr/bin/tar$, ^/var/ossec$, ^4294967295$)
which the default matcher handles correctly, as 100034 demonstrates.
5.1 Corrected, and now proven — committed 2026-08-11¶
The pattern has been flattened to per-alternative anchored literals. Same three exact paths, no widening:
- <field name="audit.exe" type="pcre2">^(/var/ossec/bin/wazuh-(modulesd|syscheckd)|/usr/sbin/aide)$</field>
+ <field name="audit.exe">^/var/ossec/bin/wazuh-modulesd$|^/var/ossec/bin/wazuh-syscheckd$|^/usr/sbin/aide$</field>
Committed 2026-08-11 — the flatten is now proven
The withholding was right on principle and is now discharged on evidence.
The flatten has been observed working — all three of 100016's
alternatives matched, each on an independent run (§3.10) — so the file is
committed, at blob 8ce43cf0 / sha1sum 8588a9ce. It could not be
validated before, exactly as this section said, because the reachability
defect made it unobservable.
Digest-algorithm mismatch — a real cost, and the document was not the one at fault
The digests in this document are sha256. The 2026-08-10 and 2026-08-11 verification sessions worked in sha1 and git blob hashes. All three are hex strings, and truncated to eight characters they are indistinguishable.
9262b352… was twice judged to "match nothing live" during the
verification work and treated as a stale value carried forward from a
session summary. It is correct. It is the sha256 of the deployed file,
exactly as §3.1 states — the same object git knows as blob 20c2badb and
sha1sum reports as 48c23898:
| Artifact | git blob | sha1sum |
sha256 |
|---|---|---|---|
| installed on MAPLE (unchanged throughout) | 20c2badb |
48c23898 |
9262b352 |
| committed fix (v2.1) | 8ce43cf0 |
8588a9ce |
75369079 |
The error was reproduced twice in one session, the second time while explicitly correcting the first — the same pattern as A1 → A1′ (§7) and as §3.6 instance 3. A comparison that fails because the two sides are not the same kind of thing looks exactly like a comparison that fails because one side is wrong, and the second reading is the tempting one when a prior artifact is already under suspicion.
Practically: state the algorithm next to every digest in this estate, and when a digest does not match, check the algorithm before concluding the digest is stale.
92885b62… in §3.5's hypothesis table is likewise a sha256, of a transient
test artifact installed on MAPLE during hypothesis-1 testing. It was never
a git object, so comparing it against the working tree was never a
meaningful test of it either way.
The near-miss worth recording
Had v1.0's staged fix shipped without this correction, it would have left
100016 — the largest share of the volume — matching nothing even once
reachable. That is the finding's own failure mode reproduced inside its
remediation, and only a trace-based acceptance check catches it (§3.7).
One silent failure was concealing another. 100016 cannot be observed
failing on its pattern, because it is never evaluated. That is why flattening
it produced no observable change, and why the flatten cannot be validated until
the reachability fix lands. Two defects stacked, and the outer one makes the
inner one unobservable — which is the same shape as everything else in this
document.
6. Impact — this is the dominant alert-fatigue contributor¶
Over 14 days to 2026-08-05, of everything at or above the alerting threshold:
| Count | Share | |
|---|---|---|
100015 |
2,576 | 95.1% of all level ≥ 10 alerts |
100010 (curl/wget, L10) |
98 | 3.6% |
40112 |
26 | 1.0% |
| everything else combined | 8 | 0.3% |
Restricted to level 12 alone, 100015 is 99.0% (2,576 of 2,603).
Finding #5 (alert fatigue) is, to a first approximation, this one rule. The remediation for it was written, reviewed, committed and deployed — and because the rules never entered the ruleset, none of that reduced the alert volume by a single event. The documentation records the problem as addressed.
6.1 Alert thresholds — confirmed, and the level-3/level-0 decision CLOSED¶
email_alert_level = 10 and log_alert_level = 3 are confirmed on the
manager — re-read directly from ossec.conf on 2026-08-10, so this is now
independently verified rather than carried evidence (§9).
This is what makes level 3 the right target for the baseline rules: at level 3
the suppressed events fall below email_alert_level, so mail stops. But
they remain at or above log_alert_level, so they are still written and
still indexed. Level 3 suppresses the notification, not the volume.
DECIDED 2026-08-11 — level 3. The fix targets INBOX volume, not indexed volume.
Level 3 is retained. log_alert_level=3 keeps the events written and
indexed; email_alert_level=10 means level 3 stops the mail. The mail
decision was confirmed observationally, not inferred — -a tags the L12
alert : mail and the L3 alert with no channel while still generating it
(§3.10).
Consequence, stated plainly so nobody reads the fix as incomplete: indexed volume stays at roughly 184 documents/day. That is by design. What drops is email. Anyone measuring success by index volume will conclude the fix failed; they will be measuring the wrong thing.
The argument that settled it is the one below: this finding exists because a control went silent and nothing noticed. A level-3 rule firing at its expected rate is positive evidence the exclusion is alive. Level 0 would make a dead agent, a dropped auditd watch and a working baseline indistinguishable — the same bet that produced VLN-010.
The decision as it stood — level 3 or level 0?
Level 3 keeps the events in the indexer. The baseline stays auditable:
you can still answer "is modulesd reading /etc/shadow at the rate it
should be?", and a sudden change in that rate is itself a signal. Cost is
~184 indexed documents per day that nobody reads.
Level 0 drops them entirely. Storage and indexer noise go to zero, and so does the ability to detect that the baseline itself has shifted.
Not decided here. The argument for 3 is that this finding exists precisely because a control went silent and nothing noticed — discarding the baseline signal is the same bet that produced VLN-010. The argument for 0 is that an unread index is not monitoring either.
7. Acceptance test for the fix¶
The fix was verified (§3.10) and is now deployed (§7.2). These conditions were fixed in advance of deployment, deliberately, so that nobody improvises acceptance at restart time under pressure to declare success. They were written before the deploy and are reproduced here unchanged — the outcome against them is recorded separately in §7.2, so the criteria cannot be read as having been fitted to the result.
Conditions 1–6 below were the pre-fix criteria. A1′ has since been falsified
by the verified fix and is replaced by A1″. The rest held.
A1 was INVERTED and is withdrawn
v1.0 and v1.1 made this the primary test:
The rule count moves
8521→8524. … Exactly three additional rules, no more and no fewer.
8521 already contained all three rules. The count could not move for
any fix, and a correct fix will leave it at 8521.
So the primary acceptance test would have rejected a working repair — not a weak test, an inverted one. It was derived from the assumption that the rules were being rejected at load, and it inherited that assumption's error wholesale (§3.7).
As written in v1.2, now also superseded: "Replaced by A1′ below. The
count is retained only as a regression check: it must stay at 8521.
A count that moves means the fix changed what loads, which is not what any
correct fix here does."
That replacement was itself wrong — see A1′ below, and A1″, which
supersedes both. The correct value is 8506. The final sentence above is
the load-bearing error: a reachability fix does change what loads, because
the count measures tree nodes and re-parenting moves them (§3.6).
-
A1′ — the rule count STAYS at
8521. — WITHDRAWN, wrong in a new direction. Superseded byA1″.A1′ was ALSO inverted. Same test, third failure.
A1 said the count must move
8521→8524. That was inverted, and A1′ replaced it with the count must stay at8521— which is also wrong. The verified fix moves it to8506, and A1′ would have rejected it.A1 assumed the fix adds rules. A1′ assumed a reachability fix cannot change what loads. Both assumed
Total rules enabledcounts rules. It counts tree nodes (§3.6, instance 3), and re-parenting changes node count by construction.Note the shape: A1′ was written in the act of correcting A1, by someone who had just been burned by this exact metric, and it inherited the same unexamined premise. Correcting a conclusion does not correct the assumption underneath it.
-
A1″ — the node count reads
8506, NOT8521. Confirmed withwazuh-logtest-legacy -d.8521at restart means the fix did not take effect. Any other value means something unanticipated changed and the deployment should stop. Derivation: three rules × five fewer tree positions each = 15; 8521 − 15 = 8506. - Each rule id appears in a
wazuh-logtest-legacy -vtrace, and fires. Being tried proves reachability; firing proves matching. These are separate properties and both are required. Use the standalone binary —wazuh-logtestanswers from the running daemon and will report the old ruleset (§3.8). - Interactive LOLBin detection still fires at level 12. Re-run the BT-003
procedure: a
/etc/shadowread with a realauidmust still produce100015at level 12. Test vectoraide-interactive-MUSTFIRE.logisolates exactly this — it is byte-identical toaide-cron.logexcept for theauid. A fix that suppresses the noise and the detection together is a regression, not a fix. 100015volume measured over 24 hours, not instantly. The generators run on cron and SCA cycles, not continuously. A drop observed minutes after a restart is a restart artifact. Target ~184/day → ~0–1/day, judged on a full day.100016specifically must be observed firing, not merely loaded — it carries the largest share and it is the rule with the second defect (§5).- The compiled rule tree is inspected, not inferred.
wazuh-logtest-legacy -d -dmust show the baseline rules positioned so they are reached — the ordering must be observed in the tree, not deduced from the file. That deduction is the §3.1 error and this test exists to prevent its recurrence.
What is NOT evidence
File presence is not evidence. A clean restart is not evidence. A passing
analysisd -t is not evidence. The rule count is not evidence. All four
were true, correct and green for the entire period this control did
nothing — the count included the rules the whole time.
Nor is an alert-volume drop sufficient on its own: 100015 going quiet
with 100016 also at zero means something upstream stopped delivering
events, not that suppression worked.
The generalised lesson
Config validation proves a ruleset parses. It does not prove a rule matches,
and it does not prove a rule was even retained. Every gate in this estate
that stops at analysisd -t will pass a rule that can never fire. A
detection is not deployed until it has been observed firing on a real event
— the same both-directions standard already applied to the sudo scope-down
work.
7.1 Deploy acceptance — fixed in advance, 2026-08-11¶
Stated before the deployment session so nobody improvises it afterwards. All six must hold. Any failure stops the deployment; none of them is optional and none is satisfied by an adjacent observation.
| # | Criterion | Passes if | Fails if |
|---|---|---|---|
| D1 | Node count | wazuh-logtest-legacy -d reads 8506 |
reads 8521 — the fix did not take effect. Any other value — stop |
| D2 | Compiled tree inspected, not inferred | -d -d shows 100016/17/18 at depth 3, one position each, below 100015's depth-2 copy |
ordering deduced from file position — that is the §3.1 error |
| D3 | Daemon vectors suppressed | all five land at level 3, under the CHILD's id (100016/100017/100018) |
any lands at 100015 L12, or at L3 under 100015 |
| D4 | BT-003 intact | aide with auid=1000 still emits 100015 level 12 |
anything lower. A fix that suppresses the noise and the detection together is a regression |
| D5 | 100016 observed firing |
seen matching a real event, not merely loaded | loaded but never matched — it carries the largest share and the §5 second defect |
| D6 | 24h volume, judged the following day | 100015 email volume drops toward ~0/day |
judged minutes after restart — that is a restart artifact, not a result |
D6 measures the INBOX. It does not measure the index.
Indexed volume stays at ~184/day and that is the intended outcome, not a
partial one. log_alert_level=3 keeps every suppressed event written and
indexed; email_alert_level=10 is what level 3 clears. The fix removes
mail, not documents (§6.1).
Anyone checking success by counting indexed 100015/100016 documents
will see no improvement and conclude the deployment failed. The metric is
email volume. Indexed volume holding steady is a pass, and 100016
appearing in the index at roughly the rate 100015 used to is positive
confirmation the exclusion is alive rather than the pipeline being dead.
Corollary from §7's "what is NOT evidence": 100015 going quiet with
100016 also at zero is a failure, not a success — it means events
stopped arriving.
Rollback
The pre-change file is preserved at /var/ossec/backup/rules/ — never in
etc/rules/, which is a configured rule_dir. Rollback is: restore the
backup, restart, confirm the count returns to 8521. Note that the restart
itself is the expensive part; the load-time capture (§4) is one-shot per
restart and should be captured on the way in.
7.2 Deploy outcome — 2026-08-11¶
Deployed 2026-08-11 18:29:50 UTC. First change to the manager since
2026-07-30. wazuh-analysisd PID 1491909 → 2361899.
The installed artifact, verified on all four measures before the restart — stating each algorithm, per the §5.1 hazard:
| Measure | Value |
|---|---|
sha256 |
753690795a4018dbaa74fb2bf9a7d71150c529654a26e948ecb57098d53f2053 |
sha1sum |
8588a9ce52d76fcc47d930e0c8ae4d87b57492b9 |
| git-blob | 8ce43cf08349d6176ef554a88e12ef5ce2aa2a46 |
| bytes | 8367 |
Pre-change backup: /var/ossec/backup/rules/gpus-lolbin-rules.xml.bak-2026-08-11,
verified against PRE sha256 9262b352… before the install was allowed to
proceed. It is in backup/rules/, not etc/rules/ — the latter is a
configured rule_dir and would load anything left in it.
D1–D5 — passed, with the source of each stated¶
| # | Result | Source |
|---|---|---|
| D1 | PASS — node count reads 8506 |
wazuh-logtest-legacy -d, on disk |
| D2 | PASS — 100016/17/18 at depth 3, one tree position each; 100015 keeps its six, reachable copy at depth 2 |
wazuh-logtest-legacy -d -d, tree inspected, not inferred |
| D3 | PASS — all five daemon vectors at level 3 under the child's id | the running daemon |
| D4 | PASS — auid=1000 still 100015 level 12, MITRE T1003 mapping intact |
the running daemon |
| D5 | PASS — 100016 observed matching, firedtimes: '1' |
the running daemon |
D3/D4/D5 were measured against the running daemon, not the file on disk,
and that distinction is the whole point. wazuh-logtest-legacy reads the file
and would have reported the new behaviour whether or not the restart took
effect. wazuh-logtest — the socket client that was the wrong tool
pre-restart (§3.8) — is the only instrument that answers from what
wazuh-analysisd actually loaded. Post-restart it becomes the right one. The
same tool, opposite verdicts on its usefulness, depending on which side of the
restart you are standing.
Verdicts as the daemon returned them:
| Vector | id | level | mail |
firedtimes |
|---|---|---|---|---|
wazuh-modulesd |
100016 |
3 | False |
1 |
wazuh-syscheckd |
100016 |
3 | False |
1 |
aide (cron) |
100016 |
3 | False |
1 |
stat |
100017 |
3 | False |
1 |
tar |
100018 |
3 | False |
1 |
aide, auid=1000 |
100015 |
12 | True |
1 |
firedtimes: '1' on every row means each rule matched, not merely loaded.
That is the distinction §7 condition 2 insisted on — being tried proves
reachability, firing proves matching — and it is what zero hits all-time
never demonstrated for the entire life of this finding.
The mail flag is the direct measurement of the Finding #5 problem
mail: False at level 3. mail: True at level 12. Reported by the daemon
itself, not inferred from email_alert_level and a level number.
This is the claim the whole fix was built to satisfy — the daemon-generated baseline stops reaching the inbox while the interactive attack path keeps its email. Every previous version of this document reasoned about it from configuration. It is now observed.
Load-time output¶
58 lines from the restart boundary, captured against the byte offset recorded before the restart. Zero errors. Zero rule-related warnings. The only 15 warnings are the known indexer-connector retry flood (§8.6), unrelated to rules.
The one-shot load-time capture is now SPENT — and §3.11 was right
§4 and the runbook both treated the next restart as a scarce, one-time
chance to capture load-time output. It has now been spent, and it did not
contain the rule count: wazuh-analysisd does not log Total rules
enabled at default verbosity.
The count came from wazuh-logtest-legacy -d, as §3.11 said it would. The
resource that was guarded for a week was never the one that held the
answer.
D6 — MEASURED 2026-09-08. All three parts PASS.¶
Not judged on the day of deployment. 47 minutes of post-restart data is precisely the restart artifact §7.1 warns against; a full day removes the argument. Three parts, all required:
| Criterion | Passes if | |
|---|---|---|
| (a) | 100015 EMAIL volume |
drops toward ~0/day |
| (b) | INDEXED volume | stays ~184/day — this is SUCCESS |
| (c) | 100016/100017/100018 |
positively non-zero at level 3 |
(b) is where this will be misread, and (a)+(c) can fail together while looking like success
The fix removes mail, not documents. log_alert_level=3 keeps every
suppressed event written and indexed. A flat indexed count is the intended
outcome. Anyone measuring success by counting indexed documents will see no
improvement and call the deployment a failure.
The compound failure: 100015 going quiet and 100016 at zero is
a failure, not a win. It means events stopped arriving — a dead agent,
a dropped auditd watch, a broken pipeline — not that suppression worked.
Silence is the failure mode this finding is about. (c) exists to
distinguish the two, and it is the only part that can.
Result — measured 2026-09-08, 26 days after the due date¶
Measured read-only against the CEDAR indexer from MAPLE: whole-day document
counts from wazuh-alerts-YYYY.MM.DD. No write to /var/ossec, no restart, no
sudo. A dash means the rule did not yet exist in a reachable position that day.
| Day | 100015 |
100016 |
100017 |
100018 |
total |
|---|---|---|---|---|---|
| 2026-08-09 (pre-fix) | 183 | 0 | 0 | 0 | 183 |
| 2026-08-10 (pre-fix) | 184 | 0 | 0 | 0 | 184 |
| 2026-08-11 (deploy, 18:29 UTC) | 149 | 34 | — | — | 183 |
| 2026-08-12 | 0 | 172 | — | — | — |
| 2026-08-13 (original due date) | 0 | 170 | 8 | 4 | 182 |
| 2026-09-05 | 0 | 170 | 8 | 4 | 182 |
| 2026-09-06 | 0 | 172 | 8 | 4 | 184 |
| 2026-09-07 | 0 | 171 | 8 | 4 | 183 |
Deploy day splits cleanly across the restart — 149 + 34 = 183 — which is itself a check: the two halves of that day sum to the baseline, so nothing was dropped in the transition.
(b) PASS. Total indexed volume is 182–184/day after the fix against
183–184/day before it. Flat, and flat is the intended outcome:
log_alert_level=3 keeps every suppressed event written and indexed. The fix
removed mail, not documents, exactly as §6.1 said it would.
(c) PASS. 100016/100017/100018 are non-zero and exclusively level
3. A rule.level aggregation over 2026-09-07 returns a single bucket for
each — {3: 171}, {3: 8}, {3: 4} — with sum_other_doc_count: 0. Not one
document at any other level.
The compound failure is excluded by measurement, not by assumption.
100015 reading zero would be worthless on its own; §7.1 is explicit that
silence is the failure mode this finding is about. Events are still arriving,
and that was watched rather than sampled once: the 2026-09-08 index held 2,397
documents at 13:52 UTC and 2,433 at 13:58:47 UTC — plus 36 in seven minutes,
observed live. The three children carry the whole ~184/day the catch-all used to
carry. The traffic moved. It did not stop.
Unasked-for result — BT-003 is now proven in production, not only in a
trace. D4 showed at deploy time that a human-triggered aide still reached
100015 at level 12. Since 2026-08-12 100015 has fired 15 times, every one
at level 12, none lower. So the detection half is alive on real events and not
merely on a synthetic one — a stronger statement than the acceptance table asked
for. The fix took the daemon noise out without touching the thing the rule
exists to catch.
(a) PASS — measured at the inbox on 2026-09-08. Mail volume dropped ~95% across the fix boundary.
The needle was wrong twice before it was right
The first two attempts counted ossec and then wazuh in
/var/log/maillog. Both returned 0, and both zeros were worthless:
postfix logs addresses, not the sending application. The envelope is
email_from alerts@greenpeace.us → email_to
rajesh.chhetry@greenpeace.us, so neither string could ever appear whether
or not mail was flowing.
A third hypothesis — that wazuh-maild speaks SMTP directly to a remote
smtp_server and bypasses local postfix entirely, making maillog the
wrong file — was eliminated by reading the live ossec.conf:
smtp_server is localhost. /var/log/maillog is the right file.
Three ways to get a meaningless zero, on the correct file, before a single correct measurement. A zero from an unvalidated needle is not evidence of absence.
Counted on the actual recipient, to=<rajesh.chhetry@greenpeace.us>:
| Log file | Count | |
|---|---|---|
maillog-20260816 |
351 | spans the 2026-08-12 fix |
maillog-20260823 |
19 | |
maillog-20260830 |
18 | |
maillog-20260906 |
25 | |
maillog (current) |
7 |
Corroborated independently by file size — 334 KB against 58 KB — and the
recipient-split counts are near-identical to the raw address counts, so forms
traffic sharing this relay is not materially inflating the post-fix figures.
That check was necessary: the forms portal delivered to
swoodley@greenpeace.org through this same postfix on 2026-09-08 at 13:59:40.
Report this as a RATE, never as an alert count
email_maxperhour is 12, so Wazuh was capped at 288 messages/day.
Pre-fix volume was ~184/day at level 12 — which would have hit that cap on
most days.
351 in a week is therefore a THROTTLED SAMPLE, not a count of alerts. The honest statement is "mail dropped ~95%". The statement "351 alerts became 19" is wrong and must not be used: the pre-fix figure is censored by the rate limit, so the true suppressed volume is higher than 351 and the real reduction is larger than the ratio suggests. The percentage survives the censoring; the absolute numbers do not.
With (a) measured, the rollback backup at /var/ossec/backup/rules/ and the
/tmp/vln010* evidence on MAPLE may be retired at the owner's discretion. The
measurement scripts are retained as the reproduction method (§7.3).
7.3 Reproduction method — retained on MAPLE¶
D6 (b) and (c) were measured by three scripts, kept on MAPLE so the numbers in §7.2 can be re-derived rather than trusted:
| Script | What it measures |
|---|---|
/tmp/d6counts.sh |
Whole-day counts of 100015–100018 per wazuh-alerts-YYYY.MM.DD, pre and post fix |
/tmp/d6lvl2.sh |
rule.level aggregation per rule — the check that (c) is exclusively level 3 |
/tmp/d6bt3.sh |
100015 hits across all indices since the deploy, with a level breakdown — the BT-003 check |
All three are reads against the CEDAR indexer, issued from MAPLE
(172.16.0.12), which is one of the two sources permitted to reach
172.16.0.13:9200. None writes to /var/ossec, none requires root, and none
restarts a daemon.
D6(a) was measured separately, as root, on /var/log/maillog — see the D6
result in §7.2.
The measurement path is itself a finding, and it is NOT this one
These scripts reach the indexer with no credentials, because the cluster has no authentication. That is recorded separately as VLN-053 and is not a defect in the Wazuh ruleset. It is noted here only so a future reader does not mistake the ease of reproduction for a property of Wazuh.
8. Side findings — each independent of the above¶
8.1 Malware detection 100026–100029 is authored but not deployed¶
✏️ Corrected 2026-08-10. The v0.1 draft of this finding stated that these
four rules "exist only as a table in EXTRACTED.md" and were "never authored",
and that earlier notes calling them in git but not deployed were wrong. That
correction was itself wrong, and is withdrawn.
Verified in the repo on 2026-08-10: 100026–100029 exist as complete XML
rules in forms-backend/wazuh-rules/gpus-forms-portal-rules.xml, lines 72–94,
each with a <match>, <description> and <group>. 100026 is a level-12
INFECTED-attachment rule with the group malware,virus_detected,pci_dss_5.1,.
EXTRACTED.md's original characterisation is the accurate one: the repo is
ahead of production. Production stops at 100025; these four are tracked but
not deployed. Malware detection for the forms pipeline is an asserted
capability with a written rule behind it that has never been loaded onto the
manager — which is the mirror image of VLN-010, not the same failure. Its own
build task.
8.2 Both forms rule files ARE tracked in git¶
✏️ Corrected 2026-08-10. The v0.1 draft stated that
gpus-forms-portal-rules.xml and forms_authz_rules.xml are "live on MAPLE and
untracked in git", and that losing MAPLE would lose those detections.
Both are tracked, verified with git ls-files:
| File | Tracked since |
|---|---|
forms-backend/wazuh-rules/gpus-forms-portal-rules.xml |
802b42c, 2026-04-20 |
soc/forms-authz-detection/forms_authz_rules.xml |
a8403a6, 2026-07-28 |
There is no single-copy-production exposure for these two files. The real
drift, per §8.1, runs the other way: the tracked copies are ahead of the
manager. Worth noting that the draft asserted the opposite direction of drift
for the same file it named in §8.1 — two mutually inconsistent claims about
gpus-forms-portal-rules.xml survived in one document, which is what an
unverified side finding looks like.
8.3 A release candidate is running as the production SIEM manager¶
✅ CONFIRMED 2026-08-10 — status changed from UNRESOLVED in v1.0.
Two independent live sources on MAPLE:
wazuh-control info → WAZUH_VERSION="v4.14.4"
WAZUH_REVISION="rc2"
WAZUH_TYPE="server"
VERSION.json → {"version":"4.14.4","stage":"rc2","commit":"5933ec9"}
EXTRACTED.md was right. The estate's SIEM manager is running a release
candidate, which puts the platform outside its vendor support boundary — an
RC is not a supported production build, and this is the component the whole
detection estate depends on. That is a standing operational exposure
independent of VLN-010 and needs an owner decision on moving to GA.
v1.0's inference was wrong — correcting it
v1.0 argued from rpm reporting VERSION=4.14.4 RELEASE=1 that a GA build
was installed, and downgraded the rc2 claim to UNRESOLVED on that basis.
RELEASE=1 is the RPM packaging release number, not the upstream build
stage. They are different fields answering different questions: RELEASE
counts repackagings of a given version, and an RC can be packaged as
-1 exactly as a GA can. It was never evidence about build stage.
The finding also withdrew "RC-build regex behaviour differs from GA" as
support for §3.5 on the strength of that inference. That withdrawal is
itself withdrawn. With rc2 confirmed, version-specific parser behaviour
is a live candidate mechanism and is carried in §3.5's Still open list.
8.4 rule.groups renders attackgpus and corrupts indexer filtering¶
Every 100015 alert carries rule.groups containing the token
attackgpus. The enclosing element is <group name="gpus,lolbin,attack">
— no trailing comma — so Wazuh concatenates it directly onto the rule's own
group list. Confirmed in the committed file, line 1.
Cosmetic in isolation, but it corrupts group-based filtering and searching in
the indexer: a search for the attack group misses these alerts and a search
for gpus misses them too. The stock ruleset and both forms rulesets use the
trailing-comma convention. It is in the file being changed, so it is nearly
free to fix in the same pass — but it is not in the staged diff.
8.5 cloudadmin is not in the wazuh group¶
Verified live 2026-08-10:
uid=1001(cloudadmin) gid=1006(cloudadmin) groups=1006(cloudadmin)
drwxr-x---. 20 root wazuh 4096 /var/ossec
cloudadmin holds neither root nor wazuh, and /var/ossec is 0750, so
every manager-side check in this diagnosis — reading a rule file, running
wazuh-logtest, reading ossec.conf — requires an interactive sudo and a
human at the keyboard. That is why §9 exists.
Read-only group membership would let routine drift checks run unattended without granting any write path. Belongs on the VLN-009 least-privilege follow-on list (§7.6), alongside the per-service key split: same argument, the account needs a narrow read capability and the only thing on offer today is full root.
8.6 The indexer-connector floods ossec.log on restart¶
The indexer-connector emits a high volume of log lines into ossec.log on
every manager restart. Operationally this buries anything else written during
the start window — including, in principle, exactly the kind of ruleset
diagnostic §4 shows does not exist. Carried from the Step 3 session; not
independently reverified (§9).
8.7 wazuh-logtest returned empty output after repeated invocations — UNTESTED¶
Observed 2026-08-11 during deploy acceptance. Recorded, not investigated.
During the acceptance re-parse, wazuh-logtest printed complete Phase 3
verdicts for all six vectors in one pass, and then — same binary, same vectors,
same daemon, minutes later — returned empty for every field, including
fields it had just printed correctly. Roughly twenty invocations had been made
against the daemon by that point.
Hypothesis: session exhaustion. wazuh-logtest opens a session against
wazuh-analysisd per invocation, and the daemon caps concurrent sessions.
This is a hypothesis and it has not been tested. It is recorded here in the
form it was observed, because this investigation's history is confident causal
claims that turned out to be wrong (§3.6), and because naming a cause before
testing it is exactly the §3.1 error.
It did not affect the deploy verdict — the raw first-pass output is the evidence, and D1/D2 came from a different binary entirely.
Why this is worth a section rather than a footnote
This is an availability property of an incident tool, not a curiosity.
wazuh-logtest is what an analyst reaches for mid-incident to ask "would
this event have alerted?" — under time pressure, on an unfamiliar log
line, often after already trying several variations.
If the tool goes silent after a handful of invocations, the failure lands at precisely the moment of maximum need, and it fails silently: empty output, not an error. An analyst who has run it fifteen times debugging a live alert would read the sixteenth empty result as "this event matches nothing" — which is the same class of wrong answer as everything else in this document, in a tool whose whole purpose is answering that question.
To test: invoke it repeatedly against a known-matching event, count the invocations to first empty result, and check whether a delay or a manager restart restores it. Do not do this during an incident.
9. Provenance — what is verified from where¶
This finding mixes evidence gathered from two different vantage points, and the distinction is load-bearing.
Verified from the workstation (repo + unprivileged SSH to MAPLE):
- the deployed-file SHA matches
HEAD(§3.1) wazuh-analysisdstart time, PID and 10-day uptime (§3.2)- rule ordering, line numbers, identical
if_sid, missing trailing comma (§3.5, §8.4) 100026–100029exist as XML; both forms rule files are tracked (§8.1, §8.2)cloudadmingroup membership and/var/ossecmode (§8.5)
Verified live on the manager on 2026-08-10, with Rajesh at an interactive
sudo session — these move OUT of carried evidence:
email_alert_level = 10andlog_alert_level = 3, re-read directly fromossec.conf(§6.1). Now independently verified, not carried.WAZUH_REVISION="rc2"fromwazuh-control info, corroborated byVERSION.json(§8.3). Now CONFIRMED.Total rules enabled: '8521', recovered from the rotated archive at/var/ossec/logs/wazuh/2026/Jul/ossec-30.log.gz— note the archive path islogs/wazuh/, notlogs/ossec/. The liveossec.logwas 3916 bytes and no longer contained the line, so the baseline exists only in the archive.- all six eliminated hypotheses and their traces (§3.5)
- the mechanism itself:
Total rules enabled: '8521'fromwazuh-logtest-legacy -d, matching the Jul 30 archive exactly; the full compiled rule tree from-d -dlisting all nine rules with their levels; and the level-descending evaluation order confirmed across two traces with differentaudit.keyvalues (§3.5) - no duplicate ids, and all four
if_sidparents present in0365-auditd_rules.xml - that the standalone loader emits no diagnostic at default verbosity, and that there is no fault for it to emit (§4)
Still carried from the 2026-08-06 Step 3 session and not reverifiable
without sudo, because everything under /var/ossec is unreadable to
cloudadmin (§8.5):
- the indexer-connector restart flood (§8.6)
Indexer counts (§2, §6) were measured 2026-08-05 and are not re-measured here; the 14-day window has since moved.
Evidence base for this revision, left in place on MAPLE:
| Artifact | Contents |
|---|---|
/tmp/legacy-debug.txt |
-d load. Contains the Total rules enabled: '8521' line at 10705 |
/tmp/legacy-debug2.txt |
-d -d load. Contains the printRuleinfo tree — every rule id with its level, in evaluation order |
/tmp/legacy-modulesd.txt |
the credential_access trace |
/tmp/lolbin-vectors/curl-100010.log |
a real 100010 event pulled from the indexer, used for the second-order confirmation |
Verified live on the manager on 2026-08-11, Rajesh at an interactive sudo
session, using the zero-touch method (§3.11) — no install, no restart, no write
to /var/ossec:
- the compiled tree before and after, from
-d -don two scratch roots: the baselines at identical depths to100015before, at depth 3 below it after (§3.10) - node count
8521→8506, and its derivation as a node count (§3.6) - all five daemon vectors resolving to the child id at level 3, each run
independently; the
auid=1000vector holding at100015L12 - the child-supersedes-parent trace, both directions (§3.10)
- the mail-channel split from
-a:: mailat L12, absent at L3 80780/80784/80790never tried on any vector, closing the single-attachment-point question (§3.10)-texiting 0 and printing nothing on both rulesets (§3.8)- re-verified against the committed artifact (
sha1sum 8588a9ce) rather than the tested candidate, since the two differ by comments: 29 assertions, 0 failures. The comment-only claim was checked, not assumed.
Throughout, wazuh-analysisd remained PID 1491909, started Jul 30 20:08:21,
and /var/ossec/etc/rules/gpus-lolbin-rules.xml remained sha1sum 48c23898.
Both were re-asserted at the end of every block.
Evidence base for the 2026-08-11 verification, left in place on MAPLE until after deployment:
| Artifact | Contents |
|---|---|
/tmp/vln010-raw/load-{baseline,candidate}.txt |
-d loads of both roots — node counts and the full printRuleinfo tree |
/tmp/vln010-raw/e3-*.txt, h3-*.txt |
-v traces, both directions |
/tmp/vln010-block{C,D,E,F,H}.txt |
the verification blocks in sequence, including the two harness defects and their fixes |
/tmp/vln010-blockI.txt |
the 29-assertion re-verification against the committed file |
/tmp/vln010, /tmp/vln010-v21 |
the scratch ruleset roots |
Verified live on the manager on 2026-08-11, during deployment (§7.2):
- the four-algorithm digest of the installed artifact, and of the pre-change backup, each checked before the restart was permitted
wazuh-analysisdPID 1491909 → 2361899 at 18:29:50 UTC- D1 node count
8506and the D2 compiled tree, fromwazuh-logtest-legacy - D3/D4/D5 from the running daemon via
wazuh-logtest, including themailflags andfiredtimes - 58 lines of load-time output from the recorded byte offset, error-free
Evidence base for the deployment, retained on MAPLE until D6 reports:
| Artifact | Contents |
|---|---|
/tmp/vln010-deploy-{1..5}-*.txt |
pre-flight, install, restart, acceptance, re-parse |
/var/ossec/backup/rules/gpus-lolbin-rules.xml.bak-2026-08-11 |
the rollback path, verified against PRE sha256 |
The pre-flight gate caught a PROMPT error, not a system error
Worth recording alongside §3.7, because it is the same principle pointing in an unusual direction.
The deployment instructions were re-issued after the deployment had already
happened. Their pre-flight required wazuh-analysisd to be PID 1491909,
started Jul 30 20:08:21, and to STOP otherwise. Live state was PID
2361899 — because the deploy had run an hour earlier. The gate refused.
Re-running would have caused real damage, not wasted effort. It would have burned a second restart on a healthy manager, and — worse — the backup step would have captured the post-change file as the "pre-change" backup, silently overwriting the only path back. The rollback would have restored the fix.
The gate was written to catch drift in the system. It caught an error in the instructions instead, and the check that saved it was the one asserting an expected value rather than merely reporting the observed one. A pre-flight that prints state is documentation; a pre-flight that asserts state is a control.
The one-shot load-time capture is UNSPENT — and turns out to be unnecessary
§4 and the runbook both treated the next manager restart as a scarce,
one-time chance to capture load-time output, because the Jul 30 output
survived only in a rotated archive. The pre-restart ossec.log snapshot
and byte offset are still in place and no restart has been performed.
It was never needed for this question. wazuh-logtest-legacy -d -d
gives the rule count and the entire compiled tree, from disk, in its own
process, with no restart and no daemon involvement. The scarce resource was
never scarce; it was just the only source anybody had thought of.
100010/100011 — settled, and my earlier explanation was also wrong
v1.1 carried these as an open question, and the previous session's working
explanation was a "first-field pre-filter" skipping rules whose audit.key
could not match. That was wrong too. The curl-100010.log trace shows
Wazuh trying every sibling in level order regardless of whether its fields
could possibly match:
100012 and 100015 are both tried against a lolbin_download event they
cannot match. There is no pre-filter. 100011 is absent from that trace
for exactly one reason: 100010 matched first and evaluation stopped.
Fully explained by evaluation order plus stop-at-first-match — the same mechanism as the finding itself. Closed, not open.
Timestamp precision — softening a v1.0 correction
v1.0 corrected the record to say wazuh-analysisd started at 20:08:21 and
that "20:08:24 is remoted". That was over-stated. Both timestamps are
real and describe different events: the process forked at 20:08:21, and
it logged Total rules enabled: '8521' at 20:08:24 — the same second
wazuh-remoted started, which is what caused the confusion. The
substantive claim is unaffected either way: the ruleset load happened
after the 20:00 file write.
That the central evidence for a detection finding cannot be re-checked without
a human running sudo is not a footnote — it is §8.5 restated as a
consequence, and it is the strongest argument for the read-only group
membership.
10. References¶
soc/wazuh-rules/gpus-lolbin-rules.xml— the rules under diagnosissoc/wazuh-rules/test-vectors/— six captured vectors, includingaide-interactive-MUSTFIRE.log, the interactive regression casesoc/wazuh-rules/EXTRACTED.md— the note whose claims §8.1 and §8.3 revisitforms-backend/wazuh-rules/gpus-forms-portal-rules.xml— rules100026–100029(§8.1) and the five furthertype="pcre2"conditions (§3.5)soc/forms-authz-detection/forms_authz_rules.xml— rule100034, the working control case in §3.5- VLN-009 — WDC passwordless sudo — §8.5 belongs on its least-privilege follow-on list
- VLN-011 — Forms HappyFox dispatch never ran — same failure spine: green checks over a control doing nothing