Skip to content

fix(amwscan): the discriminator is in the parentheses, and the => line is source - #96

Merged
thiagoluga merged 2 commits into
mainfrom
fix/the-discriminator-is-in-the-parentheses
Aug 18, 2026
Merged

thiagoluga merged 2 commits into
mainfrom
fix/the-discriminator-is-in-the-parentheses

Conversation

@thiagoluga

@thiagoluga thiagoluga commented Aug 18, 2026 •

Copy link
Copy Markdown
Owner

Found by reading the raw report from a live account — which only became possible with #94, because raw_ref pointed at the process's stdout until then.

What the real engine emits

 => [!] Function (exec) [line 147]
    - Potentially dangerous function `exec`
      => exec('kill -' . (int) $signal . ' ' . (int) $pid . ' 2>/dev/null', $out, $code)

Function is not in the rule table and never will be. exec is, and has been all along — the adapter captured it into a field used only for display, then classified on the rule name.

On that account, of 333 Function findings:

token count table says reported as
eval 199 backdoor / critical other / medium
assert 42 backdoor / high other / medium
exec 29 webshell / critical other / medium
shell_exec 12 webshell / critical other / medium

288 of 333, all landing on the generic fallback.

And the => line is not a category

It was read as the category the engine assigned, with priority over everything else. It is the source that matched. On that account the resulting "categories" read lave, tressa, ipconfig, suhosin — fragments of somebody's source code.

lave is eval backwards. AMWScan detects strrev-obfuscated calls and prints the reversed string it found, so an obfuscated backdoor classified as unknown while the plain word eval sat one pair of brackets away.

That text comes out of the scanned file, which means its author chooses it. Nothing today exploits that — the snippet is usually far too long to collide with a table key — but a finding's category should not be selectable by the thing being examined.

Scope — corrected after measuring it

I first wrote that verdicts and quarantine decisions do not change. That was too strong, and an A/B run against the real engines showed why.

For the Function (...) findings it holds: both eval and the unknown fallback are heuristic, the multiplier is identical, and only the category and severity a human reads to triage change — 199 eval hits shown as generic medium rather than backdoor/critical.

For Signature (<hash>) it does not hold. Those used to be classified by the snippet, and in the container's corpus that snippet is the word backdoor, which maps to a heuristic entry. They now fall through to signature:

before:  vote: amwscan  weight 0.80 x heuristic = 0.64  (rule Signature)
after:   vote: amwscan  weight 0.80 x signature = 0.80  (rule Signature)

A signature hit now carries the weight the project always intended for it — the rule table's own comment calls it "the only case where it claims to recognize the THREAT, not a pattern". It is a correctness fix, and it raises scores, which means a finding can cross a verdict threshold it previously did not.

In this corpus nothing crossed one. Same image, only the classification line reverted:

without the fix:  confirmed=1  likely=3  suspicious=3  clean=0
with the fix:     confirmed=1  likely=3  suspicious=3  clean=0

Flagging it rather than burying it: on a site with a signature hit plus one other weak vote, this can turn likely into confirmed, and confirmed is what authorises an automatic quarantine.

My first attempt at that comparison was worthless, and I want it on the record. I compared against a run in which the engines had failed to download, so it had no AMWScan at all — two numbers that differ because one scan had two more engines say nothing about a classification change.

The fixture, and the uncomfortable part

The old fixture is genuine. => backdoor really was printed — that hit's matched content happened to be the word "backdoor". One sample's coincidence was read as structure, and a parser plus a passing test were built on it.

PROVENANCE.md now records the lesson next to the earlier one it already carried: a real capture is necessary and not sufficient. Where a format matters, capture more than one file, from more than one site. The second capture disagreed with the first about the only thing that mattered.

New fixture: real-0.15.1-function-and-exploit.txt, sanitised (paths rewritten, long snippets truncated as the engine truncates them, nothing else altered), covering Function (eval|exec|assert|shell_exec|proc_open), Exploit (execution|hex_char), Signature (<hash>) and the two strrev hits.

TestAMWScanTheTagBeatsTheRuleName is replaced by TestAMWScanASignatureIsKnownMalware. Confirmed by mutation — classifying on the snippet again fails with exec: category "other", wanted "webshell".

Left deliberately undone

The Exploit vocabulary — execution (732), base64_long, hex_char, nano, etc_passwd, php_uname, clever_include — is not in the table. Each needs the engine's own description read before being mapped; those descriptions are in the report, on the - line under each finding, and the distribution is recorded in PROVENANCE.md for whoever does it. Mapping them from the names alone is the guess this PR is about not making.

…e is source

Found by reading the raw report from a live account — which only became possible
with #94, since raw_ref pointed at the process stdout until then.

Real AMWScan output names the rule generically and puts the specific thing in
parentheses:

    => [!] Function (exec) [line 147]
       - Potentially dangerous function `exec`
         => exec('kill -' . (int) $signal . ' ' . (int) $pid . ' 2>/dev/null', ...)

"Function" is not in the rule table and never will be. `exec` is, and has been
all along. The adapter captured it into a field used only for display and
classified on the rule name instead, so on that account 199 `eval`, 42 `assert`,
29 `exec` and 12 `shell_exec` findings — 288 of 333 — came out as
other/medium/heuristic.

Worse, the indented "=> ..." line was read as the category the engine assigned,
with priority over everything else. It is the SOURCE THAT MATCHED. On that
account the resulting "categories" read `lave`, `tressa`, `ipconfig`, `suhosin`
— fragments of somebody's source code. `lave` is `eval` backwards: AMWScan
detects strrev-obfuscated calls and prints the reversed string, so an obfuscated
backdoor classified as unknown while the word `eval` sat one pair of brackets
away. And since that text comes out of the scanned file, its author chooses it.

classify() now takes the parenthesised token first and falls back to the rule
name, which keeps Signature (<hash>) landing on known_malware. The matched
source is kept as evidence, under its own name, and never classified on.

Scope, stated precisely: confidence drives the score and is heuristic either
way, so verdicts and quarantine decisions do NOT change. What changes is the
category and severity a human reads to triage — 199 eval hits shown as generic
medium instead of backdoor/critical.

The test that asserted the opposite is replaced. It was D-022's shape exactly:
an assumption about an external format, plus a test written to confirm it.

The old fixture is genuine, and that is the part worth sitting with. `=>
backdoor` really was printed, because that hit's matched content happened to be
the word "backdoor". One sample's coincidence was read as structure. PROVENANCE
now records that a real capture is necessary and not sufficient, alongside a
second capture from a different site that disagrees with the first about the
only thing that mattered.
Records what #96 decided, including the part that is easy to miss: this raises
the effective weight of Signature findings from 0.64 to 0.80, because they used
to be classified by the matched source and now reach the signature confidence
the project always intended for them. A site with a signature plus one weak vote
can cross from likely to confirmed, and confirmed authorises quarantine. The
validation corpus does not cross, and that is measured rather than assumed.

And the lesson that is not the obvious one: the fixture it was all built on is
genuine. A real capture was taken and a wrong model was still built on it,
because one sample cannot say which parts of a line are structure and which are
content. D-022 says verify against reality; this adds that a real capture is
necessary and not sufficient.
@sonarqubecloud

Copy link
Copy Markdown

@thiagoluga
thiagoluga merged commit b27b5ef into main Aug 18, 2026
10 checks passed
@thiagoluga
thiagoluga deleted the fix/the-discriminator-is-in-the-parentheses branch August 18, 2026 19:04
@thiagoluga

Copy link
Copy Markdown
Owner Author

Deployed, and confirmed on the account it was found on

v0.1.12-9-gb27b5ef is on the validation host — sha256 verified before and after the move, panel answering, public in 0.56s.

The classification across that account's votes, before and after:

category before after
other 180 45
backdoor — 97
webshell — 38
known_malware 36 36

135 findings left the generic bucket. The 45 still in other are the unmapped Exploit vocabulary, which #97 addresses.

One detail worth recording, because it demonstrates the concern rather than arguing it

known_malware is 36 both before and after — but for different reasons, and the difference is the whole point.

On this account, a Signature hit's matched source is the word exploit, which is not in the rule table, so it fell through to the rule name and classified correctly by accident. In the validation container, the same rule's matched source is the word backdoor, which is in the table — so it classified as backdoor/heuristic and carried 0.64 instead of 0.80.

Same engine, same rule, two different wrong answers, decided by what the scanned file happened to contain. That is what "the file's author chooses the category" means in practice, and it is why the matched source is now evidence only.

thiagoluga added a commit that referenced this pull request Aug 18, 2026
…#97)

#96 made the parenthesised token the discriminator. The Function tokens were
already in the table; the Exploit ones were not, and on the account this came
from that is 779 findings, 732 of them `execution` alone, all landing on
other/medium/heuristic.

AMWScan prints what each pattern means, on the "- " line under every finding.
Each mapping quotes that description beside it, so the next person can check the
mapping against the source rather than against my reading of a token name:

  execution       "RCE ... execute PHP code on the target machine via HTTP"
                  -> backdoor / critical, the same thing eval means
  nano            "a family of PHP webshells ... code golfed to be stealthy"
                  -> webshell / critical
  clever_include  "LFI ... inject and execute arbitrary commands or code"
                  -> injection / high, matching the existing `include` entry
  infected_comment
                  "comments composed by 5 random chars usually used to detect
                  if a file is infected yet" -> other / high: a marker something
                  left behind, not a technique the file performs

The obfuscation family — base64_long, hex_char, double_var2, concat_vars_array,
concat_vars_with_spaces — sits at MEDIUM, matching the existing `encoded` entry
rather than `obfuscated`. Every one of their descriptions says the technique is
USUALLY used for malicious code, and usually is the operative word: minified
libraries, licence blobs and legitimate encoders trip the same patterns. A
heuristic that fires on ordinary vendor code at high severity is one whose
severity stops meaning anything.

Two are deliberately left unmapped, with the reasoning written beside the ones
that are:

  etc_passwd  true of an attacker, and also of every config parser, test fixture
              and tutorial that names the path
  php_uname   one information-gathering call that installers make legitimately;
              the engine files it under RCE, which is more than a single call
              earns

They stay unknown, which means other/medium and a line in the note counts, so
they keep showing up as something to decide about. An unmapped rule is not a
discarded one.

Severity does not feed the score — confidence does, and all of these are
heuristic — so no verdict moves. What moves is what a person reads when
deciding which of 779 findings to look at first.

Confirmed by mutation in both directions: removing `execution` fails with
'category "other", wanted "backdoor"', and raising hex_char to high fails with
'obfuscation alone is not a reason to act'.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant