feat(amwscan): map the Exploit vocabulary from the engine's own words - #97
Conversation
#96 made the parenthesised token the discriminator. The Function tokens were already in the table; the Exploit ones were not, and on the account this came from that is 779 findings, 732 of them `execution` alone, all landing on other/medium/heuristic. AMWScan prints what each pattern means, on the "- " line under every finding. Each mapping quotes that description beside it, so the next person can check the mapping against the source rather than against my reading of a token name: execution "RCE ... execute PHP code on the target machine via HTTP" -> backdoor / critical, the same thing eval means nano "a family of PHP webshells ... code golfed to be stealthy" -> webshell / critical clever_include "LFI ... inject and execute arbitrary commands or code" -> injection / high, matching the existing `include` entry infected_comment "comments composed by 5 random chars usually used to detect if a file is infected yet" -> other / high: a marker something left behind, not a technique the file performs The obfuscation family — base64_long, hex_char, double_var2, concat_vars_array, concat_vars_with_spaces — sits at MEDIUM, matching the existing `encoded` entry rather than `obfuscated`. Every one of their descriptions says the technique is USUALLY used for malicious code, and usually is the operative word: minified libraries, licence blobs and legitimate encoders trip the same patterns. A heuristic that fires on ordinary vendor code at high severity is one whose severity stops meaning anything. Two are deliberately left unmapped, with the reasoning written beside the ones that are: etc_passwd true of an attacker, and also of every config parser, test fixture and tutorial that names the path php_uname one information-gathering call that installers make legitimately; the engine files it under RCE, which is more than a single call earns They stay unknown, which means other/medium and a line in the note counts, so they keep showing up as something to decide about. An unmapped rule is not a discarded one. Severity does not feed the score — confidence does, and all of these are heuristic — so no verdict moves. What moves is what a person reads when deciding which of 779 findings to look at first. Confirmed by mutation in both directions: removing `execution` fails with 'category "other", wanted "backdoor"', and raising hex_char to high fails with 'obfuscation alone is not a reason to act'.
|
A/B against the real engines — the "no verdict moves" claim holdsThis finished after the merge, so posting it for the record rather than as a gate. Same image, same corpus, only the new table entries differing: And the votes are byte-identical, which is the part that actually settles it — the summary could match by coincidence, the multipliers cannot:
Worth contrasting with #96, where I made the same claim and it was half wrong: |
Deployed, and the arc closes on the account this started from
The categories across that account's votes, across today:
Confidences unchanged at The 17 still in That number is now small enough to be read rather than skimmed, which was the point. It started at 180. |



#96 made the parenthesised token the discriminator. The
Functiontokens were already in the table; theExploitones were not — and on the account this came from that is 779 findings, 732 of themexecutionalone, all landing onother/medium/heuristic.Mapped from what the engine says, not from what the names suggest
AMWScan prints a description under every finding. Each mapping quotes it in the table, so the next person can check the mapping against the source rather than against my reading of a token name.
executionnanoclever_includeinfected_commentexecutionlands whereevallands because the engine describes them as the same thing.clever_includegets exactly what the table already gives plaininclude.infected_commentis a marker something left behind rather than a technique the file performs, which is why it is not backdoor or webshell.The obfuscation family sits at medium, on purpose
base64_long,hex_char,double_var2,concat_vars_array,concat_vars_with_spaces— every one of their descriptions says the technique is usually used for malicious code, and usually is the operative word. Minified libraries, licence blobs and legitimate encoders trip the same patterns.They match the existing
encodedentry rather thanobfuscated. A heuristic that fires on ordinary vendor code at high severity is one whose severity stops meaning anything.Two left unmapped, with the reasoning next to the ones that are
etc_passwd— "an attacker who has accessed the /etc/passwd file may attempt a brute force attack". True of an attacker; also true of every config parser, test fixture and tutorial that names the path. The engine is describing what the string means when an attacker wrote it, and the pattern cannot tell who did.php_uname— the engine files it under RCE. It is one information-gathering call, which diagnostics and installers make legitimately. Calling that critical on its own would put ordinary code next to webshells in the same list.They stay unknown, which means
other/mediumand a line in the note counts, so they keep surfacing as something to decide about. An unmapped rule is not a discarded one — and a test asserts they stay that way, so mapping them later has to be a decision rather than a drift.Scope
No verdict moves. Severity does not feed the score; confidence does, and all of these are
heuristic, same as the fallback they replace. What moves is what a person reads when deciding which of 779 findings to look at first.Confirmed by mutation in both directions: