Repository navigation
Some eicar files not detected #247
Description
Activity
Sorry for the delayed response. Bit of a busy week this week. I'll address this properly soon.
Thanks for your time! No worries about the delay
Hi @Maikuolan , is there any new statement on this issue ?
Sorry, I had forgotten to reply to this earlier.
New signatures have been added since the time this issue was created to help cover some of the missed detections, but most of them remain unresolved at this time. The bottleneck for resolving those, in the case of the PDF files, is phpMussel's ability to parse PDF files to the extent of being able to properly identify and decode the parts of the PDF files which actually contain the eicar text is lacking and in need of improvement, and in the case of the XLS/X and and PPT/X files, to identify the bounds of the correct segments within the relevant files of those containers in order to be able to decode and then match against the eicar text accordingly is likewise lacking and in need of improvement.
I think I'm going to need help from others if that improvement is ever going to be properly implemented, ideally, from someone that knows the formats better than I do (but failing that, from whoever might be willing to help).
The official spec for PDF (ISO 32000-1:2008) is 748 pages long - Longer than the length of many popular novels nowadays, and a rather hefty bit of homework to digest in the context of just wanting to learn enough to know how to properly parse a PDF file in order to check it against anti-virus signatures and such (and naturally, something I've avoided doing, a good part due to the time it would require to do so to the extent that I could feel confident that I know what I would need to know and some mild feelings of dread at that prospect, especially with phpMussel being an unpaid, open-source hobby project and all). In comparison, the official spec for PE (portable executable) format (i.e., Windows EXE files) is a mere 96 pages long (and I did work properly through that one from start to finish in the past, back when I was first needing to implement support for scanning PE files to phpMussel, to be able to properly process PE sectional signatures and such, and that wasn't too difficult to do, didn't take too much time, but feels comparatively much simpler than PDF looks like it would be from this position). The current iteration of phpMussel's PDF handler is the result of trying to force myself into reading at least some of the official PDF spec and covering some of the basics, like identifying different sections within a file, when a section has been compressed, knowing to decompress it (at least, for those compressions which PHP natively supports), and it does successfully detect eicar in PDF files saved using some older versions of the spec, though not very well at all on newer versions. But yeah; needs work.
As for XLS/X and PPT/X formats: I haven't been able to even locate proper specifications for those, so I couldn't guess about their length. Though, that they're simply archive containers is a simple enough thing to figure out for anyone that bothers trying to figure them out, and that they can for the most part be treated similarly to ZIP files (for newer versions of the format, at least; older versions, not so much, and detection is worse for those older versions); phpMussel already does that much by itself anyway, as it already detects when a file is an archive container and will check all the files within those containers accordingly. But, knowing how all the files within such containers relate to and interact with one another (or, more accurately, the extent to which they could relate to and interact with one another, i.e., what's theoretically possible for the format) has mostly been guesswork, as that information would typically be covered the spec (which I haven't seen, due to not being able to find). I've had some degree of success in that guesswork, but all the same, improvement is needed.
I could just grab arbitrary chunks of raw binary from the attached files in question, write some simple signatures to check against that raw binary, and that would result in the files in question being successfully detected and blocked by phpMussel, but such an approach would be superficial at best, and simply hide the real problem at worst: That aforementioned needed improvement. Any slight modification to the files, e.g., importing the files into an editor and resaving/recompiling them as so that the raw binary changes (particularly so where compression is employed) would be sufficient for working around said signatures, rendering them useless, but would also mean that any real viruses mimicking the exact same mechanisms used by the eicar text in those specific files which aren't being detected would likewise most likely not be detected.
This document is kind of ancient, something I haven't updated in a number of years now, so is likely a fair bit outdated, but is something I wrote a long time ago to attempt to document how phpMussel fairs against other engines in regard to various kinds of eicar files: https://phpmussel.github.io/comparative.html
Might be vaguely useful to know about in the context of this issue. The eicar files in question aren't the same as those attached above, but shows some other places where phpMussel's coverage is or isn't good.
I am using the latest core version (3.7.1) in my symfony app.
I have downloaded all signatures (clamav and phpmussel) from the signatures repo and activated all signatures in config file.
But there is some files that are not detected as viruses :
PDF
eicar-adobe-acrobat-attachment.pdf
eicar-adobe-acrobat-javascript-alert.pdf
XLS
eicar-excel-dde-cmd-powershell-echo.xls
eicar-excel-dde-cmd-powershell-echo.xlsx
PowerPoint
eicar-powerpoint-action-macro-msgbox.ppt
eicar-powerpoint-action-powershell-echo.ppt
eicar-powerpoint-action-powershell-echo.pptx
Other eicar files from this repo are correctly detected