fix(import): read a priority-less syslog line, widen detection, and name the file - #51
Merged
Merged
Conversation
…ity (#50) The priority exists only on the wire. A captured file routinely has none — rsyslog's default on-disk format is a timestamp, a host and a tag — and the importer read all of it as free text: the host and the tag stayed inside the message, and the Hostname and App columns came back empty. A file you cannot filter by host or application is most of the way back to a wall of text. It also produced the report that opened #50. Anonymous mode saw the hostname still sitting at the front of the message, rewrote it to a stand-in, and the result read as though the importer had prefixed every line with "example-01.invalid". So after a timestamp is recognised, "HOST TAG[PID]: MSG" is read out of what follows, and the tag comes off through the wire parser's own extractor rather than a second definition of what a tag looks like. The year and the zone stay the format panel's, not the wire parser's clock-relative guess. The first token is the only judgement call, and it is guarded: a severity word is never a hostname, and the tag must be a real "something:" token, so "21:42:40 WARN queue: depth 812" and "INFO worker pool started" keep their message intact. Six such lines are pinned in a test. In syslog mode a priority-less line was treated as the tail of the one above it, which folded a whole rsyslog file into a handful of messages. It is now read as what it is.
…ith a corpus Detection read a timestamp at the front of a line and a severity word standing on its own. Measured against the formats people actually have on disk, that left whole families unread: the level glued to the date (klog), three integers between the timestamp and the level (logcat), a bracketed header with the year last (Apache), an epoch and nothing else (Squid), a timestamp in the middle of the line (access logs), and the two structured formats — JSON and logfmt — whose fields it never looked at because they are not where it looks. Automatic detection is now a chain, from the formats that announce themselves to the ones that must be inferred. Order is precedence, and every branch before the last is anchored and specific, so a line reaches the general detector only when nothing recognised it outright. That is what lets the panel be wide without the wide part being a guess: a JSON object, a klog line and an access line each look like exactly one thing. - shapes.go holds the four that a wider timestamp list could never reach - JSON, access and logfmt are delegated to the readers already written for them, each behind a test of what the line must start with - slash dates (Go, nginx), numeric zones (zap), a zone after a space (Serilog), day-and-month without a year (logcat), Ruby's letter-and-comma prefix, and the three-letter level spellings INF, WRN, DBG, VRB and FTL - the preview now counts the lines whose host and application were read The corpus test is the specification: 27 lines, one per format, taken from what each tool really writes — and four that must stay unrecognised, because a wide recogniser earns its width only by staying silent on a line it does not understand. An invented hostname is worse than an unparsed line: the text still shows while the field quietly lies. tools/sample-logs/auto-formats.txt is the same corpus as a file to import.
Reported: an Apache access log imports correctly, "Access log" sits in the format list, and nothing anywhere says that is what the file is. Parsing it right while the panel still reads "automatic detection" is correct and unconvincing at the same time — the operator has no way to know the tool understood the file rather than merely tolerated it. Every branch of the recogniser now says what it recognised, the shapes are counted per file, and the file is named from them: - a clear majority names the file, not a plurality. Half access lines and half something else is "mixed", because announcing an access log there would be a confident answer to a question that has none - a file that is mostly unrecognised says "mixed" when anything at all was read, and stays silent when nothing was - when the named shape has a format of its own, the selector switches to it and says it chose it. Only from automatic, and only once: a format chosen by hand is a decision, and overruling it would be the application arguing with the operator - the sample is then read again through that format, so what is on screen is what the chosen mode produces Android's brief logcat form — "E/ActivityManager( 1234): ANR in ..." — is recognised too. It was the shape the sample file actually used, and the test that reads every sample is what noticed the corpus disagreed with its own README. Verified over files rather than lines: samples_test.go reads all thirteen sample files and checks the count, the verdict, the matching mode and the fields of the first message. A file with no expectation fails the test, so adding a sample means making a claim about it. A second test declares each file's detected mode and checks nothing is lost by it — the dialog offers that switch, and an operator who takes it must not get a worse result for agreeing with us.
The list offered the six parsing ENGINES while the recogniser knew twenty formats. A file could be reported as "Kubernetes klog" and then not be in the list at all — the same gap as before, one level down, and the reason several RFCs hid behind a single entry called "Syslog". What a reader thinks in is a format's name, not an engine's. So the list is names now, grouped by what they are: Syslog, Web servers, Structured, Platforms, and one's own pattern. Several names map to one engine, and a couple carry field names with them — picking "JSON, Serilog compact" fills in @t, @l and @m rather than leaving them to be looked up. An entry's label is the same string the verdict uses, so a file cannot be called one thing above the list and another inside it. Five shapes that automatic detection read on its own are now modes too: syslog without a priority, Kubernetes klog, Android logcat, the Apache error log, and an epoch at the front of the line. A name the dialog reports has to be a name the reader can choose — and choosing it reads strictly. A file declared as klog that is not klog comes back with every line unrecognised, which is the honest answer; five such pairs are pinned in a test. The sample corpus now says so too: the mode each file is detected as is part of what samples_test.go checks, and a second test declares that mode over the same file to make sure nothing is lost by agreeing with the detector.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #50.
1. What was reported
It adds nothing. The reporter's file is what rsyslog writes to disk by default — RFC 3164 without a priority, because the priority only exists on the wire:
Run through the importer as it shipped:
The timestamp was read and everything after it was treated as free text: the Hostname and App columns stayed empty, and the hostname sat at the front of the message, where anonymous mode correctly rewrote it to a stand-in. What looked like an invented prefix was the reporter's own host.
Now
HOST TAG[PID]: MSGis read out of what follows a timestamp, with the tag coming off through the wire parser's own extractor rather than a second definition of what a tag is. The year and zone stay the format panel's. In syslog mode a priority-less line used to be treated as the tail of the one above it, folding a whole rsyslog file into a handful of messages.The first token is the only judgement call and it is guarded — a severity word is never a hostname, and the tag must be a real
something:token. Six lines that must keep their message intact are pinned, including21:42:40 WARN queue: depth 812.2. Detection, widened
Fixing that showed how much else was missed. Automatic detection is now a chain, from the formats that announce themselves to the ones that must be inferred; every branch before the last is anchored and specific, so a line reaches the general detector only when nothing recognised it outright.
INFWRNDBGVRBFTL3. The file is now named, and the dialog says so
Reported after the above: an Apache access log imports correctly, "Access log" sits in the format list, and nothing says that is what the file is.
Every branch now reports what it recognised, the shapes are counted per file, and the file is named from them:
Verified
Lines —
TestAuto_Corpus: 27 lines, one per format, taken from what each tool really writes. Four of them must stay unrecognised: a sentence with a colon, a level word followed by a colon, a stack-trace line, a line with no shape. A wide recogniser earns its width only by staying silent on what it does not understand — an invented hostname is worse than an unparsed line, because the text still shows while the field quietly lies.Files —
TestSampleFiles: all thirteen files intools/sample-logs/, checked for the message count, the verdict, the matching mode and the fields of the first message.A file with no expectation fails that test, so adding a sample means making a claim about it. A second test declares each file's detected mode and checks nothing is lost by it: the dialog offers that switch, and an operator who takes it must not get a worse result for agreeing with us.
That corpus already earned its keep — it is what caught that
logcat.txtused the brief form, which nothing recognised, while the README claimed a custom pattern was needed for it. The brief form is recognised now.go test ./...,go vet,gofmt,svelte-check --fail-on-warnings, the frontend build and the site's link check are clean. Strings added across all 8 locales (471 keys × 8, parity checked).Not in this PR
That nothing on screen said the display was anonymised is #49, in #52.