Skip to content

fix(import): read a priority-less syslog line, widen detection, and name the file - #51

Merged
Wasabules merged 4 commits into
mainfrom
fix-import-bsd-syslog
Sep 29, 2026
Merged

Wasabules merged 4 commits into
mainfrom
fix-import-bsd-syslog

Conversation

@Wasabules

@Wasabules Wasabules commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Closes #50.

1. What was reported

After import of offline logs the importer adds prefix "example-01.invalid" in the message.

It adds nothing. The reporter's file is what rsyslog writes to disk by default — RFC 3164 without a priority, because the priority only exists on the wire:

Sep 18 08:05:13 nbb-ad-01.vms.nbb.nod Microsoft-Windows-Security-Auditing[756]: An account was logged on

Run through the importer as it shipped:

hostname="" app="" proc=""
message="nbb-ad-01.vms.nbb.nod Microsoft-Windows-Security-Auditing[756]: An account was logged on"

The timestamp was read and everything after it was treated as free text: the Hostname and App columns stayed empty, and the hostname sat at the front of the message, where anonymous mode correctly rewrote it to a stand-in. What looked like an invented prefix was the reporter's own host.

Now HOST TAG[PID]: MSG is read out of what follows a timestamp, with the tag coming off through the wire parser's own extractor rather than a second definition of what a tag is. The year and zone stay the format panel's. In syslog mode a priority-less line used to be treated as the tail of the one above it, folding a whole rsyslog file into a handful of messages.

The first token is the only judgement call and it is guarded — a severity word is never a hostname, and the tag must be a real something: token. Six lines that must keep their message intact are pinned, including 21:42:40 WARN queue: depth 812.

2. Detection, widened

Fixing that showed how much else was missed. Automatic detection is now a chain, from the formats that announce themselves to the ones that must be inferred; every branch before the last is anchored and specific, so a line reaches the general detector only when nothing recognised it outright.

Added Why detection could not reach it
Kubernetes klog the level is a letter glued to the date
Android logcat, both forms integers between the timestamp and the level; or no timestamp at all
Apache error log bracketed header, year last
Squid and friends an epoch and nothing else
access logs the timestamp is in the middle of the line
JSON, logfmt the fields are not where detection looks
Go, nginx slash dates
zap, Serilog numeric zone, zone after a space
Ruby, Rails a letter and a comma before the bracket
INF WRN DBG VRB FTL three-letter level spellings

3. The file is now named, and the dialog says so

Reported after the above: an Apache access log imports correctly, "Access log" sits in the format list, and nothing says that is what the file is.

Every branch now reports what it recognised, the shapes are counted per file, and the file is named from them:

The import dialog naming the file

  • a clear majority names the file, not a plurality. Half access lines and half something else is mixed, because announcing an access log there would be a confident answer to a question that has none;
  • when the named shape has a format of its own, the selector switches to it and says it chose it — only from automatic, and only once, since a format chosen by hand is a decision and overruling it would be the application arguing with the operator;
  • the sample is then read again through that format, so what is on screen is what the chosen mode produces.

Verified

Lines — TestAuto_Corpus: 27 lines, one per format, taken from what each tool really writes. Four of them must stay unrecognised: a sentence with a colon, a level word followed by a colon, a stack-trace line, a line with no shape. A wide recogniser earns its width only by staying silent on what it does not understand — an invented hostname is worse than an unparsed line, because the text still shows while the field quietly lies.

Files — TestSampleFiles: all thirteen files in tools/sample-logs/, checked for the message count, the verdict, the matching mode and the fields of the first message.

access.txt               n=7   detected=access  mode=access   [access=7]
apache-error.txt         n=5   detected=apache  mode=         [apache=5]
archive-2019.txt         n=6   detected=bsd     mode=syslog   [bsd=6]
auto-formats.txt         n=23  detected=mixed   mode=         [access=1 apache=1 bsd=3 epoch=1
                                                               json=2 klog=1 logcat=1 logfmt=1
                                                               plain=10 syslog=2]
json-lines.txt           n=8   detected=json    mode=json     [json=8]
klog.txt                 n=6   detected=klog    mode=         [klog=6]
logcat.txt               n=9   detected=logcat  mode=         [logcat=9]
logfmt.txt               n=7   detected=logfmt  mode=logfmt   [logfmt=7]
messy.txt                n=8   detected=mixed   mode=         [bsd=1 json=1 none=4 plain=2]
nginx-error.txt          n=5   detected=plain   mode=         [plain=5]
plain-app.txt            n=9   detected=plain   mode=         [plain=9]
rsyslog-traditional.txt  n=6   detected=bsd     mode=syslog   [bsd=6]
syslog-capture.txt       n=8   detected=syslog  mode=syslog   [syslog=8]

A file with no expectation fails that test, so adding a sample means making a claim about it. A second test declares each file's detected mode and checks nothing is lost by it: the dialog offers that switch, and an operator who takes it must not get a worse result for agreeing with us.

That corpus already earned its keep — it is what caught that logcat.txt used the brief form, which nothing recognised, while the README claimed a custom pattern was needed for it. The brief form is recognised now.

go test ./..., go vet, gofmt, svelte-check --fail-on-warnings, the frontend build and the site's link check are clean. Strings added across all 8 locales (471 keys × 8, parity checked).

Not in this PR

That nothing on screen said the display was anonymised is #49, in #52.

…ity (#50)

The priority exists only on the wire. A captured file routinely has none —
rsyslog's default on-disk format is a timestamp, a host and a tag — and the
importer read all of it as free text: the host and the tag stayed inside the
message, and the Hostname and App columns came back empty. A file you cannot
filter by host or application is most of the way back to a wall of text.

It also produced the report that opened #50. Anonymous mode saw the hostname
still sitting at the front of the message, rewrote it to a stand-in, and the
result read as though the importer had prefixed every line with
"example-01.invalid".

So after a timestamp is recognised, "HOST TAG[PID]: MSG" is read out of what
follows, and the tag comes off through the wire parser's own extractor rather
than a second definition of what a tag looks like. The year and the zone stay
the format panel's, not the wire parser's clock-relative guess.

The first token is the only judgement call, and it is guarded: a severity word
is never a hostname, and the tag must be a real "something:" token, so
"21:42:40 WARN queue: depth 812" and "INFO worker pool started" keep their
message intact. Six such lines are pinned in a test.

In syslog mode a priority-less line was treated as the tail of the one above
it, which folded a whole rsyslog file into a handful of messages. It is now
read as what it is.
…ith a corpus

Detection read a timestamp at the front of a line and a severity word standing
on its own. Measured against the formats people actually have on disk, that
left whole families unread: the level glued to the date (klog), three integers
between the timestamp and the level (logcat), a bracketed header with the year
last (Apache), an epoch and nothing else (Squid), a timestamp in the middle of
the line (access logs), and the two structured formats — JSON and logfmt —
whose fields it never looked at because they are not where it looks.

Automatic detection is now a chain, from the formats that announce themselves
to the ones that must be inferred. Order is precedence, and every branch before
the last is anchored and specific, so a line reaches the general detector only
when nothing recognised it outright. That is what lets the panel be wide
without the wide part being a guess: a JSON object, a klog line and an access
line each look like exactly one thing.

- shapes.go holds the four that a wider timestamp list could never reach
- JSON, access and logfmt are delegated to the readers already written for
  them, each behind a test of what the line must start with
- slash dates (Go, nginx), numeric zones (zap), a zone after a space
  (Serilog), day-and-month without a year (logcat), Ruby's letter-and-comma
  prefix, and the three-letter level spellings INF, WRN, DBG, VRB and FTL
- the preview now counts the lines whose host and application were read

The corpus test is the specification: 27 lines, one per format, taken from what
each tool really writes — and four that must stay unrecognised, because a wide
recogniser earns its width only by staying silent on a line it does not
understand. An invented hostname is worse than an unparsed line: the text still
shows while the field quietly lies.

tools/sample-logs/auto-formats.txt is the same corpus as a file to import.
Reported: an Apache access log imports correctly, "Access log" sits in the
format list, and nothing anywhere says that is what the file is. Parsing it
right while the panel still reads "automatic detection" is correct and
unconvincing at the same time — the operator has no way to know the tool
understood the file rather than merely tolerated it.

Every branch of the recogniser now says what it recognised, the shapes are
counted per file, and the file is named from them:

- a clear majority names the file, not a plurality. Half access lines and half
  something else is "mixed", because announcing an access log there would be a
  confident answer to a question that has none
- a file that is mostly unrecognised says "mixed" when anything at all was
  read, and stays silent when nothing was
- when the named shape has a format of its own, the selector switches to it and
  says it chose it. Only from automatic, and only once: a format chosen by hand
  is a decision, and overruling it would be the application arguing with the
  operator
- the sample is then read again through that format, so what is on screen is
  what the chosen mode produces

Android's brief logcat form — "E/ActivityManager( 1234): ANR in ..." — is
recognised too. It was the shape the sample file actually used, and the test
that reads every sample is what noticed the corpus disagreed with its own
README.

Verified over files rather than lines: samples_test.go reads all thirteen
sample files and checks the count, the verdict, the matching mode and the
fields of the first message. A file with no expectation fails the test, so
adding a sample means making a claim about it. A second test declares each
file's detected mode and checks nothing is lost by it — the dialog offers that
switch, and an operator who takes it must not get a worse result for agreeing
with us.
@Wasabules Wasabules changed the title fix(import): read the host and tag of a priority-less syslog line, and widen what detection recognises fix(import): read a priority-less syslog line, widen detection, and name the file Sep 28, 2026
The list offered the six parsing ENGINES while the recogniser knew twenty
formats. A file could be reported as "Kubernetes klog" and then not be in the
list at all — the same gap as before, one level down, and the reason several
RFCs hid behind a single entry called "Syslog".

What a reader thinks in is a format's name, not an engine's. So the list is
names now, grouped by what they are: Syslog, Web servers, Structured,
Platforms, and one's own pattern. Several names map to one engine, and a couple
carry field names with them — picking "JSON, Serilog compact" fills in @t, @l
and @m rather than leaving them to be looked up.

An entry's label is the same string the verdict uses, so a file cannot be
called one thing above the list and another inside it.

Five shapes that automatic detection read on its own are now modes too: syslog
without a priority, Kubernetes klog, Android logcat, the Apache error log, and
an epoch at the front of the line. A name the dialog reports has to be a name
the reader can choose — and choosing it reads strictly. A file declared as klog
that is not klog comes back with every line unrecognised, which is the honest
answer; five such pairs are pinned in a test.

The sample corpus now says so too: the mode each file is detected as is part of
what samples_test.go checks, and a second test declares that mode over the same
file to make sure nothing is lost by agreeing with the detector.
@Wasabules
Wasabules merged commit 33dfc37 into main Sep 29, 2026
9 checks passed
@Wasabules
Wasabules deleted the fix-import-bsd-syslog branch September 29, 2026 09:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] SyslogStudio adds prefix "example-01.invalid" to messages

1 participant