Skip to content

Accept ASCII in -Validate wherever BOM-less UTF-8 is allowed - #136

Merged
amrali-eg merged 1 commit into
masterfrom
fix/validate-ascii-as-utf8
Sep 28, 2026
Merged

amrali-eg merged 1 commit into
masterfrom
fix/validate-ascii-as-utf8

Conversation

@amrali-eg

Copy link
Copy Markdown
Owner

Problem

After #132, conversion to UTF-8 without a BOM leaves ASCII files unchanged ("ASCII is already valid UTF-8 without a BOM"). -Validate still compared labels, so the same pure-ASCII file failed. Reproduced with the master build:

-Validate list Before
utf-8 Invalid / CharsetNotAllowed, exit 2 under -FailOnChanges
utf-8,utf-8-bom same
utf-8-bom same

A pipeline that converted with -Target utf-8 and then validated with -Validate "utf-8,utf-8-bom" (the example in CLI.md) failed every file whose text was plain English. It also contradicted the plan guarantee text "ASCII already counts as BOM-less UTF-8".

Change

  • ConversionPolicy exposes the rule it already applies (IsAsciiAlreadyUtf8) and its explanation (AsciiAlreadyUtf8Reason). Conversion and -Validate now both use it, so they accept the same files.
  • -Validate accepts a detected ASCII file when the list allows BOM-less utf-8, and explains the pass with "ASCII is already valid UTF-8 without a BOM." A list allowing only utf-8-bom still rejects it, because an ASCII file has no BOM.
  • Full-file validation is unchanged: the whole file is still validated as ASCII, so a non-ASCII byte past the 64 KiB detection sample is reported Invalid / StrictValidationFailed, as conversion reports it.
  • CLI.md notes the rule in the -Validate row; backlog entry BL-38.

No conversion decision changes, so the semantics version stays at 8.

Tests

In AsciiAlreadyUtf8Tests:

  • A theory over utf-8, utf-8,utf-8-bom, us-ascii and utf-8-bom, asserting result, reason code, explanation and the -FailOnChanges exit code.
  • A file with ASCII in the detection sample and a UTF-8 é after it: Invalid / StrictValidationFailed, exit 2.

Mutation checks (each restored byte-for-byte afterwards)

Mutation Tests failing
The UTF-8 rule disabled in -Validate 3
Accepted ASCII skips full-file validation 2 (including the existing ConvertAndValidateNowAgreeAboutTheSameFile)
utf-8-bom also accepts ASCII 1

Verification

  • Rebased onto master after Refuse a contradicting source choice even when it is the target #135; the backlog conflict was resolved by keeping both entries (BL-37, BL-38).
  • Release build, 0 warnings; 967/967 tests pass.
  • docs/Test-DefectBacklog.ps1 passes.
  • GUI smoke not run locally for this PR; CI runs gui-smoke. The full smoke run and the four-corpus audit will be on the final committed release candidate.

🤖 Generated with Claude Code

Conversion to UTF-8 without a BOM leaves ASCII files unchanged, but
-Validate compared labels, so the same file failed -Validate "utf-8" as
CharsetNotAllowed (exit 2 under -FailOnChanges). A pipeline that converted
with -Target utf-8 and then validated "utf-8,utf-8-bom" failed every file
whose text was plain English.

-Validate now uses the conversion rule from ConversionPolicy and explains
the pass with "ASCII is already valid UTF-8 without a BOM." Only the
BOM-less utf-8 label accepts ASCII; utf-8-bom alone still rejects it. The
whole file is still validated as ASCII, so a non-ASCII byte past the
detection sample is reported invalid, as conversion reports it.

Tests cover each allowed list, the explanation and the late byte; three
mutations (rule disabled, validation skipped, utf-8-bom accepted) fail them.

Backlog: BL-38.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@amrali-eg
amrali-eg merged commit 7cd05e2 into master Sep 28, 2026
3 checks passed
@amrali-eg
amrali-eg deleted the fix/validate-ascii-as-utf8 branch September 28, 2026 21:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant