Skip to content

[ENG-4059] Decode request values the way apps read them - #298

Merged
patchstackdave merged 2 commits into
mainfrom
fix/request-decoding
Sep 29, 2026
Merged

patchstackdave merged 2 commits into
mainfrom
fix/request-decoding

Conversation

@patchstackdave

@patchstackdave patchstackdave commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Rules now see request values decoded the way the app reads them.

  • Percent-decoding. A % that starts no escape no longer stops the escapes around it from being decoded. This applies both to normalization and to the urldecode mutation. Each run of escapes is decoded as UTF-8, and bytes that are not valid UTF-8 become U+FFFD.

  • +. The urldecode mutation reads + as a space, as form decoding does. server.REQUEST_URI and all read + in the query as a space. A + in the path, and an encoded %2B, stay a literal +.

  • HTML character references. Normalization and the htmlentitydecode mutation now share one decoder. It handles named references such as :, numeric references without a closing ;, and an upper-case hex marker. Text after a decimal reference is kept (&#58abc becomes :abc).

  • Structured values. urldecode, htmlentitydecode and base64_decode applied to an object or array decode each string inside it and keep the structure. This works for post.* values and for the output of json_decode, and structural matchers still see the shape. The walk is iterative, is bounded like the engine's leaf walk, and handles cyclic values and an own __proto__ key.

  • Nesting depth. Normalization walks nested values iteratively, to the same depth bound as the engine's leaf walk, so every value that walk reaches is normalized. A value past the bound is still matched, in its raw form. Reaching the bound is reported as a container-cap skip through onSkip and coverage(), once per request. Evaluation results carry the limits they reached as skips.

Synthetic tests in tests/protect/request-decoding.test.ts pair every value a rule must match with one it must not, and cover both sides of the depth bound.

Validation: full suite (3,466 passed, 7 skipped), typecheck, and build.

🤖 Generated with Claude Code

Percent-decoding keeps decoding the escapes around a % that starts no
escape, and decodes each run as UTF-8. The urldecode mutation reads + as
a space, as form decoding does, and REQUEST_URI and all read + in the
query as a space while keeping a + in the path literal.

HTML character references decode the same way in normalization and in
the htmlentitydecode mutation: named references such as : and
numeric references without a closing semicolon.

Text mutations (urldecode, htmlentitydecode, base64_decode) applied to a
structured value decode each string inside it and keep the structure,
within the same bounds as the leaf walk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderbuds

coderbuds Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Feature adds thorough request value decoding aligning with app behavior.

🎯 Quality: 100% Elite · 📦 Size: Large — consider splitting if possible

🤖 Authorship: Agent-written — Claude Code, going by its own attribution. Whether a person read it is unknown; coding agents can call the report-ai-usage tool to say.

📈 This month: Your 161st PR — above team average · Averaging Excellent

See how your team is trending →

…the bound

Request normalization walks nested values iteratively, to the same depth
bound as the engine's leaf walk, so every value that walk reaches is
normalized first. A value past the bound is still matched in its raw
form, and reaching the bound is reported as a `container-cap` skip
through `onSkip` and `coverage()`, once per request.

An evaluation result carries the inspection limits it reached as
`skips`, and the resolver records them with `noteSkip`.

Adds coverage for `+` in a request target given only as `url`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@patchstackdave
patchstackdave added this pull request to stack #318 September 29, 2026 08:03
@patchstackdave

Copy link
Copy Markdown
Contributor Author

/review

@patchstackdave
patchstackdave merged commit 7fbb160 into main Sep 29, 2026
18 checks passed
@patchstackdave
patchstackdave deleted the fix/request-decoding branch September 29, 2026 09:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants