[ENG-4059] Decode request values the way apps read them - #298
Merged
Merged
Conversation
Percent-decoding keeps decoding the escapes around a % that starts no escape, and decodes each run as UTF-8. The urldecode mutation reads + as a space, as form decoding does, and REQUEST_URI and all read + in the query as a space while keeping a + in the path literal. HTML character references decode the same way in normalization and in the htmlentitydecode mutation: named references such as : and numeric references without a closing semicolon. Text mutations (urldecode, htmlentitydecode, base64_decode) applied to a structured value decode each string inside it and keep the structure, within the same bounds as the leaf walk. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Feature adds thorough request value decoding aligning with app behavior. 🎯 Quality: 100% Elite · 📦 Size: Large — consider splitting if possible 🤖 Authorship: Agent-written — Claude Code, going by its own attribution. Whether a person read it is unknown; coding agents can call the 📈 This month: Your 161st PR — above team average · Averaging Excellent |
…the bound Request normalization walks nested values iteratively, to the same depth bound as the engine's leaf walk, so every value that walk reaches is normalized first. A value past the bound is still matched in its raw form, and reaching the bound is reported as a `container-cap` skip through `onSkip` and `coverage()`, once per request. An evaluation result carries the inspection limits it reached as `skips`, and the resolver records them with `noteSkip`. Adds coverage for `+` in a request target given only as `url`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
patchstackdave
added this pull request to stack #318
September 29, 2026 08:03
Contributor
Author
|
/review |
devlob
approved these changes
Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rules now see request values decoded the way the app reads them.
Percent-decoding. A
%that starts no escape no longer stops the escapes around it from being decoded. This applies both to normalization and to theurldecodemutation. Each run of escapes is decoded as UTF-8, and bytes that are not valid UTF-8 become U+FFFD.+. Theurldecodemutation reads+as a space, as form decoding does.server.REQUEST_URIandallread+in the query as a space. A+in the path, and an encoded%2B, stay a literal+.HTML character references. Normalization and the
htmlentitydecodemutation now share one decoder. It handles named references such as:, numeric references without a closing;, and an upper-case hex marker. Text after a decimal reference is kept (:abcbecomes:abc).Structured values.
urldecode,htmlentitydecodeandbase64_decodeapplied to an object or array decode each string inside it and keep the structure. This works forpost.*values and for the output ofjson_decode, and structural matchers still see the shape. The walk is iterative, is bounded like the engine's leaf walk, and handles cyclic values and an own__proto__key.Nesting depth. Normalization walks nested values iteratively, to the same depth bound as the engine's leaf walk, so every value that walk reaches is normalized. A value past the bound is still matched, in its raw form. Reaching the bound is reported as a
container-capskip throughonSkipandcoverage(), once per request. Evaluation results carry the limits they reached asskips.Synthetic tests in
tests/protect/request-decoding.test.tspair every value a rule must match with one it must not, and cover both sides of the depth bound.Validation: full suite (3,466 passed, 7 skipped), typecheck, and build.
🤖 Generated with Claude Code