Skip to content

Normalize string encodings before column edits - #3

Merged
quinnj merged 1 commit into
mainfrom
investigate/string-ownership
Sep 23, 2026
Merged

quinnj merged 1 commit into
mainfrom
investigate/string-ownership

Conversation

@quinnj

@quinnj quinnj commented Sep 22, 2026

Copy link
Copy Markdown
Member

StringVector copied arbitrary AbstractString code units as UTF-8. UTF-16/32 inputs failed even though DataString(input) worked, and Latin-1 inputs could silently produce incorrect text.

Normalize other string implementations through String at the shared append boundary. String, SubString{String}, DataString, and SubString{DataString} keep their direct UTF-8 paths. This fixes construction and all column edits without a new dependency or API. Other string implementations may allocate a temporary converted String.

Validation: 3,691 full-suite assertions on Julia 1.12 and 3,689 on Julia 1.10; five static-compilation/execution checks on Julia 1.13. New tests cover UTF-16/32, Latin-1, retained values, and zero-allocation UTF-8 edits with spare capacity. Original and fixed direct paths both allocate zero bytes in repeated warmed checks.

Co-authored by Codex

@quinnj
quinnj merged commit a762d73 into main Sep 23, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant