Use iconv() in utf8.cpp module on non-Windows platforms. - #174
Open
daliborfox wants to merge 1 commit into
Open
daliborfox wants to merge 1 commit into
daliborfox wants to merge 1 commit into
Conversation
|
Development builds of d8915fb:
The links work without a GitHub account. Artifacts expire after 90 days, and this comment follows the latest successful build. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The UTF-8 module contains a Best_Fit_Index() function which maps Unicode values to Windows-1252 or OEM 437 values, and uses WideCharToMultiByte() for this purpose, combined with a caching mechanism.
This commit adds support for this function on non-Windows platforms by using iconv() to perform the character conversion.
I've prepared the following test program to check that it works: utf8-test.cpp
Comparing the differences between Linux and Windows, the only one I see is that Linux / glibc's iconv() converts u2022, which corresponds to the 0x07 bullet in OEM 437, to a lower-case 'o'. I'd imagine this is benign, since the values below 0x20 are discarded anyways (is this intentional for OEM 437? It does contain graphical symbols there):

The one somewhat not-nice thing is that since the UTF8 module doesn't seem to contain a shutdown routine / destructor, the conversion state variables never get freed.