feat(string): Implement UTF-8 string conversion and validation functions - #2528
Conversation
|
| Filename | Overview |
|---|---|
| Core/Libraries/Source/WWVegas/WWLib/utf8.cpp | Implements width-aware UTF-8 transcoding with checks for malformed, overlong, surrogate, and out-of-range sequences. |
| Core/Libraries/Source/WWVegas/WWLib/utf8.h | Defines the portable conversion API, invalid-input sentinel, sizing contracts, and destination-buffer behavior. |
| Core/GameEngine/Source/Common/System/AsciiString.cpp | Replaces lossy wide-to-byte casting with correctly sized UTF-8 encoding. |
| Core/GameEngine/Source/Common/System/UnicodeString.cpp | Adds validated UTF-8 decoding while preserving legacy non-UTF-8 data through byte-wise fallback. |
| Core/GameEngine/Source/GameNetwork/GameSpy/Thread/ThreadUtils.cpp | Migrates GameSpy string conversion from Win32 APIs to WWLib and handles invalid UTF-8 through the documented fallback. |
| Core/Libraries/Source/WWVegas/WWLib/CMakeLists.txt | Adds the portable UTF-8 source and header to the WWLib target. |
Flowchart
%%{init: {'theme': 'neutral'}}%%
flowchart TD
A[Engine or GameSpy text] --> B{Conversion direction}
B -->|UTF-8 to wide| C[Utf8_To_Wide_Len]
C --> D{Valid UTF-8?}
D -->|Yes| E[Utf8_To_Wide]
D -->|No| F[Legacy one-to-one byte fallback]
B -->|Wide to UTF-8| G[Wide_To_Utf8_Len]
G --> H[Wide_To_Utf8]
E --> I[UnicodeString or std::wstring]
F --> I
H --> J[AsciiString or std::string]
Reviews (17): Last reviewed commit: "build: Use qualified include path for WW..." | Re-trigger Greptile
|
xezon
left a comment
There was a problem hiding this comment.
Get_Utf8_Size should not include the null terminator in its size.
|
The diff now shows unrelated changes. |
39d7229 to
40393b8
Compare
Try again cleaned up the commits and force pushed |
| } | ||
| ensureUniqueBufferOfSize((Int)size + 1, false, nullptr, nullptr); | ||
| char* buf = peek(); | ||
| if (!Unicode_To_Utf8(buf, src, srcLen, size)) |
There was a problem hiding this comment.
So is this translating UTF16LE from windows into UTF8 that is then stored within AsciiString?
If so this may help with the paths issue with usernames and paths not using Latin characters, but file handling functions will need updating to use unicode variants instead of Ascii.
There was a problem hiding this comment.
Yes. It does not fix path handling on its own though, the file APIs still take the narrow string. Should that be a separate change?
|
I wonder if we should also add a flag to state that the Ascii string is holding a UTF8 string? I guess all normal ascii characters will display properly, it's just extended character sets that will look garbled. |
…ming, null-terminate when room
|
Slight nitpick on function naming, shouldn't it be Utf16 rather than Wchar as Wchar is only Utf16 on windows yet we will likely want the conversion functions on other platforms as well for at least csf parsing if nothing else. |
Yeah, but it's consistent with the other naming as it is for now. How about we keep the naming and using a uint16_t/char16_t type internally rather than wchar_t when we make non windows paths? Or would you rather we rename to something like Utf16_To_Utf8? |
If anything it would be Utf16Le_To_Utf8, windows uses the little endian utf16 format. Not sure if Utf16Be is used much anywhere but worth being concise with it. |
|
Pushed updates addressing the open comments:
Tested the codec in isolation: ASCII / 2- / 3- / 4-byte round-trips on both wide widths, plus RFC 3629 rejection of overlong, surrogate, out-of-range, and truncated sequences. Ready for another pass |
|
Reworked the conversion interface: Utf8_To_Wide_Len / Utf8_To_Wide return UTF8_INVALID on bad input. 0 now means empty only, so callers can tell the two apart. |
xezon
left a comment
There was a problem hiding this comment.
I think it is a good idea to make this custom implementation and be independent of Windows API.
OmniBlade
left a comment
There was a problem hiding this comment.
Not a fan of having different encodings on different platforms personally and went with an ICU4C approach as suggested by CryoTheRenegade for OpenW3D which keeps the representation as utf-16 for Wide strings on all platforms. The refactor to support that is larger though and it adds a dependency.
|
Can we use ICU for non VC6 and whatever else for VC6 ? |
Kind of, but we'd want to replace wchar_t with UTF-16 first. I think we merge this, then use OmniBlade’s OpenW3D approach to refactor WideChar/UnicodeString to fixed 16-bit storage. Then we can use ICU because both VC6 and non VC6 builds would share the same internal representation and public contract, only the backend would differ. |
xezon
left a comment
There was a problem hiding this comment.
Looks good. A few things can be polished.
| do | ||
| { | ||
| c = wcschr(dest, L'\n'); | ||
| c = wcschr(&ret[0], L'\n'); |
There was a problem hiding this comment.
Any idea why this is removing \r\n ? Probably should have been a different function doing that.
| do | ||
| { | ||
| c = wcschr(dest, L'\r'); | ||
| c = wcschr(&ret[0], L'\r'); |
There was a problem hiding this comment.
While at it, maybe optimize these loops. It is searching for the character from the beginning of the string on every iteration. Or better in a follow up change.
Adds portable UTF-8/wide-character conversion helpers to WWLib and uses them for engine string translation and the existing GameSpy conversion paths.
This picks up the work from #2045 with API and behavior changes based on review feedback.
WWLib/utf8.h/utf8.cppAdds four conversion functions:
Wide_To_Utf8_LenUtf8_To_Wide_LenWide_To_Utf8Utf8_To_WideThe
_Lenfunctions return the required destination size without counting a null terminator. The conversion functions return the number of units written and write a null terminator when the destination has spare capacity.Malformed UTF-8 is reported as
UTF8_INVALID, keeping it distinct from a successful empty conversion, which returns0.The decoder validates UTF-8 according to RFC 3629, rejecting:
The encoder substitutes U+FFFD for unpaired UTF-16 surrogates and wide values that have no valid UTF-8 representation. Its output therefore always decodes successfully through
Utf8_To_Wide.The implementation uses no platform text APIs. It adapts to the platform width of
wchar_t:wchar_tis handled as UTF-16, including surrogate pairs;wchar_tis handled as UTF-32 codepoints.AsciiString::translateReplaces the original 7-bit-only wide-to-narrow translation with UTF-8 encoding.
UnicodeString::translateDecodes valid UTF-8 into the platform wide-character representation.
When the input is not valid UTF-8, translation falls back to the previous unsigned-byte-to-wide-character behavior. This preserves legacy byte values instead of rejecting or replacing the input. It is a compatibility fallback, not a complete code-page-specific decoder.
Because
AsciiStringdoes not carry encoding metadata, valid UTF-8 is preferred over the legacy fallback.ThreadUtils.cppReplaces the GameSpy-specific Win32 conversion calls with the shared WWLib functions.
MultiByteToWideCharSingleLineuses the same invalid-UTF-8 fallback asUnicodeString::translate. Newline and carriage-return replacement remains unchanged.Non-goals
This change does not make narrow filesystem APIs Unicode-aware. Paths are still passed to the existing narrow platform APIs; updating those APIs belongs in a separate change.
Verification
wchar_tpaths.