fix(ci): the invisible-character gate never matched anything - #107
Conversation
MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.
ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.
grep -P '\xc2\xa0' -> miss
grep -P '\x{a0}' -> MATCH
Only \x00 worked, being single-byte in both readings.
FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.
The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.
Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📜 Recent review details⏰ Context from checks skipped due to timeout. (1)
|
| Layer / File(s) | Summary |
|---|---|
Unicode pattern and text scan .github/workflows/dogfood-gate.yml |
The pattern uses Unicode codepoint escapes and a C0 control-character range. The grep scan treats files as text and retains null-byte detection. |
Estimated code review effort: 1 (Trivial) | ~5 minutes
Merge Risk: ⚪ Minimal · up to 71b9a
This localized workflow fix corrects invisible-character detection behavior and adds handling for additional control characters; no actionable merge-blocking risk remains after normal checks and review.
Poem
A rabbit checks each hidden mark
Unicode guides the careful scan
Control bytes now leave a spark
Text files join the checking plan
The gate sees what it must detect
🚥 Pre-merge checks | ✅ 4 | ❌ 1
❌ Failed checks (1 warning)
| Check name | Status | Explanation | Resolution |
|---|---|---|---|
| Linked Issues check | The change addresses the codepoint escapes, C0 control detection, and grep -a requirements in [#70]. However, the linked issue also requires a separate leading-BOM check and matching updates to stdlib… |
Add the separate byte-wise leading-BOM check and update stdlib/ByteDetector.affine and config.ncl with the matching C0-control logic, or narrow the linked issue scope to the CI workflow change only. |
✅ Passed checks (4 passed)
| Check name | Status | Explanation |
|---|---|---|
| Title check | ✅ Passed | The title clearly identifies the main change: fixing the CI gate for invisible-character detection. |
| Description check | ✅ Passed | The description directly explains the detection failure, root cause, implemented fixes, and verification results. |
| Out of Scope Changes check | ✅ Passed | The changed workflow logic is related to the invisible-character detection requirements in [#70]. No unrelated changes are shown. |
| Docstring Coverage | ✅ Passed | No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0… |
Full details: Linked Issues check
Explanation
The change addresses the codepoint escapes, C0 control detection, and grep -a requirements in [#70]. However, the linked issue also requires a separate leading-BOM check and matching updates to stdlib/ByteDetector.affine and config.ncl, which are not shown in this pull request.
Full details: Docstring Coverage
Explanation
No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)
- Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
- Create stacked PR
- Commit on current branch
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.
Comment @coderabbitai help to get the list of available commands.
Up to standards ✅🟢 Issues
|
There was a problem hiding this comment.
Pull Request Overview
This PR successfully updates the CI gate to use Unicode codepoint escapes and the -a flag, resolving a bug where invisible characters were not being detected. While the logic for the patterns is sound and Codacy results are up to standards, there is a performance bottleneck in how find invokes grep, and a potential locale-dependency issue in the PCRE pattern.
A significant gap is the absence of automated regression tests. The PR addresses several invisible-character cases that were previously missed, but no sample files containing these characters were added to the repository to ensure the gate remains functional in the future.
About this PR
- The PR fixes detection for specific invisible characters, but it does not include any sample files or automated tests containing these characters. To prevent future regressions and verify the CI gate logic, consider adding a directory of 'malicious' files containing the targeted Unicode and C0 control characters.
Test suggestions
- Missing recommended test scenario: Verify detection of Non-Breaking Space (U+00A0) in a source file
- Missing recommended test scenario: Verify detection of Zero-Width Space (U+200B) in a source file
- Missing recommended test scenario: Verify detection of C0 Control characters (e.g., Backspace \x08) in a source file
- Missing recommended test scenario: Verify detection of NUL byte (\x00) and ensure file is not skipped by grep
- Missing recommended test scenario: Ensure standard whitespace (Tab \x09, LF \x0A, CR \x0D) does not trigger the gate
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Missing recommended test scenario: Verify detection of Non-Breaking Space (U+00A0) in a source file
2. Missing recommended test scenario: Verify detection of Zero-Width Space (U+200B) in a source file
3. Missing recommended test scenario: Verify detection of C0 Control characters (e.g., Backspace \x08) in a source file
4. Missing recommended test scenario: Verify detection of NUL byte (\x00) and ensure file is not skipped by grep
5. Missing recommended test scenario: Ensure standard whitespace (Tab \x09, LF \x0A, CR \x0D) does not trigger the gate
TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback
| -o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \ | ||
| -o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \ | ||
| -exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | ||
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
There was a problem hiding this comment.
🟡 MEDIUM RISK
Suggestion: The -r flag is redundant when grep is executed via find on individual file paths. Additionally, switching from \; to + significantly improves performance by batching file arguments into fewer grep processes. Note that EL_EXIT=$? will capture the exit status of the find command; ensure your logic remains reliant on the FINDINGS count for detection.
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | |
| -exec grep -aPl "$PATTERNS" {} + > /tmp/empty-lint-results.txt 2>/dev/null |
| # non-breaking spaces, null bytes, and other invisible Unicode in source files. | ||
| set +e | ||
| PATTERNS='\xc2\xa0|\xe2\x80\x8b|\xe2\x80\x8c|\xe2\x80\x8d|\xef\xbb\xbf|\xc2\xad|\xe2\x80\x8e|\xe2\x80\x8f|\xe2\x80\xaa|\xe2\x80\xab|\xe2\x80\xac|\xe2\x80\xad|\xe2\x80\xae|\x00' | ||
| PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}' |
There was a problem hiding this comment.
⚪ LOW RISK
Suggestion: To ensure the PCRE engine correctly interprets the Unicode hex escapes (like \x{200b}) across different environments and locales, it is safer to prefix the pattern with (*UTF8).
| PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}' | |
| PATTERNS='(*UTF8)\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}' |
Measured 2026-08-27: this gate caught 0 of 6 invisible-character test cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi override or word joiner.
Root cause
The pattern used UTF-8 byte sequences (
\xc2\xa0) whilegrep -Pmatches characters. Bytesc2 a0are one character U+00A0;\xc2\xa0asks for two, U+00C2 then U+00A0 — never present.Only
\x00worked, being single-byte in both readings. The gate ran, passed, and could not see what it exists to see.Fixed
\x01-\x08,\x0B,\x0C,\x0E-\x1Fadded (TAB/LF/CR excluded)grep -a— without it grep skips any NUL-bearing file as binaryThe C0 range matters: a stray backspace byte made a workflow unparseable in
developer-ecosystem, so it never ran — and this linter called it clean.Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
Verified: YAML re-parsed, and the corrected pattern was confirmed to catch a real NBSP before the change was kept.