Bypassing WAFs with Unicode Compatibility
When a WAF inspects input before Unicode normalization but the backend processes it after, compatibility characters can slip through.
Bypassing WAFs with Unicode Compatibility#
Web Application Firewalls (WAFs) often rely on blacklists. They block <script>, javascript:, and alert(. But what if we can write these words without using standard ASCII characters?
The Magic of Unicode Normalization#
Many systems normalize input before processing it. This means they convert "fancy" characters into their standard ASCII equivalents.
<(Fullwidth Less-Than) becomes<script(Fullwidth Latin) becomesscript℡(Telephone Sign) might becomeTEL
The Attack#
If the WAF checks the input before normalization, but the backend application processes it after normalization, we have a bypass.
Example: XSS#
WAF Rule: Block <script>
Payload: <script>alert(1)</script>
Flow:
- WAF: Sees
<script>. This does not match<script>. PASS. - Backend: Normalizes input.
<script>becomes<script>. - Execution: The browser executes the script.
Finding Compatible Characters#
You can use the IDNA (Internationalizing Domain Names in Applications) standard to find these mappings.
Ican be represented byⅠ(Roman Numeral One)Kcan be represented byK(Kelvin Sign)
Conclusion#
Unicode is vast and complex. Whenever you face a WAF, check if the application performs normalization. It might be your golden ticket.
Normalization Forms#
Unicode defines several normalization forms: NFD splits characters into base plus mark, NFKD further applies compatibility decomposition, NFC recomposes. Compatibility characters are why < maps to <: it is a presentation form of the ASCII less-than.
import unicodedata print(unicodedata.normalize('NFKD', '<script>')) # <script>
NFKD is the dangerous one for filters. It folds fullwidth, ligatures, and some circled letters into ASCII.
Real-World Variations#
- Fullwidth Latin:
ABC->ABC - Roman numerals:
Ⅰ Ⅱ Ⅲ->I II III - Kelvin sign:
K->K - Ligatures:
fi->fi - Parenthesized letters:
⒜->a
Bypass Conditions#
The bypass needs an ordering gap:
- WAF decodes and pattern-matches.
- Backend normalizes later (framework, database collation, or browser).
- Browser or DB executes the normalized result.
If the WAF normalizes too, the trick collapses. Test by sending the fullwidth form and checking whether the backend error message echoes the normalized ASCII.
Detection on the Defensive Side#
Normalize at the edge. Convert every input to NFKC at the reverse proxy, then evaluate rules. Log the raw and normalized forms side by side; a mismatch between them flags a probe.
# Conceptual: do not ship as-is; normalization belongs in code, not nginx. # Log raw vs normalized at the app entry instead.
Also deny or flag control-character runs and mixed-script identifiers (a homoglyph а from Cyrillic next to Latin a). Mixed scripts in a single token is a strong probe signal.
Lab Exercise#
# 1. Start a vulnerable target docker run -p 8080:80 vulnerable/app:latest # 2. Send the fullwidth payload through a proxy curl 'http://localhost:8080/search?q=%EF%BC%9Cscript%EF%BC%9E' # 3. Compare response with the ASCII equivalent curl 'http://localhost:8080/search?q=<script>'
If the fullwidth request returns a processed page and the ASCII one returns a WAF block, you have confirmed the gap.
Homoglyph Hunt Spots#
- Scope or origin checks that compare
example.combut acceptеxample.com(Cyrillic e). - Path allowlists where the normalize step runs after the match.
- Header values that pass IDNA-unaware filters.
IDNA and Unicode#
Domain names go through a Punycode conversion (IDNA2008). Some lookalikes stay visually plausible in the browser: раураl.com (Cyrillic) points somewhere else entirely. Browsers now flag high-risk mixes with a warning, but internal dashboards may not.
Filter Gap Analysis#
raw = "<script>" fullwidth = "<script>" def naive_filter(s): return "<script>" not in s print(naive_filter(raw)) # False -> blocked print(naive_filter(fullwidth)) # True -> passes filter
The gap exists when normalize runs after the filter. Test in a lab, not on a live service.
Encoding Stack#
| Step | Example |
|---|---|
| Percent-decode | %EF%BC%9C -> < |
| Unicode normalize | < -> < |
| Route match | path rule hits <script> |
Each step is a place where a filter can sit. Map the pipeline order on your target; the bypass lives in the gap.
Charset Confusion#
UTF-7 and legacy charsets produce similar gaps. Input declared as UTF-8 but parsed as UTF-7 can let +ADw-script+AD4- pass a UTF-8 regex and render as <script> in an old IE context. Modern browsers mitigate, but kiosk and legacy apps still ship the bug.
Practical Reporting#
Include the raw request (hex of the fullwidth bytes), the normalized response body, and the exact pipeline order that your testing observed. A WAF vendor reads that and can fix the ordering, not just the rule.
Mixed-Script Detection#
A token that mixes Latin and Cyrillic characters is rare in legitimate input. Flag them:
def mixed(s): scripts = set() for ch in s: if 'a' <= ch.lower() <= 'z': scripts.add('latin') if 'а' <= ch.lower() <= 'я': scripts.add('cyrillic') return len(scripts) > 1
A probe that uses mixed scripts is either a test or an attack; either way it deserves a log line.
ZAP and Normalization#
When ZAP reports a fullwidth payload as a hit, verify whether the response actually normalized it. Scanner output reports the match on the request side; the backend behavior decides whether it is real.
IDNA Display#
Some old browsers show a Unicode domain's ASCII lookalike, not its real characters. Operators reviewing certificates can miss this. Use whois and Punycode form (xn--...) to confirm what the cert names.
Normalization in Frameworks#
- Node:
String.prototype.normalize('NFKD'). - Python:
unicodedata.normalize. - Java:
Normalizer.normalize. - Go:
golang.org/x/text/unicode/norm.
Whichever form your stack uses decides which bypasses exist. Test the library's default on a scratch string before calling any filter sufficient.
Charset Test List#
Send the same payload in UTF-8, UTF-16LE, UTF-16BE, and GBK where the target accepts those encodings. Different decoders normalize differently, and the WAF usually tests only one.
Lab Matrix#
For each candidate, run:
- Fullwidth payload through the WAF and the direct backend
- NFKD normalization before vs after the filter
- Mixed-script token detection on the watchlist
- UTF-7 or alternate codec fallback test
- IDNA normalization on any URL parameter
- Charset comparison across UTF-8 and UTF-16
Reference Tables#
| Form | Result |
|---|---|
| Fullwidth | <script> becomes <script> |
| Ligature | fi folds into fi |
| Roman numeral | I folds into Ⅰ |
| Kelvin sign | K folds into K |
Reminders#
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
- confirm before reporting
Command Cheatsheet#
curl 'https://target/?q=%EF%BC%9Cscript%EF%BC%9E' curl 'https://target/?q=<script>' python3 -c "print('<script>')"
Final Notes#
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
- confirm, document, report
Short Notes#
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
- document the observation, then the conclusion
Decoder Ordering Map#
| Step | Example |
|---|---|
| Percent-decode | %EF%BC%9C -> < |
| Normalize | < -> < |
| Match | filter or route compare |
| Execute | backend runs the normalized string |
Closing Notes#
- Verify each observation with a second independent check.
- Prefer packet, request/response, or log evidence over prose.
- Tie the finding to the control that should have caught it.
- Redact secrets but keep the request shape visible.
- Never quote a payload you have not actually tested.
- Keep one repro per file; name it clearly.
- Update the matrix when a new tool or a new gadget appears.
- Shorten CI runs; the tool should fit in a normal deploy.
- Confirm a false positive before you report it.
- A finding without evidence is folklore.
- Document the exact time delta or row count.
- Record the binary version or framework version in the report.
- Every tool output in the report ties to a command.
- The evidence tree should let a second reader replay the finding.
- Log the encoding and the raw bytes with the finding.
- If a fix closes one probe, re-run the others to be sure.
- Closing one instance rarely closes the class.
- Rotate pretexts and rate limits so the test keeps signal.
- Track report rate, not click rate, for awareness campaigns.
- Keep capture files short; slice the interesting window.
- ZAP baselines belong in CI, full scans on staging.
- Burst the auth endpoint, then verify no lockout.
- Diff lockfiles before running a supply-chain claim.
- Parameter pollution hides in headers and cookies too.
- Checksum your dependencies and rotate keys on a schedule.
- Time-based blind is slow; always pair it with a boolean check.
- Show the GRANTS output; it proves reachability limits.
- DOM sinks are data-flow endpoints, not just payload targets.
- Trusted Types and CSP sit on the same defense line.
- Stash the tshark JSON slice with the pcap for reference.
- A short pcap with notes beats a ten-minute capture.
- Name files with date, host, and window.
- A flat periodic line in the IO graph is a beacon candidate.
- Spike bursts in Burp are usually manual, not scanner.
- Tune Nuclei rate limits so the range is not blocked.
- sqlmap risk/level flags widen the payload classes.
- Arjun finds parameters that do not appear in the URL.
- Keep WordPress plugin nonces in a separate test case.
- XML-RPC is the slow, quiet way in. Disable it if unused.
- Backup files in the docroot are findings by themselves.
- Upload extension checks should run on the normalized name.
- Homoglyphs hide inside scope lists and allowlists.
- Normalize at the edge, log raw, never filter before decode.
- UTF-7 and legacy codecs keep bypassing naive filters.
- A fullwidth probe that passes is a bug in the pipeline.
- Rotate keys, pin versions, and rotate them again after a patch.
- Prototype pollution turns one key into a global change.
__proto__is the first key to block; checkconstructortoo.- Run
npm lsdeep; the vulnerable path is rarely the top level. - Refuse
constructor.prototypekeys at every merge point. - Freeze Object.prototype for the server if you must merge.
What do you think?
React to show your appreciation