将易混淆字符规范化为安全文本
规范化易混淆字符,就是把每个仿冒字符替换为它所模仿的普通字符,使看起来相同的两个字符串在比较时也相等。在比较用户名、匹配黑名单或对标识符去重之前进行这一步尤其有用。
实例解析
- 输入
- Cоnfig file: аdmin2, Noёl, café
- 检测到的文字系统
- 拉丁文 (Latn), 西里尔文 (Cyrl)
- 已标记字符
- 7
| 位置 | 字符 | 码位 | 文字系统 | 替换为 | 规则 |
|---|---|---|---|---|---|
| 0 | C | U+FF23 | 拉丁文 | C | NFKC 规范化 |
| 1 | о | U+043E | 西里尔文 | o | 易混淆字符映射表 |
| 7 | fi | U+FB01 | 拉丁文 | fi | NFKC 规范化 |
| 10 | : | U+FF1A | 通用 | : | 易混淆字符映射表 |
| 12 | а | U+0430 | 西里尔文 | a | 易混淆字符映射表 |
| 17 | 2 | U+FF12 | 通用 | 2 | 易混淆字符映射表 |
| 22 | ё | U+0451 | 西里尔文 | ë | 易混淆字符映射表 |
- 保留可读的 Unicode
- Config file: admin2, Noël, café
- 严格的 ASCII 后备
- Config file: admin2, Noel, café
工作原理
- 首先按内置映射表替换已知的仿冒字符。映射表之外的字符使用 NFKC 规范化,把全角形式、连字和其他兼容字符转换为标准等价字符。
- “保留可读的 Unicode”会在映射表定义了带重音字母时保留它,例如西里尔字母 ё → ë。“严格的 ASCII 后备”则改用普通 ASCII 字母(ё → e)。
- 既不在映射表中、也不会被 NFKC 改变的字母会原样保留,因此 café 在两种模式下都保留重音。规范化后的文本只是比较用的键,而不是安全结论:请同时保存原文,并检查被标记的字符。
转换是尽力而为:映射的易混淆项和 NFKC 折叠是确定性的,但某些合法的 Unicode 不会被标记。
您的文字
粘贴或键入 — 结果会在您键入时更新(对于长输入会稍微去抖)。
已扫描 30 个字符
7 个可疑项
严格的 ASCII 后备
原文(可疑字符已标记)
原始视图中的可疑字符带有下划线并标记为“可疑”。除了突出颜色。
suspicious character Csuspicious character оnfig suspicious character filesuspicious character : suspicious character аdminsuspicious character 2, Nosuspicious character ёl, café
清理后的输出
字符分析
| 索引(从0开始) | 原字符 | 替换为 | 代码点 | 原因 |
|---|---|---|---|---|
| 0 | C | C | U+FF23 | NFKC normalization changed this character (compatibility or width folding). |
| 1 | о | o | U+043E | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 3 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 4 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 5 | g | g | U+0067 | Not flagged as a confusable or compatibility character. |
| 6 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 7 | fi | fi | U+FB01 | NFKC normalization changed this character (compatibility or width folding). |
| 8 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 9 | e | e | U+0065 | Not flagged as a confusable or compatibility character. |
| 10 | : | : | U+FF1A | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 11 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 12 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 13 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 14 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 15 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 16 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 17 | 2 | 2 | U+FF12 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 18 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 19 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 20 | N | N | U+004E | Not flagged as a confusable or compatibility character. |
| 21 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 22 | ё | e | U+0451 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 23 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 24 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 25 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 26 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 27 | a | a | U+0061 | Not flagged as a confusable or compatibility character. |
| 28 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 29 | é | é | U+00E9 | Not flagged as a confusable or compatibility character. |