혼동 문자를 안전한 텍스트로 정규화하기
혼동 문자 정규화란 닮은꼴 문자를 각각 흉내 내는 일반 문자로 바꿔, 똑같아 보이는 두 문자열이 비교할 때도 같도록 만드는 것입니다. 사용자 이름 비교, 차단 목록 대조, 식별자 중복 제거 전에 하면 좋습니다.
예제 풀이
- 입력
- Cоnfig file: аdmin2, Noёl, café
- 발견된 문자 체계
- 로마자 (Latn), 키릴 문자 (Cyrl)
- 표시된 문자 수
- 7
| 위치 | 문자 | 코드 포인트 | 문자 체계 | 대체 문자 | 규칙 |
|---|---|---|---|---|---|
| 0 | C | U+FF23 | 로마자 | C | NFKC 정규화 |
| 1 | о | U+043E | 키릴 문자 | o | 혼동 문자 매핑표 |
| 7 | fi | U+FB01 | 로마자 | fi | NFKC 정규화 |
| 10 | : | U+FF1A | 일반 문자 | : | 혼동 문자 매핑표 |
| 12 | а | U+0430 | 키릴 문자 | a | 혼동 문자 매핑표 |
| 17 | 2 | U+FF12 | 일반 문자 | 2 | 혼동 문자 매핑표 |
| 22 | ё | U+0451 | 키릴 문자 | ë | 혼동 문자 매핑표 |
- 읽을 수 있는 유니코드 유지
- Config file: admin2, Noël, café
- 엄격한 ASCII 대체
- Config file: admin2, Noel, café
작동 방식
- 먼저 알려진 닮은꼴 문자를 내장 매핑표에 따라 바꿉니다. 매핑표에 없는 문자는 NFKC로 정규화되어 전각 형태, 합자, 기타 호환 문자가 표준 문자로 바뀝니다.
- '읽을 수 있는 유니코드 유지'는 매핑표에 정의된 경우 악센트가 있는 문자를 남깁니다(예: 키릴 문자 ё → ë). '엄격한 ASCII 대체'는 대신 일반 ASCII 문자를 사용합니다(ё → e).
- 매핑표에 없고 NFKC로도 바뀌지 않는 문자는 그대로 유지되므로 café는 두 모드 모두에서 악센트를 유지합니다. 정규화된 텍스트는 비교용 키일 뿐 보안 판정이 아닙니다. 원문도 함께 보관하고 표시된 문자를 검토하세요.
변환은 최선의 노력입니다. 매핑된 혼란 항목 및 NFKC 접기는 결정적이지만 일부 합법적인 유니코드는 플래그가 지정되지 않습니다.
당신의 텍스트
붙여넣기 또는 입력 - 입력할 때 결과가 업데이트됩니다(긴 입력의 경우 약간 디바운싱됨).
30자 스캔됨
7개 의심 항목
엄격한 ASCII 대체
원본(의심스러운 문자가 표시됨)
원래 보기의 의심스러운 문자에는 밑줄이 그어져 있고 'susp'라는 라벨이 붙어 있습니다. 하이라이트 컬러 외에도
suspicious character Csuspicious character оnfig suspicious character filesuspicious character : suspicious character аdminsuspicious character 2, Nosuspicious character ёl, café
정리된 출력
문자 분석
| 인덱스(0 기반) | 원본 | 교체 | 코드 포인트 | 이유 |
|---|---|---|---|---|
| 0 | C | C | U+FF23 | NFKC normalization changed this character (compatibility or width folding). |
| 1 | о | o | U+043E | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 3 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 4 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 5 | g | g | U+0067 | Not flagged as a confusable or compatibility character. |
| 6 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 7 | fi | fi | U+FB01 | NFKC normalization changed this character (compatibility or width folding). |
| 8 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 9 | e | e | U+0065 | Not flagged as a confusable or compatibility character. |
| 10 | : | : | U+FF1A | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 11 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 12 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 13 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 14 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 15 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 16 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 17 | 2 | 2 | U+FF12 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 18 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 19 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 20 | N | N | U+004E | Not flagged as a confusable or compatibility character. |
| 21 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 22 | ё | e | U+0451 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 23 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 24 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 25 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 26 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 27 | a | a | U+0061 | Not flagged as a confusable or compatibility character. |
| 28 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 29 | é | é | U+00E9 | Not flagged as a confusable or compatibility character. |