Every symbol in an HTML page can be written in up to four ways. The copyright sign can be typed as ©, or written as the named reference ©, the decimal reference © or the hexadecimal reference ©. All four give the same character. In a page saved as UTF-8 you can type most symbols directly, so references are needed in two cases: characters that HTML itself treats as markup, and characters that are hard to type, invisible or easy to confuse with something else.
The tables below list the symbols people look up most often. They are generated from two primary sources, not copied from other reference pages:
- the WHATWG named character reference list, as published in entities.json (file last modified 12 November 2025). It has 2,231 entries: 2,125 names that end in a semicolon and 106 older names that also work without one.
- the Unicode Character Database 18.0, for the code points and the character names.
When a symbol has no name in the list, the table says “none”. Those symbols still work, but only as numeric references or as the character itself.
The five characters HTML needs escaped
Only five characters can change how HTML is parsed. Everything else is just text.
| Character | Why it matters | Write it as |
|---|---|---|
& | Starts a character reference everywhere | & |
< | Starts a tag in text content | < |
> | Ends a tag; escaping it avoids confusion in hand-written markup | > |
" | Ends a double-quoted attribute value | " |
' | Ends a single-quoted attribute value | ' or ' |
The browser applies the same rule when it turns a DOM tree back into markup. The HTML serialization algorithm (what innerHTML returns) replaces &, <, > and the no-break space U+00A0 in text and in attribute values, and also " inside attribute values. It never writes any other named reference.
For text that comes from users or other systems, the OWASP Cross Site Scripting Prevention Cheat Sheet recommends encoding all five, with ' as ', so the same output is safe in text and in quoted attributes. This escaping covers HTML text and quoted attribute values only. JavaScript strings, URLs and CSS each have their own escaping rules.
Encoding the string Tom's "Fish & Chips" <Special> with the HTML Entity Encoder / Decoder gives:
Tom's "Fish & Chips" <Special>
The status line reports “Encoded — 6 characters converted”: one apostrophe, two double quotes, one ampersand and the two angle brackets.
Named, decimal and hex references
A named reference is & + a name from the WHATWG list + ;. Names are easier to read in source code than numbers, but they exist for only 1,446 single characters (plus 93 names that produce a pair of code points, covered below). Names are case-sensitive: Δ is Δ and δ is δ. © is a valid second name for ©, but &Copy; is not a name at all and stays as plain text.
A decimal reference is &# + the code point in base 10 + ;. A hexadecimal reference is &#x + the code point in base 16 + ;. The hex form matches the U+ notation used in character charts, so U+2713 CHECK MARK is ✓, and the decimal form of the same character is ✓. Numeric references work for every Unicode character, including emoji: 😀 (U+1F600) is 😀 or 😀.
Always end a reference with a semicolon. The 106 legacy names (©, &, <,  , ½ and others) are still recognized without one, but the parser reports an error and applies the special rules shown in the parser section below.
HTML symbol tables
Each row gives every name the WHATWG list defines for that character. The names in a row are interchangeable, so pick the one your team finds easiest to read. Copy from the Symbol column if you want the character itself.
Spaces and invisible characters
These are the symbols where a reference helps most, because you cannot see the character in the source. keeps two words on the same line. ­ (soft hyphen) is shown only when the browser breaks a long word at that point. The zero-width space has five names; four of them (​ and similar) are long names you will rarely see in ordinary pages.
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
| (space) |   |   |   | NO-BREAK SPACE |
| (space) |   |   |   | EN SPACE |
| (space) |   |   |   | EM SPACE |
| (space) |     |   |   | THIN SPACE |
| (space) |     |   |   | HAIR SPACE |
| (invisible) | ​ ​ ​ ​ ​ | ​ | ​ | ZERO WIDTH SPACE |
| (invisible) | ‌ | ‌ | ‌ | ZERO WIDTH NON-JOINER |
| (invisible) | ‍ | ‍ | ‍ | ZERO WIDTH JOINER |
| (invisible) | ­ | ­ | ­ | SOFT HYPHEN |
| (invisible) | ⁠ | ⁠ | ⁠ | WORD JOINER |
Quotes, dashes and punctuation
Curly quotes, dashes and the ellipsis are the characters word processors insert automatically, so they appear in most text pasted from documents. The en dash is for ranges (1990–2026), the em dash for breaks in a sentence. The hyphen U+2010 is different from the hyphen-minus - on the keyboard.
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
– | – | – | – | EN DASH |
— | — | — | — | EM DASH |
‐ | ‐ ‐ | ‐ | ‐ | HYPHEN |
‘ | ‘ ‘ | ‘ | ‘ | LEFT SINGLE QUOTATION MARK |
’ | ’ ’ ’ | ’ | ’ | RIGHT SINGLE QUOTATION MARK |
‚ | ‚ ‚ | ‚ | ‚ | SINGLE LOW-9 QUOTATION MARK |
“ | “ “ | “ | “ | LEFT DOUBLE QUOTATION MARK |
” | ” ” ” | ” | ” | RIGHT DOUBLE QUOTATION MARK |
„ | „ „ | „ | „ | DOUBLE LOW-9 QUOTATION MARK |
« | « | « | « | LEFT-POINTING DOUBLE ANGLE QUOTATION MARK |
» | » | » | » | RIGHT-POINTING DOUBLE ANGLE QUOTATION MARK |
‹ | ‹ | ‹ | ‹ | SINGLE LEFT-POINTING ANGLE QUOTATION MARK |
› | › | › | › | SINGLE RIGHT-POINTING ANGLE QUOTATION MARK |
… | … … | … | … | HORIZONTAL ELLIPSIS |
• | • • | • | • | BULLET |
· | · · · | · | · | MIDDLE DOT |
† | † | † | † | DAGGER |
‡ | ‡ ‡ | ‡ | ‡ | DOUBLE DAGGER |
¶ | ¶ | ¶ | ¶ | PILCROW SIGN |
§ | § | § | § | SECTION SIGN |
′ | ′ | ′ | ′ | PRIME |
″ | ″ | ″ | ″ | DOUBLE PRIME |
¡ | ¡ | ¡ | ¡ | INVERTED EXCLAMATION MARK |
¿ | ¿ | ¿ | ¿ | INVERTED QUESTION MARK |
Legal, units and marks
Å is a second name for Å U+00C5, not for the separate ANGSTROM SIGN U+212B. The service mark ℠ and degree Celsius ℃ have no names.
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
© | © © | © | © | COPYRIGHT SIGN |
® | ® ® ® | ® | ® | REGISTERED SIGN |
™ | ™ ™ | ™ | ™ | TRADE MARK SIGN |
℠ | none | ℠ | ℠ | SERVICE MARK |
℗ | ℗ | ℗ | ℗ | SOUND RECORDING COPYRIGHT |
° | ° | ° | ° | DEGREE SIGN |
‰ | ‰ | ‰ | ‰ | PER MILLE SIGN |
№ | № | № | № | NUMERO SIGN |
℃ | none | ℃ | ℃ | DEGREE CELSIUS |
Å | Å Å | Å | Å | LATIN CAPITAL LETTER A WITH RING ABOVE |
Currency symbols
The older currency signs have names. The rupee, won, ruble and bitcoin signs do not, so they need numeric references or the character itself. The fullwidth yen sign used in Japanese text is a separate character, U+FFE5 (¥).
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
$ | $ | $ | $ | DOLLAR SIGN |
¢ | ¢ | ¢ | ¢ | CENT SIGN |
£ | £ | £ | £ | POUND SIGN |
¤ | ¤ | ¤ | ¤ | CURRENCY SIGN |
¥ | ¥ | ¥ | ¥ | YEN SIGN |
€ | € | € | € | EURO SIGN |
₹ | none | ₹ | ₹ | INDIAN RUPEE SIGN |
₩ | none | ₩ | ₩ | WON SIGN |
₽ | none | ₽ | ₽ | RUBLE SIGN |
₿ | none | ₿ | ₿ | BITCOIN SIGN |
Math and comparison
The minus sign − (U+2212) is a different character from the hyphen-minus - on the keyboard, so text search and code treat 5 − 3 and 5 - 3 as different strings. ∆ INCREMENT (U+2206) has no name, but the Greek capital Δ does (Δ).
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
− | − | − | − | MINUS SIGN |
× | × | × | × | MULTIPLICATION SIGN |
÷ | ÷ ÷ | ÷ | ÷ | DIVISION SIGN |
± | ± ± ± | ± | ± | PLUS-MINUS SIGN |
≠ | ≠ ≠ | ≠ | ≠ | NOT EQUAL TO |
≈ | ≈ ≈ ≈ ≈ ≈ ≈ | ≈ | ≈ | ALMOST EQUAL TO |
≡ | ≡ ≡ | ≡ | ≡ | IDENTICAL TO |
≤ | ≤ ≤ | ≤ | ≤ | LESS-THAN OR EQUAL TO |
≥ | ≥ ≥ ≥ | ≥ | ≥ | GREATER-THAN OR EQUAL TO |
∞ | ∞ | ∞ | ∞ | INFINITY |
√ | √ √ | √ | √ | SQUARE ROOT |
∑ | ∑ ∑ | ∑ | ∑ | N-ARY SUMMATION |
∏ | ∏ ∏ | ∏ | ∏ | N-ARY PRODUCT |
∫ | ∫ ∫ | ∫ | ∫ | INTEGRAL |
∂ | ∂ ∂ | ∂ | ∂ | PARTIAL DIFFERENTIAL |
∆ | none | ∆ | ∆ | INCREMENT |
∇ | ∇ ∇ | ∇ | ∇ | NABLA |
∈ | ∈ ∈ ∈ ∈ | ∈ | ∈ | ELEMENT OF |
∉ | ∉ ∉ ∉ | ∉ | ∉ | NOT AN ELEMENT OF |
∩ | ∩ | ∩ | ∩ | INTERSECTION |
∪ | ∪ | ∪ | ∪ | UNION |
⊂ | ⊂ ⊂ | ⊂ | ⊂ | SUBSET OF |
⊃ | ⊃ ⊃ ⊃ | ⊃ | ⊃ | SUPERSET OF |
∀ | ∀ ∀ | ∀ | ∀ | FOR ALL |
∃ | ∃ ∃ | ∃ | ∃ | THERE EXISTS |
∴ | ∴ ∴ ∴ | ∴ | ∴ | THEREFORE |
Fractions, superscripts and Greek letters
MICRO SIGN µ (U+00B5, µ) and GREEK SMALL LETTER MU μ (U+03BC, μ) look the same in most fonts but are different characters. A search for one does not find the other.
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
½ | ½ ½ | ½ | ½ | VULGAR FRACTION ONE HALF |
¼ | ¼ | ¼ | ¼ | VULGAR FRACTION ONE QUARTER |
¾ | ¾ | ¾ | ¾ | VULGAR FRACTION THREE QUARTERS |
⅓ | ⅓ | ⅓ | ⅓ | VULGAR FRACTION ONE THIRD |
⅔ | ⅔ | ⅔ | ⅔ | VULGAR FRACTION TWO THIRDS |
¹ | ¹ | ¹ | ¹ | SUPERSCRIPT ONE |
² | ² | ² | ² | SUPERSCRIPT TWO |
³ | ³ | ³ | ³ | SUPERSCRIPT THREE |
µ | µ | µ | µ | MICRO SIGN |
α | α | α | α | GREEK SMALL LETTER ALPHA |
β | β | β | β | GREEK SMALL LETTER BETA |
γ | γ | γ | γ | GREEK SMALL LETTER GAMMA |
δ | δ | δ | δ | GREEK SMALL LETTER DELTA |
ε | ε ε | ε | ε | GREEK SMALL LETTER EPSILON |
θ | θ | θ | θ | GREEK SMALL LETTER THETA |
λ | λ | λ | λ | GREEK SMALL LETTER LAMDA |
μ | μ | μ | μ | GREEK SMALL LETTER MU |
π | π | π | π | GREEK SMALL LETTER PI |
σ | σ | σ | σ | GREEK SMALL LETTER SIGMA |
φ | φ | φ | φ | GREEK SMALL LETTER PHI |
ω | ω | ω | ω | GREEK SMALL LETTER OMEGA |
Δ | Δ | Δ | Δ | GREEK CAPITAL LETTER DELTA |
Σ | Σ | Σ | Σ | GREEK CAPITAL LETTER SIGMA |
Ω | Ω Ω | Ω | Ω | GREEK CAPITAL LETTER OMEGA |
Arrows
The plain arrows have up to five names each. → is the shortest name for →, and ⇒ is the shortest for the double arrow ⇒.
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
← | ← ← ← ← ← | ← | ← | LEFTWARDS ARROW |
→ | → → → → → | → | → | RIGHTWARDS ARROW |
↑ | ↑ ↑ ↑ ↑ | ↑ | ↑ | UPWARDS ARROW |
↓ | ↓ ↓ ↓ ↓ | ↓ | ↓ | DOWNWARDS ARROW |
↔ | ↔ ↔ ↔ | ↔ | ↔ | LEFT RIGHT ARROW |
↕ | ↕ ↕ ↕ | ↕ | ↕ | UP DOWN ARROW |
⇐ | ⇐ ⇐ ⇐ | ⇐ | ⇐ | LEFTWARDS DOUBLE ARROW |
⇒ | ⇒ ⇒ ⇒ ⇒ | ⇒ | ⇒ | RIGHTWARDS DOUBLE ARROW |
⇔ | ⇔ ⇔ ⇔ ⇔ | ⇔ | ⇔ | LEFT RIGHT DOUBLE ARROW |
↵ | ↵ | ↵ | ↵ | DOWNWARDS ARROW WITH CORNER LEFTWARDS |
↩ | ↩ ↩ | ↩ | ↩ | LEFTWARDS ARROW WITH HOOK |
⇧ | none | ⇧ | ⇧ | UPWARDS WHITE ARROW |
Check marks, stars, card suits and shapes
The light check mark ✓ has a name and the heavy ✔ does not. The same happens with shapes: □ △ ○ have names, while the filled ■ ▲ ● do not. Some of these characters also have an emoji form. Unicode’s emoji variation sequences list ☎, the four card suits and ✔ with both styles: add U+FE0E (︎) after the character to ask for the plain text style, or U+FE0F (️) for the color emoji style.
| Symbol | Named references | Decimal | Hex | Unicode name |
|---|---|---|---|---|
✓ | ✓ ✓ | ✓ | ✓ | CHECK MARK |
✔ | none | ✔ | ✔ | HEAVY CHECK MARK |
✗ | ✗ | ✗ | ✗ | BALLOT X |
✘ | none | ✘ | ✘ | HEAVY BALLOT X |
★ | ★ ★ | ★ | ★ | BLACK STAR |
☆ | ☆ | ☆ | ☆ | WHITE STAR |
♠ | ♠ ♠ | ♠ | ♠ | BLACK SPADE SUIT |
♣ | ♣ ♣ | ♣ | ♣ | BLACK CLUB SUIT |
♥ | ♥ ♥ | ♥ | ♥ | BLACK HEART SUIT |
♦ | ♦ ♦ | ♦ | ♦ | BLACK DIAMOND SUIT |
♪ | ♪ | ♪ | ♪ | EIGHTH NOTE |
♫ | none | ♫ | ♫ | BEAMED EIGHTH NOTES |
♀ | ♀ | ♀ | ♀ | FEMALE SIGN |
♂ | ♂ | ♂ | ♂ | MALE SIGN |
☎ | ☎ | ☎ | ☎ | BLACK TELEPHONE |
◊ | ◊ ◊ | ◊ | ◊ | LOZENGE |
■ | none | ■ | ■ | BLACK SQUARE |
□ | □ □ □ | □ | □ | WHITE SQUARE |
▲ | none | ▲ | ▲ | BLACK UP-POINTING TRIANGLE |
△ | △ △ | △ | △ | WHITE UP-POINTING TRIANGLE |
▼ | none | ▼ | ▼ | BLACK DOWN-POINTING TRIANGLE |
▽ | ▽ ▽ | ▽ | ▽ | WHITE DOWN-POINTING TRIANGLE |
● | none | ● | ● | BLACK CIRCLE |
○ | ○ | ○ | ○ | WHITE CIRCLE |
A reference only selects a character. Whether it shows up as a glyph depends on the fonts on the reader’s device. If a font has no glyph for the character, the browser tries other fonts, and if none has it you see an empty box. Writing 🧬 instead of 🧬 does not change that.
Symbols with several names, and names that make two characters
Most characters with names have only one, but some have many. ≈ ALMOST EQUAL TO has six: ≈, ≈, ≈, ≈, ≈ and ≈. → has five, and the zero-width space has five. A decoder accepts all of them. An encoder that writes names has to choose one, which is why two tools can give different output for the same input.
93 names produce two code points. They are mathematical symbols with a combining overlay. ≂̸ is U+2242 followed by U+0338 COMBINING LONG SOLIDUS OVERLAY (≂̸), and <⃒ is < followed by U+20D2 COMBINING LONG VERTICAL LINE OVERLAY. If your code counts characters after decoding, these names count as two.
Where the parser surprises you
The HTML parser accepts the 106 legacy names without a semicolon, and it takes the longest name that matches. That creates a well-known bug with query strings in visible text:
Source: example.com/?lang=en®ion=us
Shown: example.com/?lang=en®ion=us
® is a legacy name for ®, so in text the parser replaces it even though no semicolon follows. Inside an attribute such as href, a legacy name followed by = or a letter or digit is left alone, so the link itself still works and only the visible text is wrong. Writing &region fixes it in both places. The specification gives the same example: I'm ¬it; I tell you displays as “I’m ¬it; I tell you”, because ¬ is matched, while ∉ is a complete name for ∉.
Numeric references have their own rules (numeric character reference end state):
€toŸdo not give the C1 control characters. The parser maps them with the Windows-1252 table, because old pages written on Windows used these numbers for typographic characters.–becomes – (en dash),’becomes ’ and€becomes €.�, numbers aboveand surrogates such as�become U+FFFD REPLACEMENT CHARACTER (�).- An emoji must be written as one reference with its real code point. Encoders that loop over UTF-16 code units write two surrogate references instead, and each of them turns into �.
Double escaping is the other common fault. When text that is already escaped is escaped again, < becomes &lt;, and readers see the literal characters < on the page. Escape once, at the point where the text is written into HTML, and store the raw text everywhere else.
Checking symbols with the HTML Entity tool
The HTML Entity Encoder / Decoder runs in your browser and does two things:
- Encode writes the five markup characters as
&,<,>,"and', and every character above U+007F as a decimal reference, one reference per code point. It never writes named references, so use the tables above when you want a name. - Decode gives the input to the browser’s own HTML parser, so it follows every rule in the previous section, including legacy names without a semicolon and the Windows-1252 mapping.
For example, the text Early bird: 5 € — “save 20%” ✓ encodes to:
Early bird: 5 € — “save 20%” ✓
The output is pure ASCII, which helps when a template system or mail gateway does not handle UTF-8 correctly. In a UTF-8 page you only need the first five characters escaped, and the rest can stay as typed. If you paste example.com/?lang=en®ion=us into the input and click Decode, you will see the ® from the previous section appear.
Escaping text in templates and scripts
Most template engines escape text for you. Vue’s text interpolation ({{ }}) inserts data as plain text, and in the browser, setting textContent instead of innerHTML never creates markup. When you write the escaping yourself, use the standard library function of your language.
Python’s html.escape escapes the five characters by default (quote=True), with ' as ':
import html
html.escape("Tom's \"Fish & Chips\" <Special>")
# 'Tom's "Fish & Chips" <Special>'
html.unescape("© 2026 — ✓")
# '© 2026 — ✓'
In PHP, htmlspecialchars uses ENT_QUOTES | ENT_SUBSTITUTE | ENT_HTML401 by default since PHP 8.1.0, so single quotes become '. Before 8.1 the default was ENT_COMPAT, which left single quotes unescaped. Pass the flags explicitly if your code may run on older versions.
JavaScript has no built-in escape function. A replacement function must replace & first, otherwise it escapes its own output:
const escapeHtml = (s) => s
.replaceAll('&', '&')
.replaceAll('<', '<')
.replaceAll('>', '>')
.replaceAll('"', '"')
.replaceAll("'", ''');
escapeHtml('AT&T <b>"ok"</b>'); // 'AT&T <b>"ok"</b>'
None of these functions turns © or → into names. In UTF-8 output that is correct: the character itself is valid HTML. Use the named or numeric codes from the tables when you write markup by hand, when the character is invisible, or when you need to tell look-alike characters apart in the source.