A URL has at most seven parts, and two documents define them. RFC 3986 (2005) is the generic syntax that most server-side libraries follow. The WHATWG URL Standard describes what browsers actually do, including how they repair input that RFC 3986 would reject. The parts have the same names in both, but a few inputs split differently, and those differences are where most URL bugs come from.
RFC 3986 §3 draws the structure like this:
foo://example.com:8042/over/there?name=ferret#nose
\_/ \______________/\_________/ \_________/ \__/
| | | | |
scheme authority path query fragment
The authority itself has three parts: optional user info before an @, the host, and an optional port after a :. The rest of this guide takes one realistic URL apart and then covers each part in order.
One URL, every part
https://shop.example.co.uk:8443/catalog/shoes/?color=red&size=9&size=10&q=trail+runner#reviews
| Part | Value in this URL | Defined in | Browser property | Sent to the server |
|---|---|---|---|---|
| Scheme | https | RFC 3986 §3.1 | protocol → https: | Decides the protocol |
| User info | none | §3.2.1 | username, password | Not in HTTP; see below |
| Host | shop.example.co.uk | §3.2.2 | hostname | Yes, in DNS and the Host header |
| Port | 8443 | §3.2.3 | port | Decides the TCP port |
| Path | /catalog/shoes/ | §3.3 | pathname | Yes |
| Query | color=red&size=9&size=10&q=trail+runner | §3.4 | search | Yes |
| Fragment | reviews | §3.5 | hash → #reviews | No |
Pasting this URL into the URL Parser gives exactly these fields, four query parameters (size twice, and q decoded to trail runner) and no notes, which means the browser accepted the URL without changing it. The same values from code:
const u = new URL('https://shop.example.co.uk:8443/catalog/shoes/?color=red&size=9&size=10&q=trail+runner#reviews');
console.log(u.hostname); // shop.example.co.uk
console.log(u.port); // 8443
console.log(u.pathname); // /catalog/shoes/
console.log(u.searchParams.getAll('size')); // ["9","10"]
console.log(u.searchParams.get('q')); // trail runner
console.log(u.hash); // #reviews
console.log(u.origin); // https://shop.example.co.uk:8443
Scheme
The scheme is everything before the first :. RFC 3986 §3.1 says it starts with a letter, may contain letters, digits, +, - and ., and is case-insensitive with lowercase as the canonical form, so HTTPS://Shop.EXAMPLE.com/A/B becomes https://shop.example.com/A/B in a browser. The path keeps its case; only the scheme and the host are case-insensitive.
The URL Standard singles out six “special” schemes: http, https, ws, wss, ftp and file. Only these get a host that is normalized as a domain name, the backslash treatment described later, and a default port. Other schemes such as mailto:, data: or javascript: often have no // at all, and then the whole rest of the URL is an opaque path: in mailto:a@b.example there is no host, only the path a@b.example.
A scheme-less string is not a URL. example.com/path makes new URL() throw. localhost:3000/api is worse: it parses, but as the scheme localhost with the path 3000/api and no host, because localhost is a valid scheme name.
User info
https://user:pass@example.com/ carries a user name and password in the authority. RFC 3986 §3.2.1 states that the user:password format “is deprecated” and that applications should not render the text after the colon in clear text. The Fetch Standard goes further: new Request() and fetch() throw a TypeError when the URL includes credentials.
The @ is also a phishing tool. In https://paypal.com@evil.example/login, paypal.com is only a user name; the host is evil.example.
Host
RFC 3986 §3.2.2 allows three kinds of host: a registered name (example.com), an IPv4 address (192.0.2.16) and an IP literal in square brackets ([2001:db8::7]). Hosts are case-insensitive.
Browsers accept more IPv4 spellings than RFC 3986 does. The URL Standard’s IPv4 parser reads hexadecimal, octal and shortened forms, so http://0x7f.1/admin and http://2130706433/ both mean 127.0.0.1. To RFC 3986, 0x7f.1 is just a name. Code that blocks requests to 127.0.0.1 by comparing strings misses these forms.
Non-ASCII host names are internationalized domain names (IDN). DNS only carries ASCII, so each label is converted to Punycode (RFC 3492) with an xn-- prefix, after mapping rules from Unicode UTS #46 that also turn full-width letters into ASCII. https://日本語.jp/ is requested as xn--wgv71a119e.jp, and bücher.example as xn--bcher-kva.example.
The same mechanism allows look-alike domains. https://раураl.com/signin uses Cyrillic р, а and у next to a Latin l, and the browser sends xn--l-7sba6dbr.com. Chrome’s IDN policy shows such mixed-script names in Punycode in the address bar, following the restriction levels in UTS #39.
Subdomain, registrable domain and public suffix
Neither RFC 3986 nor the URL parser tells you which part of a host is the “domain”. In shop.example.co.uk the registrable domain is example.co.uk, because co.uk is a public suffix under which anyone can register names. The only way to know that is the Public Suffix List, a maintained list of such suffixes. Browsers use it to decide what counts as the same site, and RFC 6265 §5.3 uses it to reject cookies set for a public suffix such as co.uk. Splitting on the last two dots gives the wrong answer for co.uk, com.au or github.io; use a library that loads the list instead.
Port
The port follows the host after a :. A scheme can define a default: 80 for http and ws, 443 for https and wss, 21 for ftp. RFC 3986 §3.2.3 says producers “should omit the port component” when it equals the default, and browsers do: https://example.com:443/ becomes https://example.com/, and url.port is an empty string. The request still goes to port 443. Code that reads url.port to find the port must fall back to the scheme default.
Path
The path is a sequence of segments separated by /. Two segments are special: . means “this directory” and .. means “the parent”, and RFC 3986 §5.2.4 defines how to remove them. Browsers apply that algorithm to every URL, so /a/./b/../c becomes /a/c. An encoded slash is different: in /a%2Fb/c the first segment is a/b, and servers disagree about whether to decode it before routing.
Characters that may not appear raw are percent-encoded as UTF-8 bytes. The browser does this for you: a space in a path becomes %20, and お問い合わせ becomes %E3%81%8A%E5%95%8F%E3%81%84%E5%90%88%E3%82%8F%E3%81%9B. A trailing slash is significant; /catalog/shoes/ and /catalog/shoes are different paths, even if a framework routes them to the same handler.
Query
RFC 3986 §3.4 defines the query only as the text between ? and #. It does not define keys or values; the RFC merely notes that queries “are often used to carry identifying information in the form of key=value pairs”. The pairs come from the application/x-www-form-urlencoded format of HTML forms, which the URL Standard specifies and URLSearchParams implements:
- Pairs are separated by
&, and each pair splits at its first=. A pair without=is a key with an empty value. +means a space, and%2Bmeans a literal plus, soc%2B%2Bisc++.- A key can repeat.
size=9&size=10has two values;get('size')returns the first,getAll('size')both. - Bytes are decoded as UTF-8. Bytes that are not valid UTF-8, such as GBK or Shift_JIS from older sites, become U+FFFD.
Libraries differ at the edges. Python’s parse_qs also split on ; before Python 3.10. In Go 1.27, url.Parse("…?q=a+b&q=2&flag&x=%zz").Query() returned q and flag but dropped x, because %zz is not a valid escape; browsers keep %zz as text, and decodeURIComponent('%zz') throws URIError: URI malformed.
Rebuilding a query with URLSearchParams.toString() is not neutral either. new URLSearchParams('flag&x=%zz&a=%E4%B8').toString() returns flag=&x=%25zz&a=%EF%BF%BD: every pair is re-encoded, and the invalid UTF-8 escape becomes a replacement character for good. The URL Parser keeps untouched pairs byte for byte when you edit one parameter.
Fragment
The fragment is everything after #. RFC 3986 §3.5 says it is “separated from the rest of the URI prior to a dereference”, so the browser never sends it to the server. It selects an element on the page, and single-page apps with hash routing keep their whole route there: in https://m.example.cn/#/goods/detail?id=8 the query is empty and id=8 lives in the fragment.
Origin
The origin is the scheme, host and port together. It is the unit of the browser’s same-origin policy: https://example.com and https://example.com:443 are the same origin, while http://example.com and https://api.example.com are different ones. Paths, queries and fragments play no part in it.
Where browsers and RFC 3986 disagree
The URL Standard exists because browsers must accept URLs that people actually type and paste. It strips leading and trailing spaces, deletes tabs and newlines, and for special schemes reads a backslash as a slash. RFC 3986 allows none of these characters, and parsers built on it split such input differently. The sharpest case:
http://evil.example\@good.example/login
| Reader (tested 2026-10-01) | Host it uses |
|---|---|
Browser / new URL() (WHATWG) | evil.example; the path is /@good.example/login |
| RFC 3986 Appendix B regular expression | good.example; evil.example\ is user info |
Python 3.12 urllib.parse.urlsplit() | good.example |
| curl 8.7.1 | good.example (Host: good.example) |
Go 1.27 net/url.Parse | error: invalid userinfo |
When one component validates a URL and another fetches it, a gap like this lets a request reach a host that was never allowed. Orange Tsai’s Black Hat USA 2017 talk “A New Era of SSRF” collected many such parser differences. The defence is to parse once, with one parser, and pass the parsed result rather than the string. The URL Parser marks this case with a backslash note and an RFC 3986 host warning.
A checklist before trusting a URL
- Parse it with the same parser that will make the request, and compare the host, not the string.
- Check for user info. Anything before
@is not the site. - Treat an empty
portas the scheme’s default, not as “no port”. - Read the host in Unicode and Punycode, and look twice at mixed scripts.
- Decode the query with form rules (
+is a space) and expect repeated keys. - Do not look for the fragment on the server; it never arrives.
The URL Parser runs these checks in your browser: it shows each part, every change the parser made and every place an RFC 3986 parser would disagree. For a single value, URL Encode / Decode handles percent-encoding, and DNS Lookup shows where a host actually points.