A MIME type, or media type in current standards, is the label that says what format a sequence of bytes is in: text/html, image/png, application/json. HTTP sends it in the Content-Type header so the browser knows whether to render a page, decode an image, run a script or save a file. Email uses the same labels for attachments, which is where the name came from: MIME is Multipurpose Internet Mail Extensions.
This guide covers how a type is written, who registers types, what two popular web servers send by default, how Chrome reacts when the type is wrong, and how to tell a file’s real format from its bytes. The server and browser results were measured on 2 October 2026 with Nginx 1.30.5 and Apache httpd 2.4.68 in Docker and Chrome 152 on macOS. The lookup and sniffing examples come from ZeroTool’s MIME Type Lookup, and this site’s tests recompute them from the tool’s code.
How a MIME type is written
RFC 9110 §8.3.1 gives the HTTP grammar:
media-type = type "/" subtype parameters
The type and subtype are case-insensitive. Parameters follow after semicolons as name=value pairs; parameter names are case-insensitive, and whether a value is depends on the parameter. RFC 9110 lists four equivalent ways to say “HTML in UTF-8” and prefers the first:
text/html;charset=utf-8
Text/HTML;Charset="utf-8"
text/html; charset="utf-8"
text/html;charset=UTF-8
RFC 6838 §4.2 restricts registered names further: letters, digits and ! # $ & - ^ _ . +, starting with a letter or digit, up to 127 characters, and it asks for 64 or fewer. Two characters carry meaning:
- The part before the first dot is a tree. No dot is the standards tree (
application/json).vnd.is the vendor tree (application/vnd.ms-excel),prs.is the personal tree, andx.is the unregistered tree, which “cannot be registered” and is “strongly discouraged”. Names starting withx-(application/x-7z-compressed) are not in thex.tree: RFC 6838 §3.4 says they “are no longer considered to be members of this tree”, and widely used ones may be registered under their current name. - The part after the last plus is a structured syntax suffix.
application/ld+jsonandimage/svg+xmltell a generic parser the bytes are JSON or XML underneath.
Common parameters are charset for text, boundary for multipart/form-data, and profile or version for some JSON-based types.
The registry and its top-level types
IANA’s Media Types registry, last updated 24 September 2026 when downloaded for this guide, has eleven top-level types. The counts are the rows in each registry’s CSV:
| Top-level type | Registered subtypes | Notes |
|---|---|---|
application | 1,807 | 22 marked obsolete or deprecated |
audio | 165 | |
example | 0 | Reserved for documentation by RFC 4735; no subtype can be registered |
font | 6 | Added by RFC 8081 in 2017: collection, otf, sfnt, ttf, woff, woff2 |
haptics | 3 | Added by RFC 9695 |
image | 88 | |
message | 27 | |
model | 44 | 3D formats such as model/gltf+json |
multipart | 17 | |
text | 107 | |
video | 97 |
Registered and common are different lists. The MIME Type Lookup tool has 155 entries chosen for web work; checked against the IANA files, 106 of them are registered and 49 are not. The unregistered ones include several that every web developer uses:
video/webmandaudio/webm.audio/wavandaudio/x-wav; the audio registry has no WAV entry at all.image/x-icon; the registered form isimage/vnd.microsoft.icon.text/yaml; the registered type isapplication/yaml(RFC 9512).application/x-7z-compressed,video/x-matroskaand thetext/x-*source-code types.
The tool also keeps two registered types that IANA marks obsolete: application/javascript and application/ecmascript. RFC 9239 (2022) made text/javascript the one JavaScript type and moved the others to “OBSOLETED in favor of text/javascript”. Searching .js in the tool returns both application/javascript and text/javascript because both appear on real servers.
A few extensions map to more than one type in the tool’s table, and those are the ones that cause bugs: .ts is MPEG transport stream video (video/mp2t) and TypeScript source (text/x-typescript); .xml is application/xml and text/xml; .ico is image/x-icon and image/vnd.microsoft.icon.
What Nginx and Apache send by default
Each server maps file extensions to types with a table. These are the Content-Type headers the official Docker images (nginx:stable-alpine, httpd:2.4-alpine) sent for files with each extension, with no configuration changes:
| Extension | Nginx 1.30.5 | Apache httpd 2.4.68 |
|---|---|---|
.js | application/javascript | text/javascript |
.mjs | application/octet-stream | text/javascript |
.json | application/json | application/json |
.wasm | application/wasm | application/wasm |
.svg | image/svg+xml | image/svg+xml |
.webp | image/webp | image/webp |
.avif | image/avif | image/avif |
.heic | application/octet-stream | image/heic |
.jxl | application/octet-stream | image/jxl |
.woff2 | font/woff2 | font/woff2 |
.csv | application/octet-stream | text/csv |
.xml | text/xml | application/xml |
.m4a | audio/x-m4a | audio/mp4 |
.ts | video/mp2t | video/mp2t |
.webmanifest | application/octet-stream | no Content-Type header |
.md, .yaml, .toml, .map | application/octet-stream | no Content-Type header |
| unknown extension | application/octet-stream | no Content-Type header |
Nginx falls back to default_type application/octet-stream from its nginx.conf. Apache 2.4 has no default type, so a file it does not recognize goes out with no Content-Type at all and the browser has to guess.
The .mjs row matters most. A page on the Nginx container loading <script type="module" src="/mod.mjs"> never ran the module, and Chrome logged:
Failed to load module script: Expected a JavaScript-or-Wasm module script but the server responded with a MIME type of "application/octet-stream". Strict MIME type checking is enforced for module scripts per HTML spec.
The same page on the Apache container ran it. To fix Nginx, add the types after the stock table. Putting types { … } on its own inside a server block looks right but replaces the whole inherited table: in the test, .svg and .json were then served as application/octet-stream. This version keeps the rest of the table:
server {
include /etc/nginx/mime.types;
types {
text/javascript js mjs;
}
}
Nginx 1.30.5 warns duplicate extension "js", content type: "text/javascript", previous content type: "application/javascript" on start and then uses the later mapping. On Apache, AddType text/javascript .mjs does the same job for servers whose table lacks it.
How Chrome treats a wrong type
Browsers do not trust the extension in the URL, but they do look at the type. Two rules from the Fetch standard decide most cases:
- A script request is blocked if the response type starts with
audio/,image/orvideo/, or istext/csv(Fetch §2.10). - If the response also has
X-Content-Type-Options: nosniff, a script is blocked unless the type is a JavaScript MIME type, and a stylesheet is blocked unless the type istext/css(Fetch §3.6.1).
A test page loaded the same one-line script under six types. Chrome 152 ran two of them:
| Response headers for the script | Result |
|---|---|
Content-Type: text/plain | Ran |
Content-Type: application/json | Ran |
Content-Type: text/plain + nosniff | Blocked: “MIME type (‘text/plain’) is not executable, and strict MIME type checking is enabled” |
Content-Type: application/octet-stream + nosniff | Blocked, same message |
Content-Type: image/png | Blocked: “MIME type (‘image/png’) is not executable” |
Content-Type: text/csv | Blocked: “MIME type (‘text/csv’) is not executable” |
A stylesheet served as text/plain from a page in standards mode was not applied either. nosniff is worth sending on every response: it turns a mislabelled script from a silent success into an error you notice, and stops a browser from running an uploaded text file as a script.
Reading a file’s real type from its bytes
Most binary formats begin with a fixed byte sequence, often called magic bytes: PNG starts with 89 50 4E 47 0D 0A 1A 0A, PDF with %PDF, ZIP with PK\x03\x04. The WHATWG MIME Sniffing standard defines the patterns browsers use, and file(1) on Unix uses a much larger database.
The tool’s Sniff file tab reads the first 64 bytes of a file, tries 37 signatures in order, and reports the first match. If none matches it shows the browser’s File.type, or application/octet-stream when the browser has none. These eleven files were dropped onto it in Chrome 152 and also run through file --mime-type (file 5.41, macOS):
| File | First bytes | Tool result | Chrome File.type | file 5.41 |
|---|---|---|---|---|
photo.jpg (a PNG renamed) | 89 50 4E 47 0D 0A 1A 0A | image/png | image/jpeg | image/png |
report.docx | 50 4B 03 04 0A 00 00 00 | application/zip + ZIP container note | application/vnd.openxmlformats-officedocument.wordprocessingml.document | same as Chrome |
book.epub | 50 4B 03 04 14 00 00 00 | application/zip + ZIP container note | application/epub+zip | application/epub+zip |
pic.avif | 00 00 00 1C 66 74 79 70 61 76 69 66 | image/avif | image/avif | image/avif |
tone.mp3 (no ID3 tag) | FF FB 50 C4 | audio/mpeg | audio/mpeg | audio/mpeg |
voice.m4a | 00 00 00 1C 66 74 79 70 4D 34 41 20 | video/mp4 | audio/x-m4a | audio/x-m4a |
ls-binary (macOS /bin/ls) | CA FE BA BE 00 00 00 03 | application/java-vm | none | application/x-mach-binary |
parts.csv | 42 4D 57 20 (BMW ) | image/bmp | text/csv | text/plain |
logo.svg | 3C 73 76 67 (<svg) | image/svg+xml (no signature; from File.type) | image/svg+xml | image/svg+xml |
user.json | 7B 22 69 64 ({"id) | application/json (no signature; from File.type) | application/json | application/json |
app.ts | 63 6F 6E 73 (cons) | application/octet-stream (no signature, no File.type) | none | text/plain |
What the table shows:
- The extension and the bytes disagree, and the bytes win. Chrome derives
File.typefrom the file name, sophoto.jpgreportsimage/jpegeven though it is a PNG. This is why an upload handler must not copy the browser’s type into storage. - ZIP-based formats look identical at the start. DOCX, XLSX, EPUB, JAR and APK all begin with a ZIP local file header.
filetells DOCX and EPUB apart by reading further into the archive; the tool stops at 64 bytes and says the file may be one of several ZIP-based formats. - Short signatures produce false matches.
CA FE BA BEis the Java class file magic and also the header of a macOS universal binary, so/bin/lscomes out asapplication/java-vm. A CSV whose first line starts withBMmatches the two-byte BMP signature.M4Ais an MP4 brand for audio, and the tool reports every MP4-family brand it knows asvideo/mp4. - Text formats have no signature. JSON, CSV, SVG and source code can only be identified by parsing. Chrome gives no type for
.ts, so the tool falls back toapplication/octet-stream.
Formats whose marker is further in, such as tar (ustar at byte 257) and ISO 9660 images (CD001 at byte 32769), are outside the 64-byte window. Use file for those.
Checking uploads on the server
A safe upload check decides the type from the bytes, compares it with the extension, and ignores the Content-Type the client sent. This Node module allows four formats and rejects everything else, including a PNG named .jpg:
// check-upload.mjs
import { open } from 'node:fs/promises';
const ALLOWED = [
{ mime: 'image/png', exts: ['.png'], test: (b) => b.subarray(0, 8).equals(Buffer.from('89504e470d0a1a0a', 'hex')) },
{ mime: 'image/jpeg', exts: ['.jpg', '.jpeg'], test: (b) => b[0] === 0xff && b[1] === 0xd8 && b[2] === 0xff },
{ mime: 'image/webp', exts: ['.webp'], test: (b) => b.toString('latin1', 0, 4) === 'RIFF' && b.toString('latin1', 8, 12) === 'WEBP' },
{ mime: 'application/pdf', exts: ['.pdf'], test: (b) => b.toString('latin1', 0, 5) === '%PDF-' },
];
export async function checkUpload(path, originalName) {
const file = await open(path);
const head = Buffer.alloc(64);
const { bytesRead } = await file.read(head, 0, 64, 0);
await file.close();
const bytes = head.subarray(0, bytesRead);
const match = ALLOWED.find((t) => t.test(bytes));
if (!match) return { ok: false, reason: 'file type not allowed' };
const dot = originalName.lastIndexOf('.');
const ext = dot < 0 ? '' : originalName.slice(dot).toLowerCase();
if (!match.exts.includes(ext)) return { ok: false, reason: `content is ${match.mime} but the name ends in "${ext}"` };
return { ok: true, mime: match.mime };
}
For a wider set of formats, libraries such as file-type for Node or python-magic (a wrapper around libmagic, the library behind file) read more of the file and know far more signatures. Whatever detects the type, store the detected value and send it back as the Content-Type when serving the file, together with X-Content-Type-Options: nosniff.
Detecting the format does not make a file safe. image/svg+xml is XML and can contain <script>; a correctly identified PDF can still carry a malicious payload. Type checks decide which handler gets the file, not whether its content is trustworthy.
charset and the Excel BOM are separate problems
The charset parameter tells a browser how to decode text it receives: Content-Type: text/csv; charset=utf-8 makes the browser read the response as UTF-8. It applies to that HTTP response and nothing else.
When the user downloads the CSV and opens it in Excel, the HTTP header is gone; only the file is saved. Microsoft’s support article Opening CSV UTF-8 files correctly in Excel explains that a UTF-8 CSV with a byte order mark opens correctly by double-click, and gives an import route for files without one. So a CSV export for Excel users needs both: charset=utf-8 in the header for the browser, and the bytes EF BB BF at the start of the file for Excel.
When application/octet-stream is the right answer
application/octet-stream means arbitrary binary data. RFC 2046 §4.5.1 says the recommended action for it is to offer to put the data in a file, which is what browsers do. That makes it correct for encrypted blobs, firmware images and other downloads with no more specific type.
It is the wrong default for files whose type you know. An image or PDF stored as application/octet-stream is downloaded instead of displayed. If you want a known type to download rather than open, send the real type with Content-Disposition: attachment; filename="report.pdf" (RFC 6266) instead of hiding the type.
Sources
- RFC 6838: media type specifications and registration procedures
- RFC 9110 §8.3: Content-Type and the media type grammar
- IANA Media Types registry (updated 2026-09-24)
- RFC 9239: updates to ECMAScript media types
- Fetch standard §2.10 and §3.6.1; MIME Sniffing standard
- Nginx
typesdirective - RFC 2046 §4.5.1, RFC 6266
Related tools: File Hash Checker to verify a download once you know what it is, Image to Base64, which also identifies images from their bytes before building a data URI, and HTTP Status Codes for the rest of the response.