A common way to lose a table of contents: the README renders fine on GitHub, then the same file is viewed on Bitbucket or built by a Jekyll site with a different slug rule, and the TOC links stop jumping anywhere. The Markdown is identical. The heading ids are not.
Markdown table-of-contents generation looks like a one-liner — read the headings, slugify them, emit links — and the first 80% of the work really is. The trap is the last 20%: every renderer has its own slugify rules, its own duplicate handling, and its own opinion on Unicode. A TOC built for the wrong target is worse than no TOC at all, because the links look correct in your editor and break the moment they hit the renderer.
Why anchors aren’t standardized
The Markdown spec doesn’t talk about heading anchors. CommonMark deliberately left HTML rendering of headings to the implementation, and every renderer that ships a “slug from heading” feature wrote its own algorithm. The dialects diverged enough that you can’t write a universal TOC generator without picking a target.
The four dialects this tool offers:
| Renderer | Where you see it | Anchor style |
|---|---|---|
| GitHub | github.com READMEs, issues and wikis | lowercase, remove punctuation except - and _, spaces to hyphens, keep Unicode letters |
| GitLab | gitlab.com, GitLab self-managed | same as GitHub, then two or more hyphens collapse into one |
kramdown (input: kramdown) | Jekyll sites configured with kramdown’s own parser | drop leading non-letters, keep only ASCII letters, digits, spaces and hyphens, lowercase; an empty result becomes section |
| Bitbucket Cloud | bitbucket.org | markdown-header- prefix on every id, _N for duplicates |
Two notes on the table. Jekyll’s default configuration uses the GFM parser (kramdown: input: GFM), whose header ids follow the GitHub rules and keep Unicode (kramdown-parser-gfm); the ASCII-only row applies only when a site switches to input: kramdown. And the tool’s Bitbucket option currently writes a markdown- prefix, while Bitbucket Cloud renders ids as markdown-header-...; add header- after markdown- in the generated links until the tool is fixed. Hugo (goldmark, the default renderer since v0.60) and MkDocs have their own id rules; check the id in the rendered page before committing a TOC for a target not listed here.
The slugify algorithm, four ways
Walk through a single heading and watch the four dialects produce different anchors:
Heading: ## Quick Start: Setting up SSO (Auth 2.0)
| Style | Anchor |
|---|---|
| GitHub | quick-start-setting-up-sso-auth-20 |
| GitLab | quick-start-setting-up-sso-auth-20 |
| Jekyll | quick-start-setting-up-sso-auth-20 |
| Bitbucket Cloud | markdown-header-quick-start-setting-up-sso-auth-20 |
So far so consistent. Now try a Unicode heading: ## 快速开始
| Style | Anchor |
|---|---|
| GitHub | 快速开始 |
| GitLab | 快速开始 |
kramdown (input: kramdown) | section (nothing ASCII is left) |
| Bitbucket Cloud | not shown here: its Python-Markdown slugger removes non-ASCII characters, so check the rendered id |
This is where kramdown’s own parser falls over: every heading without ASCII letters becomes section, section-1, and so on (kramdown auto_ids). The tool’s Jekyll option leaves such anchors empty instead, so the link goes nowhere either way. If your Jekyll site has non-ASCII headings, use the default input: GFM and generate the TOC with the GitHub style.
One more case — repeated headings:
## Examples
### Curl
## Examples
### Python
| Style | Anchors |
|---|---|
| GitHub / GitLab / kramdown | examples, curl, examples-1, python |
| Bitbucket Cloud | markdown-header-examples, markdown-header-curl, markdown-header-examples_1, markdown-header-python |
Bitbucket is the odd one out: it uses an underscore for the duplicate counter instead of a hyphen. Every other major renderer uses -N.
Marker mode: keep the TOC fresh without git noise
A common workflow trap with TOCs is the “stale TOC” diff: you add a section to a 2,000-line operations runbook, the TOC at the top drifts out of sync, and reviewers spend two PRs noticing. There are two ways out:
- Generator-managed comment markers. Wrap the TOC region with
<!-- toc -->…<!-- /toc -->and re-run the generator on every change. The markers stay in the file; only the body between them moves. - Pre-commit hook. Same as above, run automatically before commit.
The marker mode in this tool implements pattern 1. Paste the document, enable Marker mode, and you get the full Markdown back with a regenerated TOC inside the markers. If your document doesn’t have markers yet, the tool inserts a fresh block right before the first heading so you can commit it once and use markers from then on.
<!-- toc --> / <!-- /toc --> is the marker pair used by the npm package markdown-toc, so a file prepared here can later be maintained with markdown-toc -i. HTML comments are not displayed by GitHub or GitLab. Bitbucket Cloud escapes raw HTML, so the markers would show up as text there.
Five traps you’ll hit eventually
A list of failure modes that actually happen in real projects, ordered by how often they bite:
-
Code blocks with
#comments. Bash, Ruby, and Python comments start with#. A naive heading parser reads# This is a commentinside a fenced code block as an H1. Any TOC generator worth using skips fenced code blocks. The tool here handles```and~~~fences correctly; if you’re rolling your own, watch this case. -
Setext headings. Markdown supports two heading styles: ATX (
# Title) and setext (Title\n=====). Older READMEs still use setext for H1 and H2. A generator that only handles ATX silently skips them. -
Inline formatting in headings.
## **Important**: Backupsshould produce a TOC entry whose label includes the bold but whose anchor uses the plain text. For example,## \code` exampleshould link to#code-example`, not to an anchor containing backticks. Strip inline syntax for the slug; preserve it for the label. -
Repeated headings across H levels.
## Examplesand### Examplesboth slug toexamples, thenexamples-1for the second one. Not all generators get the cross-level dedupe right. Renderers count every heading in the document, including ones you leave out of the TOC; this tool counts only the headings inside the selected level range, so if you exclude H1 and the document has# Examplesand## Examples, the tool links to#exampleswhile GitHub assignsexamples-1to the H2. -
The marker-already-present case. If a document already has
<!-- toc -->…<!-- /toc -->, naive in-place insertion adds a second TOC. The right behavior is to detect the existing markers and replace what’s between them, not append.
Code recipes
Pre-commit hook: regenerate the TOC on commit
The cleanest way to keep TOCs in sync is to fail the commit if the TOC is stale, then have a one-key fix. With markdown-toc (npm) and a pre-commit hook (the script uses GNU md5sum; on macOS use md5 -q instead):
#!/bin/sh
# .git/hooks/pre-commit
for f in $(git diff --cached --name-only --diff-filter=AM | grep '\.md$'); do
before=$(md5sum "$f")
npx markdown-toc -i "$f"
after=$(md5sum "$f")
if [ "$before" != "$after" ]; then
echo "TOC out of date: $f — staged the regenerated version."
git add "$f"
fi
done
Replace npx markdown-toc -i with whatever generator you prefer; the contract is that it edits the file in place and reads the marker block.
GitHub Actions: TOC drift check on PR
name: TOC drift
on: [pull_request]
jobs:
check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: npx markdown-toc -i README.md
- run: |
if ! git diff --quiet README.md; then
echo "::error::README.md TOC is out of date. Run 'npx markdown-toc -i README.md' locally and commit."
exit 1
fi
Jekyll: make sure your slugger matches your TOC
If your Jekyll site has non-ASCII headings, keep (or restore) the GFM parser, which is Jekyll’s default:
# _config.yml
kramdown:
input: GFM
Then generate your TOC with the GitHub style — it will match the resulting anchors.
How this tool fits
The behaviors that matter when picking a TOC generator, and where this one lands:
| Behavior | This tool |
|---|---|
Skips fenced code blocks (``` and ~~~) | yes |
| ATX and setext headings | both |
Replaces existing <!-- toc --> block in place | yes |
| Four anchor styles in a single page | GitHub / GitLab / Jekyll (kramdown parser) / Bitbucket (prefix needs header- added, see above) |
Bitbucket _N dedupe | yes |
| Dedupe counts headings outside the level filter | no (see trap 4) |
| Strips inline Markdown for the anchor, keeps it for the label | yes |
| Heading-level filter and H1 toggle | yes |
| Runs in your browser | yes |
If your workflow is server-side and you want a CLI, markdown-toc (npm, GitHub-style) uses the same markers. For ad-hoc paste-and-copy on the web, the four-style support and marker-aware replacement here saves a round trip when your target isn’t GitHub.
Further reading
- GitHub Flavored Markdown spec — the spec doesn’t define anchor slugs (those come from the GitHub renderer’s behavior), but it pins down the heading and code-block parsing the slug rules build on
- kramdown auto_ids — Jekyll’s slug rules in their canonical form
- GitLab Flavored Markdown reference — the GitLab additions on top of CommonMark
- Python-Markdown toc extension — the slugify and
_Ndedupe behavior Bitbucket Cloud’s ids follow - kramdown-parser-gfm — the GFM parser Jekyll uses by default and its GitHub-style header ids
- ZeroTool Slugify, Markdown Linter, Markdown Table Generator — sister tools for the rest of your Markdown workflow