
What is font subsetting and how does it reduce file size?
- Sajjad
- Typography
- 11 Oct, 2026
Font subsetting is the process of creating a smaller version of a font file that contains only the characters, and optionally the OpenType features, that your site actually uses. Many fonts include glyphs for dozens of languages, symbol sets and alternate forms, while a typical English or Western European site uses a small fraction of them. Removing the rest reduces the file size, often dramatically, because outlines, metrics and kerning data for unused glyphs make up most of a large font's bytes. You can subset fonts with command-line tools such as pyftsubset from fonttools, and serve several subsets efficiently using the unicode-range descriptor so each page only downloads what it needs.
Smaller font files download faster, reach the screen sooner and reduce the chance of a visible font swap. Subsetting is usually the single largest saving you can make on font weight after switching to WOFF2. In this article you'll learn what's inside a font file, how subsetting works, how to choose which characters to keep, how to run pyftsubset and glyphhanger, how to split fonts into multiple subsets with unicode-range, and the licensing and quality issues to check before you ship.
What Takes Up Space in a Font
A font file is made up of tables. The ones that grow with the number of glyphs are:
- Glyph outlines: The
glyforCFFtable, holding the shape of every character. Usually the largest part. - Metrics: The
hmtxtable, with the advance width and side bearings of each glyph. - Character mapping: The
cmaptable, linking Unicode code points to glyphs. - Layout features:
GSUBandGPOS, which hold ligatures, alternates, kerning pairs and mark positioning. Kerning between many glyphs can be surprisingly large. - Hinting: TrueType instructions that adjust outlines at small sizes.
A font that supports Latin, Cyrillic, Greek and Vietnamese might contain well over a thousand glyphs. If your content only uses the basic Latin alphabet, punctuation and a few symbols, most of those glyphs are dead weight.
How Subsetting Works
A subsetting tool takes a list of characters to keep, follows every glyph those characters depend on, and writes a new font without the rest:
- Map characters to glyphs: Look up each requested code point in the
cmaptable. - Follow dependencies: Keep any glyphs reachable through the OpenType features you choose to retain, such as ligatures and small caps, and component glyphs used in composite characters like accented letters.
- Prune tables: Remove unused glyphs from outlines, metrics, kerning and features.
- Rewrite the file: Optionally compress it as WOFF2.
The result is a valid font that renders the kept characters exactly as the original did.
Choosing Which Characters to Keep
This is the main decision. Keep too few and some text falls back to another font; keep too many and you lose the benefit.
Subset by Script
For most sites, subsetting by script or language is the safest approach. Google Fonts splits families into named subsets such as latin, latin-ext, cyrillic, greek and vietnamese. Its latin range is a useful starting point for English and most Western European languages:
U+0000-00FF, U+0131, U+0152-0153, U+02BB-02BC, U+02C6, U+02DA, U+02DC,
U+0304, U+0308, U+0329, U+2000-206F, U+20AC, U+2122, U+2191, U+2193,
U+2212, U+2215, U+FEFF, U+FFFD
That covers ASCII, the Latin-1 Supplement with common accented letters, general punctuation such as curly quotes and dashes, the euro sign, the trade mark sign and a few other common symbols.
Subset by Actual Usage
For a small, mostly static site, or for a display font used only in headings, you can subset to exactly the characters that appear in your content. This produces the smallest files but needs repeating whenever content changes, and it's risky for any font used in user-generated or frequently edited text.
What People Often Forget
- Curly quotes and apostrophes:
’,“and”live in the General Punctuation block, not in ASCII. - En and em dashes: Also in General Punctuation.
- Currency symbols: The pound sign is in Latin-1, but the euro sign is at
U+20AC. - Non-breaking space:
U+00A0, which is common in CMS output. - Accented names: Customer names, place names and borrowed words often need Latin Extended characters such as
ł,őorş. - Form input: If the font is used in input fields, users can type anything.
Subsetting With pyftsubset
pyftsubset is part of fonttools, the standard Python library for working with fonts. Install it with WOFF2 support:
pip install fonttools brotli
A Basic Latin Subset
pyftsubset SourceSans3-Regular.ttf \
--unicodes="U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD" \
--layout-features="*" \
--flavor=woff2 \
--output-file=source-sans-3-latin-400.woff2
The options mean:
- --unicodes: The code points to keep, as single values or ranges.
- --layout-features: Which OpenType features to keep.
"*"keeps all of them, which is safest. By default, pyftsubset keeps a standard set including kerning, standard ligatures and mark positioning, but drops many optional features such as small caps and stylistic sets. - --flavor=woff2: Writes WOFF2 output directly.
- --output-file: The name of the result.
Compare the sizes:
ls -lh SourceSans3-Regular.ttf source-sans-3-latin-400.woff2
The exact numbers depend on the font, but a Latin-only WOFF2 is typically a small fraction of the original TTF.
Keeping Specific Features
If you want a smaller file and know exactly which features you use, list them. This example keeps kerning, ligatures, contextual alternates and tabular figures:
pyftsubset SourceSans3-Regular.ttf \
--unicodes-file=latin-unicodes.txt \
--layout-features="kern,liga,calt,tnum" \
--flavor=woff2 \
--output-file=source-sans-3-latin-400.woff2
The --unicodes-file option reads the ranges from a file, which keeps long commands readable and lets you reuse the same list for every weight.
Subsetting to Specific Text
For a heading-only display font, pass the exact text:
pyftsubset Fraunces-Bold.ttf \
--text="ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789 &.,!?'’-" \
--layout-features="kern,liga" \
--flavor=woff2 \
--output-file=fraunces-bold-caps.woff2
Use this approach only when you control every string the font will render.
Optional Extra Savings
- --no-hinting: Removes TrueType hinting. Modern high-resolution screens and macOS ignore most hinting, but Windows at small sizes can look noticeably worse without it. Test before using.
- --desubroutinize: For CFF fonts, can improve WOFF2 compression in some cases.
- --name-IDs and --drop-tables: Strip metadata you don't need. Keep the licence and copyright entries.
Batch Processing All Weights
A shell loop applies the same subset to every file in a folder:
mkdir -p dist
for f in fonts/*.ttf; do
name=$(basename "$f" .ttf | tr '[:upper:]' '[:lower:]')
pyftsubset "$f" \
--unicodes-file=latin-unicodes.txt \
--layout-features="*" \
--flavor=woff2 \
--output-file="dist/${name}-latin.woff2"
done
ls -lh dist
Finding the Characters You Use With glyphhanger
glyphhanger, a Node.js tool by Zach Leatherman, crawls pages in a headless browser and reports the Unicode ranges actually used. It can also call pyftsubset for you. It needs fonttools installed:
npm install -g glyphhanger
glyphhanger https://www.example.com/ --spider --spider-limit=20
U+20-7E,U+A9,U+2013,U+2014,U+2019,U+201C,U+201D
The output is a list of ranges you can add to your subset. To subset a font in the same step:
glyphhanger https://www.example.com/ --spider \
--subset=fonts/SourceSans3-Regular.ttf \
--formats=woff2
Treat crawled ranges as a minimum. Add the full Latin set or a safety margin for content that may change.
Serving Multiple Subsets With unicode-range
Instead of choosing a single subset, you can split a font into several files and let the browser decide which to download. Declare one @font-face rule per subset, all with the same family name, and give each a unicode-range:
/* Latin */
@font-face {
font-family: "Source Sans 3";
src: url("/fonts/source-sans-3-latin-400.woff2") format("woff2");
font-weight: 400;
font-style: normal;
font-display: swap;
unicode-range: U+0000-00FF, U+0131, U+0152-0153, U+02BB-02BC, U+02C6,
U+02DA, U+02DC, U+0304, U+0308, U+0329, U+2000-206F, U+20AC, U+2122,
U+2191, U+2193, U+2212, U+2215, U+FEFF, U+FFFD;
}
/* Cyrillic */
@font-face {
font-family: "Source Sans 3";
src: url("/fonts/source-sans-3-cyrillic-400.woff2") format("woff2");
font-weight: 400;
font-style: normal;
font-display: swap;
unicode-range: U+0301, U+0400-045F, U+0490-0491, U+04B0-04B1, U+2116;
}
The browser checks the page's text against each range and only downloads the files whose characters actually appear. An English page downloads the Latin file; a Russian page downloads both, because it still contains Latin punctuation and digits.
There's one trade-off. If a page uses characters from several subsets, the browser makes several requests, and glyphs in different files can't kern against each other. Keep related characters, like a language's letters and its punctuation, in the same subset.
Subsetting Variable Fonts
Variable fonts can be subset with the same pyftsubset commands, and the result is still variable. You can also reduce their size by limiting the design space. fonttools' instancer can pin an axis or narrow its range:
fonttools varLib.instancer SourceSans3-VariableFont_wght.ttf \
wght=400:700 \
-o SourceSans3-wght400-700.ttf
Then subset the result as usual. Narrowing the weight range from 200 to 900 down to 400 to 700 removes variation data you don't use.
Checking the Result
After subsetting, confirm nothing broke:
- Render your real content: Look at pages with accented names, quotes, prices and special symbols.
- Check Rendered Fonts: In DevTools, inspect text and look for any characters drawn by a fallback font, shown as a separate entry under Rendered Fonts.
- List what the subset contains: Use fonttools to print the character map.
python3 -c "
from fontTools.ttLib import TTFont
font = TTFont('source-sans-3-latin-400.woff2')
cmap = font.getBestCmap()
print(len(cmap), 'characters')
print(''.join(chr(c) for c in sorted(cmap) if 0x20 < c < 0x17F))
"
Licensing Considerations
Subsetting creates a modified copy of the font, so check the licence first:
- Commercial licences: Many web licences allow subsetting for optimisation; some don't. Read the EULA or ask the foundry.
- SIL Open Font License: Permits modification. The licence includes a Reserved Font Name provision for some fonts, and the OFL FAQ discusses how subsetting for web delivery relates to it. Read the FAQ if your font declares a reserved name.
- Keep licence metadata: Don't strip copyright or licence entries from the
nametable.
FAQ: Font Subsetting
It depends on how many glyphs the original contains. A font covering many scripts can shrink to a small fraction of its size when cut down to Latin only, while a font that already only covers Latin will see a smaller saving.
Yes. The Google Fonts API serves families split into script-based subsets with unicode-range rules, so browsers only download the subsets a page needs. If you self-host, you need to do this yourself or use a package that does.
The browser draws that character with the next font in your font-family list, so it appears in a different typeface. That's why it's safer to subset by script than by the exact characters on a page.
Only if you remove the features. pyftsubset keeps kerning and standard ligatures by default, and passing --layout-features with an asterisk keeps every feature for the glyphs you retain.
Yes. pyftsubset preserves variation data for the kept glyphs, and fonttools' varLib.instancer can narrow or pin axes to remove variation data you don't use.
It depends on the licence. Open-source licences such as the SIL OFL allow modification, and many commercial web licences allow subsetting, but some don't. Check the EULA before you modify a commercial font.
Conclusion
Font subsetting removes the glyphs, metrics and features your site never uses, which is why it's one of the most effective ways to cut font weight. Most of a large font's bytes belong to scripts and symbols a typical page doesn't need, and tools like pyftsubset can strip them out and write WOFF2 in one command.
Subset by script for text that changes, by exact characters only for fixed headings and logos, and use unicode-range when you need to support several languages without making everyone download all of them. Check your real content afterwards, keep licence data intact, and confirm the licence permits modification. Done carefully, subsetting gives you much lighter font files without any visible difference.


