String Byte Size Calculator
Character count, UTF-16 code unit length, and UTF-8 byte size, all at once.
Every calculation runs locally in your browser. Nothing you type here is sent to a server.
For informational and educational purposes only — not professional or technical advice, and not a substitute for consulting a qualified professional about your specific situation. TrueMeasureKit is not liable for decisions made based on these results. See our Terms of Service.
Three numbers, three different questions
"Characters" counts what a person actually sees — an emoji is one character even if it's technically built from multiple Unicode code points. JavaScript's built-in string.length counts UTF-16 code units instead, which is why an emoji often reports as length 2 in JS — it silently over-counts, a common source of off-by-one bugs when truncating strings for a database column or a UI limit. UTF-8 bytes is what actually matters for storage and network payload size, and it can differ from both — a single CJK character is 3 bytes in UTF-8 despite being 1 character and 1 UTF-16 code unit.
Frequently asked questions
Why is UTF-8 byte size different from character count?
UTF-8 uses a variable number of bytes per character — plain ASCII characters take 1 byte, but many accented letters, symbols, and emoji take 2-4 bytes each, so text with non-ASCII characters is larger in bytes than its character count suggests.
What's the difference between character count and UTF-16 code units?
Most characters are one UTF-16 code unit, but characters outside the Basic Multilingual Plane (many emoji, for instance) are represented as a surrogate pair — two code units — so JavaScript's string.length (UTF-16 based) can exceed the actual character count.