How to Convert Text to Binary

Turning text into binary isn't one step — it's a three-step pipeline: characters become numbers, numbers become bytes, and bytes become bits. Understand this pipeline and you'll understand how every text file, message, and web page is actually stored. Let's walk through it with real examples, including emoji.

Step 1: characters become numbers (code points)

Every character has a code point — its ID number in the Unicode standard. For the basic Latin letters these match the old ASCII codes: "H" is 72, "i" is 105, "A" is 65. Emoji live much higher: 😀 is code point 128512. The code point is still just a number — not yet binary.

Step 2: numbers become bytes (UTF-8)

UTF-8 is the encoding that turns code points into bytes, and it's the standard for virtually all text on the internet. Its rules:

So "H" (72) becomes the single byte 72, while 😀 (128512) becomes four bytes: 240, 159, 152, 128. A middle case: é (code point 233) encodes as two bytes, 195 and 169 — binary 11000011 10101001. The leading byte starts with 110 (UTF-8's marker for "two-byte character ahead") and the continuation byte starts with 10, the same signature you'll see in the 😀 table below.

Step 3: bytes become 8-bit binary groups

Each byte (0–255) is written as exactly 8 binary digits, padding with leading zeros. That's the binary you see.

Worked example: "Hi" to binary

CharacterCode pointUTF-8 bytesBinary
H727201001000
i10510501101001

"Hi" in binary is 01001000 01101001. Two characters, two bytes, sixteen bits.

Worked example: 😀 to binary

ByteDecimalHexBinary
1240F011110000
21599F10011111
31529810011000
41288010000000

So 😀 in binary is 11110000 10011111 10011000 10000000 — 32 bits for a single character. Notice bytes 2–4 all start with 10: that's UTF-8's signature marking them as continuation bytes, which is how decoders know where each character's bytes end.

Why "8 bits per character" is a myth

Tutorials often say "each character becomes 8 bits." That's only true for ASCII text. The accurate statement is each byte becomes 8 bits — and a character can be 1 to 4 bytes in UTF-8. The word "café" is 5 bytes (the é takes 2), and "Hi 👋" is 7 bytes. Any tool that splits text into fixed 8-bit chunks per character will corrupt emoji — which is why our text to binary converter works byte-by-byte instead.

What about other encodings?

UTF-8 won, but you'll still meet alternatives. UTF-16 uses a minimum of 2 bytes per character — it's what JavaScript strings and Windows internals use, which is why "😀".length is 2 in JavaScript. Latin-1 (ISO-8859-1) maps bytes 0–255 one-to-one onto the first 256 Unicode code points, which is why old Western-European text "just works" as raw bytes. The golden rule: bytes are meaningless without knowing the encoding. Byte 233 is é in Latin-1 but the start of a two-byte sequence in UTF-8 — mislabeled text is exactly what produces garbled mojibake.

A concrete consequence: the string "café" looks like 4 characters but encodes as 5 UTF-8 bytes — c(1) a(1) f(1) é(2). In Python, len("café") is 4 while len("café".encode("utf-8")) is 5, a classic source of bugs when allocating buffers or validating input lengths.

Going the other way

Decoding reverses the pipeline: split the bits into 8-bit groups, convert each group to its byte value, then decode the byte sequence as UTF-8. Paste 01001000 01101001 into our binary to text converter and you'll get "Hi" back. For a more compact text encoding often used in URLs and APIs, see Base64, which represents binary data using 64 printable characters.

Convert any text to binary instantly:

Open the Text to Binary Converter

Frequently asked questions

How do I convert text to binary?

Look up each character's UTF-8 bytes, then write each byte as 8 binary digits. For example, 'H' is byte 72, which is 01001000 in binary.

How many bits is one character in binary?

It depends on the character. ASCII letters and digits take 8 bits (1 byte) in UTF-8, but characters like é take 16 bits and emoji take 32 bits.

Why does emoji take 32 bits in binary?

Emoji have large Unicode code points that don't fit in one byte, so UTF-8 encodes them as 4 bytes — 32 bits total. For example, 😀 is 11110000 10011111 10011000 10000000.

Can binary be converted back to text?

Yes — group the bits into bytes, convert each byte to its decimal value, and decode the bytes as UTF-8. Our binary to text converter does this automatically.

More converters

Latest from the blog