ENES
base64Engineering Guide

Base64 Encoding Explained: The 6-Bit Math, Padding Rules, and When Not to Use It (2026)

AS
Published on 2026-10-10ยท12 min readยทDaily Toolbox Engineering

Every developer has encountered a wall of gibberish like iVBORw0KGgoAAAANSUhEUgAA... pasted into source code, a config file, or an API response. That is Base64 โ€” one of the most widely used and most misunderstood encodings in software. It shows up in email attachments, JWT tokens, data URIs, API keys, and countless config files, yet most developers treat it as a black box: bytes go in, ASCII comes out.

This article opens the box. You will learn exactly how the 6-bit grouping math works (with bit-level worked examples), why the = padding exists, how the URL-safe variant differs, where Base64 legitimately belongs in your architecture โ€” and the critical ways it gets misused, including the dangerously common mistake of treating it as encryption.

Why Base64 Exists: Binary Data Meets Text-Only Channels

The problem Base64 solves is old and fundamental: many systems that move data around were designed for text, not binary.

The original motivation was email. SMTP, the protocol that still carries email today, was built for 7-bit ASCII text. If you tried to send a binary file โ€” a photo, an executable, a PDF โ€” raw bytes would pass through mail relays that might strip the high bit, interpret control characters, or mangle line endings. The file would arrive corrupted.

MIME (Multipurpose Internet Mail Extensions, RFC 2045) solved this by defining content transfer encodings: standard ways to represent arbitrary binary data using only safe printable characters. Base64 (RFC 4648) became the workhorse encoding for this job.

The same constraint appears constantly in modern systems:

  • JSON is a text format. It has no native binary type, so APIs that need to return a file, an image, or a cryptographic signature inside JSON must encode the bytes as text first.
  • URLs and filenames can only safely contain a limited character set. Raw binary would break parsing.
  • HTML and CSS are text. Embedding a small image directly in a stylesheet requires a text representation of its bytes.
  • HTTP headers (like Authorization: Basic ...) are text-only.

Base64's answer to all of these: take any sequence of bytes and represent it using exactly 64 safe ASCII characters. The output is slightly larger than the input (about 33% overhead), but it survives every text-only channel intact.

How It Works: The 6-Bit Grouping

Here is the core idea in one sentence: Base64 re-slices a byte stream from 8-bit groups into 6-bit groups, then maps each 6-bit value (0โ€“63) to a printable character.

Why 6 bits? Because 2โถ = 64, and 64 distinct characters fit comfortably in printable ASCII. Each output character carries exactly 6 bits of the original data.

The Alphabet

The standard Base64 alphabet (RFC 4648 ยง4) is:

Value Char Value Char Value Char Value Char
0 A 16 Q 32 g 48 w
1 B 17 R 33 h 49 x
2 C 18 S 34 i 50 y
3 D 19 T 35 j 51 z
4 E 20 U 36 k 52 0
5 F 21 V 37 l 53 1
6 G 22 W 38 m 54 2
7 H 23 X 39 n 55 3
8 I 24 Y 40 o 56 4
9 J 25 Z 41 p 57 5
10 K 26 a 42 q 58 6
11 L 27 b 43 r 59 7
12 M 28 c 44 s 60 8
13 N 29 d 45 t 61 9
14 O 30 e 46 u 62 +
15 P 31 f 47 v 63 /

So Aโ€“Z cover values 0โ€“25, aโ€“z cover 26โ€“51, 0โ€“9 cover 52โ€“61, and + and / cover 62 and 63.

Worked Example: "Man" โ†’ "TWFu"

Let's encode the 3-byte string Man step by step. This is the canonical example from RFC 4648 itself.

Step 1: Convert each byte to 8 bits.

Char Decimal Binary
M 77 01001101
a 97 01100001
n 110 01101110

Step 2: Concatenate into a 24-bit stream.

01001101 01100001 01101110

Step 3: Re-slice into four 6-bit groups.

010011 | 010110 | 000101 | 101110

Step 4: Convert each group to decimal, then look up the character.

6-bit group Decimal Character
010011 19 T
010110 22 W
000101 5 F
101110 46 u

Result: TWFu.

Notice why 3 bytes is the natural input unit: 3 bytes = 24 bits, and 24 is the least common multiple of 8 and 6. Three 8-bit bytes divide evenly into four 6-bit groups with nothing left over. The encoder processes input in 3-byte chunks, emitting 4 characters per chunk.

The 33% Overhead, Explained

Because 3 bytes become 4 characters, Base64 output is always 4/3 the size of the input (for inputs whose length is a multiple of 3). That is a 33.3% size overhead โ€” a hard mathematical floor, not an implementation detail. A 3 MB image becomes a 4 MB string. This matters enormously when deciding where Base64 is appropriate, as we will see in the data URI section.

Padding: Why the = Exists

What happens when the input length is not a multiple of 3? The encoder cannot form a complete final group of 24 bits, so it pads the last group with zero bits โ€” and then marks how many padding characters were added using =.

There are exactly two padding cases:

Input length mod 3 Remaining bytes Output chars Padding
0 3 (full chunk) 4 none
2 2 3 + = =
1 1 2 + == ==

Worked Example: "Ma" โ†’ "TWE="

Two bytes remain: M (01001101) and a (01100001). Concatenated: 16 bits.

010011 | 010110 | 0001

The third group has only 4 bits (0001), so it is padded with two zero bits to make 000100 = 4 = E. We emit three real characters and one =:

TWE=

Worked Example: "M" โ†’ "TQ=="

One byte remains: M (01001101). Just 8 bits.

010011 | 01

The second group has only 2 bits (01), padded with four zero bits to make 010000 = 16 = Q. We emit two real characters and two = signs:

TQ==

Why Padding Matters

Padding serves two purposes. First, it tells the decoder exactly how many real bytes the final quantum contains โ€” without it, a decoder cannot distinguish trailing zero bits from real data. Second, it keeps every Base64 output a multiple of 4 characters, which simplifies streaming decoders and length validation: if a supposedly-Base64 string's length is not divisible by 4, something is wrong.

One caveat: some variants (notably Base64URL as used in JWTs) deliberately omit padding. Decoders for those variants must infer the padding from the string length instead. When you write a decoder, always know which convention your input follows.

Base64URL: The URL-Safe Variant

Standard Base64 has a problem in URLs, filenames, and some protocols: the characters + and / are reserved. In a URL query string, + means "space"; / separates path segments. A standard Base64 string dropped into a URL would be misinterpreted.

RFC 4648 ยง5 defines base64url, which makes two changes:

Position Standard URL-safe
Value 62 + -
Value 63 / _
Padding = required usually omitted

So the same bytes encode slightly differently:

Standard:  a+b/c==
Base64URL: a-b_c      (padding stripped)

You will encounter base64url everywhere tokens travel in URLs: JWTs, OAuth code_challenge values in PKCE flows, and URL-safe identifiers. When converting between the two, remember it is a pure character substitution (plus padding handling) โ€” the underlying bits are identical.

A common bug: feeding a base64url string (with - and _) into a standard Base64 decoder. Most decoders reject it outright; a few silently produce garbage. If your decoder chokes on a JWT segment, check the alphabet first.

Critical: Base64 Is Not Encryption

This is the single most important section of this article, because this is where Base64 causes real security incidents.

Base64 provides zero confidentiality. It is an encoding, like hexadecimal or URL-encoding โ€” a reversible transformation with no key, no secret, and no security properties whatsoever. Anyone who recognizes Base64 (and every developer does) can decode it in under a second, using tools built into every operating system:

# "Decrypting" Base64 takes one command. There is no key.
$ echo "cGFzc3dvcmQxMjM=" | base64 -d
password123

The string cGFzc3dvcmQxMjM= looks opaque. It is the literal text password123. There is no step between "looks secret" and "is public" other than running a decoder.

How This Goes Wrong in Practice

  • "Encrypted" API tokens that are just Base64-encoded JSON. Anyone who decodes them sees user IDs, roles, and permissions in plaintext.
  • Obfuscated credentials in config files or mobile apps. Base64-encoding a hardcoded API key does not protect it from anyone who downloads the app โ€” it only protects it from people who cannot recognize Base64.
  • "Secure" URL parameters carrying Base64-encoded user data. They are readable by every proxy, log, and browser history entry between the user and your server.

The rule is simple: if you would not print it on a billboard, do not Base64-encode it and call it protected. Encoding is not encryption.

What to use instead depends on the actual goal:

Goal Use this, not Base64-alone
Keep data secret in transit TLS (which you get from HTTPS)
Keep data secret at rest Real encryption, e.g. AES-256-GCM with a proper key
Store passwords A password-hashing function (bcrypt, scrypt, or Argon2) โ€” never reversible
Prove data integrity HMAC or a digital signature

Note that Base64 often appears alongside real cryptography โ€” encrypted ciphertext is binary, so it is routinely Base64-encoded for storage in JSON or databases. That is a legitimate use: Base64 is the packaging, AES is the protection. Confusing the two is the error.

Data URIs: Embedding Binary in Text

One of the most visible uses of Base64 is the data URI scheme (RFC 2397), which lets you embed small binary resources directly in HTML or CSS:

data:[<mediatype>][;base64],<data>

A real example โ€” a 1ร—1 transparent PNG embedded in CSS:

.icon {
  background-image: url("data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg==");
}

No separate HTTP request, no extra file to deploy. For a tiny icon, this is genuinely convenient.

When Data URIs Win โ€” and When They Lose

They win for very small assets (icons, tiny patterns): you eliminate an HTTP round trip, which on high-latency connections can matter more than the 33% size penalty. They are also useful for single-file HTML documents (email templates, standalone reports) where external assets are not an option.

They lose as assets grow, for several compounding reasons:

  1. The 33% bloat is real. A 100 KB image becomes ~133 KB of text inside your HTML or CSS.
  2. It blocks parsing. Base64 in your HTML must be downloaded and parsed before the page renders; an <img> tag, by contrast, loads asynchronously.
  3. No separate caching. Change one byte of your HTML and the browser re-downloads the embedded image too. A separate file would have been cached.
  4. No progressive rendering or responsive variants. You cannot srcset a data URI easily, and the image cannot display until fully downloaded.

As a heuristic many frontend engineers use: data URIs for assets under a few kilobytes; real files (ideally modern formats like WebP or AVIF) for everything else. The exact threshold depends on your performance budget โ€” measure, do not guess.

JWT: Base64URL in the Wild

JSON Web Tokens (RFC 7519) are the most prominent real-world consumer of base64url. A JWT has three parts separated by dots:

eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9   โ† header (base64url, no padding)
.
eyJzdWIiOiIxMjM0NTY3ODkwIiwibmFtZSI6IkpvaG4gRG9lIiwiaWF0IjoxNTE2MjM5MDIyfQ   โ† payload
.
SflKxwRJSMeKKF2QT4fwpMeJf36POk6yJV_adQssw5c   โ† signature

Decode the header segment and you get:

{"alg":"HS256","typ":"JWT"}

Two things worth internalizing:

  1. The payload is not encrypted. It is base64url-encoded and signed. Anyone who holds the token can decode the payload with zero keys โ€” try pasting any JWT into a decoder. Never put secrets, passwords, or sensitive PII in a JWT payload. The signature proves the token was not tampered with; it does not hide its contents.

  2. Verify the signature, and verify the algorithm. Early in JWT's history, several libraries accepted tokens with "alg":"none" โ€” meaning "this token is unsigned, trust it anyway" โ€” which allowed trivial token forgery. Modern libraries reject none by default, but the lesson generalizes: when you validate a JWT, pin the expected algorithm explicitly rather than trusting the token's own alg header.

When to Use Base64 โ€” and When Not To

โœ… Legitimate uses โŒ Misuses
Email attachments (MIME Content-Transfer-Encoding: base64) "Encrypting" or "hiding" secrets
Binary data inside JSON/XML APIs Compressing data (it expands data by 33%)
Data URIs for tiny assets (icons) Large images or files in HTML
HTTP Basic auth credentials (Basic dXNlcjpwYXNz) Storing blobs in a database text column (use a binary/BLOB column)
Packaging ciphertext or signatures as text Replacing proper encryption
Embedding fonts in CSS (@font-face with small woff2) Anything where the 33% overhead matters and alternatives exist

A useful mental model: Base64 is a packaging format, not a processing step. It answers "how do I carry these bytes through a text-only channel?" If your problem is not "bytes must survive a text channel," you probably do not need Base64.

Try It

Reading about the 6-bit grouping is one thing; watching it happen is better. DailyToolbox's free Base64 encoder/decoder runs entirely in your browser โ€” paste text, see the Base64 output instantly, flip to decode mode, and try the worked examples from this article (Man โ†’ TWFu, Ma โ†’ TWE=) to verify them yourself. It also handles the URL-safe variant, so you can convert between standard and base64url alphabets without writing a script.

๐Ÿ‘‰ Try the Base64 Encoder/Decoder โ†’

#base64#encoding#data-uri#jwt#web-development#binary-data#mime#rfc-4648
AS
Written by Alex Sun

Alex Sun is the developer behind Daily Toolbox. He writes these guides while building the tools themselves โ€” every claim tested against the real thing.

Try the free tools mentioned above

Try the Base64 Encoder/Decoder โ†’