Every developer has encountered a wall of gibberish like iVBORw0KGgoAAAANSUhEUgAA... pasted into source code, a config file, or an API response. That is Base64 โ one of the most widely used and most misunderstood encodings in software. It shows up in email attachments, JWT tokens, data URIs, API keys, and countless config files, yet most developers treat it as a black box: bytes go in, ASCII comes out.
This article opens the box. You will learn exactly how the 6-bit grouping math works (with bit-level worked examples), why the = padding exists, how the URL-safe variant differs, where Base64 legitimately belongs in your architecture โ and the critical ways it gets misused, including the dangerously common mistake of treating it as encryption.
Why Base64 Exists: Binary Data Meets Text-Only Channels
The problem Base64 solves is old and fundamental: many systems that move data around were designed for text, not binary.
The original motivation was email. SMTP, the protocol that still carries email today, was built for 7-bit ASCII text. If you tried to send a binary file โ a photo, an executable, a PDF โ raw bytes would pass through mail relays that might strip the high bit, interpret control characters, or mangle line endings. The file would arrive corrupted.
MIME (Multipurpose Internet Mail Extensions, RFC 2045) solved this by defining content transfer encodings: standard ways to represent arbitrary binary data using only safe printable characters. Base64 (RFC 4648) became the workhorse encoding for this job.
The same constraint appears constantly in modern systems:
- JSON is a text format. It has no native binary type, so APIs that need to return a file, an image, or a cryptographic signature inside JSON must encode the bytes as text first.
- URLs and filenames can only safely contain a limited character set. Raw binary would break parsing.
- HTML and CSS are text. Embedding a small image directly in a stylesheet requires a text representation of its bytes.
- HTTP headers (like
Authorization: Basic ...) are text-only.
Base64's answer to all of these: take any sequence of bytes and represent it using exactly 64 safe ASCII characters. The output is slightly larger than the input (about 33% overhead), but it survives every text-only channel intact.
How It Works: The 6-Bit Grouping
Here is the core idea in one sentence: Base64 re-slices a byte stream from 8-bit groups into 6-bit groups, then maps each 6-bit value (0โ63) to a printable character.
Why 6 bits? Because 2โถ = 64, and 64 distinct characters fit comfortably in printable ASCII. Each output character carries exactly 6 bits of the original data.
The Alphabet
The standard Base64 alphabet (RFC 4648 ยง4) is:
| Value | Char | Value | Char | Value | Char | Value | Char |
|---|---|---|---|---|---|---|---|
| 0 | A | 16 | Q | 32 | g | 48 | w |
| 1 | B | 17 | R | 33 | h | 49 | x |
| 2 | C | 18 | S | 34 | i | 50 | y |
| 3 | D | 19 | T | 35 | j | 51 | z |
| 4 | E | 20 | U | 36 | k | 52 | 0 |
| 5 | F | 21 | V | 37 | l | 53 | 1 |
| 6 | G | 22 | W | 38 | m | 54 | 2 |
| 7 | H | 23 | X | 39 | n | 55 | 3 |
| 8 | I | 24 | Y | 40 | o | 56 | 4 |
| 9 | J | 25 | Z | 41 | p | 57 | 5 |
| 10 | K | 26 | a | 42 | q | 58 | 6 |
| 11 | L | 27 | b | 43 | r | 59 | 7 |
| 12 | M | 28 | c | 44 | s | 60 | 8 |
| 13 | N | 29 | d | 45 | t | 61 | 9 |
| 14 | O | 30 | e | 46 | u | 62 | + |
| 15 | P | 31 | f | 47 | v | 63 | / |
So AโZ cover values 0โ25, aโz cover 26โ51, 0โ9 cover 52โ61, and + and / cover 62 and 63.
Worked Example: "Man" โ "TWFu"
Let's encode the 3-byte string Man step by step. This is the canonical example from RFC 4648 itself.
Step 1: Convert each byte to 8 bits.
| Char | Decimal | Binary |
|---|---|---|
| M | 77 | 01001101 |
| a | 97 | 01100001 |
| n | 110 | 01101110 |
Step 2: Concatenate into a 24-bit stream.
01001101 01100001 01101110
Step 3: Re-slice into four 6-bit groups.
010011 | 010110 | 000101 | 101110
Step 4: Convert each group to decimal, then look up the character.
| 6-bit group | Decimal | Character |
|---|---|---|
010011 |
19 | T |
010110 |
22 | W |
000101 |
5 | F |
101110 |
46 | u |
Result: TWFu.
Notice why 3 bytes is the natural input unit: 3 bytes = 24 bits, and 24 is the least common multiple of 8 and 6. Three 8-bit bytes divide evenly into four 6-bit groups with nothing left over. The encoder processes input in 3-byte chunks, emitting 4 characters per chunk.
The 33% Overhead, Explained
Because 3 bytes become 4 characters, Base64 output is always 4/3 the size of the input (for inputs whose length is a multiple of 3). That is a 33.3% size overhead โ a hard mathematical floor, not an implementation detail. A 3 MB image becomes a 4 MB string. This matters enormously when deciding where Base64 is appropriate, as we will see in the data URI section.
Padding: Why the = Exists
What happens when the input length is not a multiple of 3? The encoder cannot form a complete final group of 24 bits, so it pads the last group with zero bits โ and then marks how many padding characters were added using =.
There are exactly two padding cases:
| Input length mod 3 | Remaining bytes | Output chars | Padding |
|---|---|---|---|
| 0 | 3 (full chunk) | 4 | none |
| 2 | 2 | 3 + = |
= |
| 1 | 1 | 2 + == |
== |
Worked Example: "Ma" โ "TWE="
Two bytes remain: M (01001101) and a (01100001). Concatenated: 16 bits.
010011 | 010110 | 0001
The third group has only 4 bits (0001), so it is padded with two zero bits to make 000100 = 4 = E. We emit three real characters and one =:
TWE=
Worked Example: "M" โ "TQ=="
One byte remains: M (01001101). Just 8 bits.
010011 | 01
The second group has only 2 bits (01), padded with four zero bits to make 010000 = 16 = Q. We emit two real characters and two = signs:
TQ==
Why Padding Matters
Padding serves two purposes. First, it tells the decoder exactly how many real bytes the final quantum contains โ without it, a decoder cannot distinguish trailing zero bits from real data. Second, it keeps every Base64 output a multiple of 4 characters, which simplifies streaming decoders and length validation: if a supposedly-Base64 string's length is not divisible by 4, something is wrong.
One caveat: some variants (notably Base64URL as used in JWTs) deliberately omit padding. Decoders for those variants must infer the padding from the string length instead. When you write a decoder, always know which convention your input follows.
Base64URL: The URL-Safe Variant
Standard Base64 has a problem in URLs, filenames, and some protocols: the characters + and / are reserved. In a URL query string, + means "space"; / separates path segments. A standard Base64 string dropped into a URL would be misinterpreted.
RFC 4648 ยง5 defines base64url, which makes two changes:
| Position | Standard | URL-safe |
|---|---|---|
| Value 62 | + |
- |
| Value 63 | / |
_ |
| Padding | = required |
usually omitted |
So the same bytes encode slightly differently:
Standard: a+b/c==
Base64URL: a-b_c (padding stripped)
You will encounter base64url everywhere tokens travel in URLs: JWTs, OAuth code_challenge values in PKCE flows, and URL-safe identifiers. When converting between the two, remember it is a pure character substitution (plus padding handling) โ the underlying bits are identical.
A common bug: feeding a base64url string (with - and _) into a standard Base64 decoder. Most decoders reject it outright; a few silently produce garbage. If your decoder chokes on a JWT segment, check the alphabet first.
Critical: Base64 Is Not Encryption
This is the single most important section of this article, because this is where Base64 causes real security incidents.
Base64 provides zero confidentiality. It is an encoding, like hexadecimal or URL-encoding โ a reversible transformation with no key, no secret, and no security properties whatsoever. Anyone who recognizes Base64 (and every developer does) can decode it in under a second, using tools built into every operating system:
# "Decrypting" Base64 takes one command. There is no key.
$ echo "cGFzc3dvcmQxMjM=" | base64 -d
password123
The string cGFzc3dvcmQxMjM= looks opaque. It is the literal text password123. There is no step between "looks secret" and "is public" other than running a decoder.
How This Goes Wrong in Practice
- "Encrypted" API tokens that are just Base64-encoded JSON. Anyone who decodes them sees user IDs, roles, and permissions in plaintext.
- Obfuscated credentials in config files or mobile apps. Base64-encoding a hardcoded API key does not protect it from anyone who downloads the app โ it only protects it from people who cannot recognize Base64.
- "Secure" URL parameters carrying Base64-encoded user data. They are readable by every proxy, log, and browser history entry between the user and your server.
The rule is simple: if you would not print it on a billboard, do not Base64-encode it and call it protected. Encoding is not encryption.
What to use instead depends on the actual goal:
| Goal | Use this, not Base64-alone |
|---|---|
| Keep data secret in transit | TLS (which you get from HTTPS) |
| Keep data secret at rest | Real encryption, e.g. AES-256-GCM with a proper key |
| Store passwords | A password-hashing function (bcrypt, scrypt, or Argon2) โ never reversible |
| Prove data integrity | HMAC or a digital signature |
Note that Base64 often appears alongside real cryptography โ encrypted ciphertext is binary, so it is routinely Base64-encoded for storage in JSON or databases. That is a legitimate use: Base64 is the packaging, AES is the protection. Confusing the two is the error.
Data URIs: Embedding Binary in Text
One of the most visible uses of Base64 is the data URI scheme (RFC 2397), which lets you embed small binary resources directly in HTML or CSS:
data:[<mediatype>][;base64],<data>
A real example โ a 1ร1 transparent PNG embedded in CSS:
.icon {
background-image: url("data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg==");
}
No separate HTTP request, no extra file to deploy. For a tiny icon, this is genuinely convenient.
When Data URIs Win โ and When They Lose
They win for very small assets (icons, tiny patterns): you eliminate an HTTP round trip, which on high-latency connections can matter more than the 33% size penalty. They are also useful for single-file HTML documents (email templates, standalone reports) where external assets are not an option.
They lose as assets grow, for several compounding reasons:
- The 33% bloat is real. A 100 KB image becomes ~133 KB of text inside your HTML or CSS.
- It blocks parsing. Base64 in your HTML must be downloaded and parsed before the page renders; an
<img>tag, by contrast, loads asynchronously. - No separate caching. Change one byte of your HTML and the browser re-downloads the embedded image too. A separate file would have been cached.
- No progressive rendering or responsive variants. You cannot
srcseta data URI easily, and the image cannot display until fully downloaded.
As a heuristic many frontend engineers use: data URIs for assets under a few kilobytes; real files (ideally modern formats like WebP or AVIF) for everything else. The exact threshold depends on your performance budget โ measure, do not guess.
JWT: Base64URL in the Wild
JSON Web Tokens (RFC 7519) are the most prominent real-world consumer of base64url. A JWT has three parts separated by dots:
eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9 โ header (base64url, no padding)
.
eyJzdWIiOiIxMjM0NTY3ODkwIiwibmFtZSI6IkpvaG4gRG9lIiwiaWF0IjoxNTE2MjM5MDIyfQ โ payload
.
SflKxwRJSMeKKF2QT4fwpMeJf36POk6yJV_adQssw5c โ signature
Decode the header segment and you get:
{"alg":"HS256","typ":"JWT"}
Two things worth internalizing:
The payload is not encrypted. It is base64url-encoded and signed. Anyone who holds the token can decode the payload with zero keys โ try pasting any JWT into a decoder. Never put secrets, passwords, or sensitive PII in a JWT payload. The signature proves the token was not tampered with; it does not hide its contents.
Verify the signature, and verify the algorithm. Early in JWT's history, several libraries accepted tokens with
"alg":"none"โ meaning "this token is unsigned, trust it anyway" โ which allowed trivial token forgery. Modern libraries rejectnoneby default, but the lesson generalizes: when you validate a JWT, pin the expected algorithm explicitly rather than trusting the token's ownalgheader.
When to Use Base64 โ and When Not To
| โ Legitimate uses | โ Misuses |
|---|---|
Email attachments (MIME Content-Transfer-Encoding: base64) |
"Encrypting" or "hiding" secrets |
| Binary data inside JSON/XML APIs | Compressing data (it expands data by 33%) |
| Data URIs for tiny assets (icons) | Large images or files in HTML |
HTTP Basic auth credentials (Basic dXNlcjpwYXNz) |
Storing blobs in a database text column (use a binary/BLOB column) |
| Packaging ciphertext or signatures as text | Replacing proper encryption |
Embedding fonts in CSS (@font-face with small woff2) |
Anything where the 33% overhead matters and alternatives exist |
A useful mental model: Base64 is a packaging format, not a processing step. It answers "how do I carry these bytes through a text-only channel?" If your problem is not "bytes must survive a text channel," you probably do not need Base64.
Try It
Reading about the 6-bit grouping is one thing; watching it happen is better. DailyToolbox's free Base64 encoder/decoder runs entirely in your browser โ paste text, see the Base64 output instantly, flip to decode mode, and try the worked examples from this article (Man โ TWFu, Ma โ TWE=) to verify them yourself. It also handles the URL-safe variant, so you can convert between standard and base64url alphabets without writing a script.