URL Encoder / Decoder
Encode text into a URL-safe string or decode one back, live on every keystroke. RFC 3986 and form modes, a character map of exactly what changed, and nothing leaves your browser.
11 characters in → 13 characters out
Character map
11 charactersOne row per character in the text input, with its Unicode code point, its class and its RFC 3986 encoded form.
| Character | Code point | Class | Encoded |
|---|---|---|---|
| h | U+0068 | unreserved | h |
| e | U+0065 | unreserved | e |
| l | U+006C | unreserved | l |
| l | U+006C | unreserved | l |
| o | U+006F | unreserved | o |
| ␣ | U+0020 | space | %20 |
| w | U+0077 | unreserved | w |
| o | U+006F | unreserved | o |
| r | U+0072 | unreserved | r |
| l | U+006C | unreserved | l |
| d | U+0064 | unreserved | d |
How it works
- 1
Pick the direction
The Encode tab turns text into a URL-safe string; the Decode tab turns a percent-encoded string back into text. Both run on every keystroke, so there is no convert button — switch tabs and the result is already there.
- 2
Choose the mode
RFC 3986 writes a space as %20; form mode, matching application/x-www-form-urlencoded, writes it as +. That single substitution is the entire difference between the two modes, and the selector applies to both directions.
- 3
Read the live result
The encoded or decoded value appears under the input with a copy button. A decode that cannot succeed — a stray %, or escapes that are not valid UTF-8 — says so plainly instead of producing garbage.
- 4
Inspect the character map
The panel under the result lists one row per character in the input: the character itself, its Unicode code point, its class (unreserved, reserved, space, percent or unsafe) and its encoded form. Long inputs stop at 500 rows with a note telling you how many more there are.
Two modes, one alphabet
Percent-encoding exists because a URL is a restricted alphabet wearing a longer one's clothes. RFC 3986 allows a URI to carry unreserved characters bare — letters, digits, and the four punctuation marks - . _ ~ — and defines everything else as either reserved, meaning it may carry structural meaning, or simply unsafe outside a URL context. The escape mechanism is uniform: take the character's UTF-8 bytes and write each one as a percent sign followed by two hexadecimal digits. A decoder runs the same rule backwards, which is what makes the scheme lossless for any text at all.
In practice two dialects of the scheme are in service, and they differ in exactly one place: how a space is written. RFC 3986 percent-encoding — what this page's Encode tab produces by default — writes it as %20. The older form dialect, application/x-www-form-urlencoded, was standardized around HTML form submission and writes the same space as +. Every other character encodes identically in both, which is why the mode selector changes so little and yet matters so much. The same input in each mode shows the shape of it:
| Input | RFC 3986 mode | Form mode |
|---|---|---|
| hello world | hello%20world | hello+world |
| a+b | a%2Bb | a%2Bb |
| !'()* | %21%27%28%29%2A | %21%27%28%29%2A |
| 😀 | %F0%9F%98%80 | %F0%9F%98%80 |
Three of those rows are identical across the modes, and they repay a second look. The plus in a+b has to become %2B in both modes — in form mode because a bare plus would be read back as a space, destroying the value, and in RFC 3986 mode because the plus is reserved and this tool never leaves reserved characters bare. The punctuation row !'()* exists because JavaScript's encodeURIComponent quietly leaves those five unescaped, while RFC 3986's unreserved set does not include them; this page escapes them so its output survives the strictest parser. And the emoji row shows the byte arithmetic at work: four UTF-8 bytes, one escape each.
Decoding mirrors the same split. In form mode the decoder reads every plus as a space before resolving escapes, so a+b comes back as 'a b'; in RFC 3986 mode the plus is an ordinary reserved character and comes back a plus. That asymmetry is the single most common cause of a wrongly decoded URL parameter — a value whose pluses were literal, decoded with form mode, arrives with spaces where pluses used to be. When in doubt, decode a short sample both ways; the wrong mode announces itself immediately.
What gets escaped: the three tiers
Every character an encoder meets falls into one of three tiers, and the tier decides the output. Unreserved characters — A-Z, a-z, 0-9, - . _ ~ — pass through untouched; encoding them would only make URLs longer for no benefit. Reserved characters — : / ? # [ ] @ ! $ & ' ( ) * + , ; = — are legal in a URI but meaningful, so an encoder escapes them whenever they appear as data rather than structure. Everything else is unsafe in a URL by definition: spaces, quotation marks, angle brackets, accented letters, CJK characters, emoji, control characters. Percent-encoding exists to carry that third tier, and the character map panel on this page labels each character of your input with its tier as you type.
| Class | Members | What encoding does |
|---|---|---|
| unreserved | A-Z a-z 0-9 - . _ ~ | left exactly as they are |
| reserved | : / ? # [ ] @ ! $ & ' ( ) * + , ; = | escaped when they appear as data |
| space | the space character | %20, or + in form mode |
| percent | % | %25 — the escape introducer, escaped like anything else |
| unsafe | everything else, including all non-ASCII | one escape per UTF-8 byte |
The percent row deserves its own line because it is the mechanism watching itself. A percent sign in your input can never be emitted bare: %25 is the only way a decoder can tell an intentional percent from the start of an escape. That is also why the character map gives % a class of its own — a % sitting in text you are about to encode is usually a sign that the text was already encoded once, and the map makes it visible before you create a %2520.
The character map lists one row per character, ordered as they appear in the input, with the character, its Unicode code point, its class and its encoded form. Emoji and other characters outside the Basic Multilingual Plane occupy a single row each, because the map counts code points rather than the surrogate pairs JavaScript uses internally to store them. Inputs longer than 500 characters stop the table there and note how many further characters were not listed — the encoding itself is unaffected, only the display is capped.
The double-encoding trap
Double encoding is the most common self-inflicted wound in URL handling, and it follows from one fact: the percent sign is itself a character that must be encoded. Encode 'a b' and you get a%20b. Encode that result again, and the encoder does exactly what it is told — it sees a percent sign that needs protecting and writes %25 in its place, giving a%2520b. The string is now stable under further encoding, but it no longer decodes to what anyone expects: one pass of decoding yields a%20b, the literal characters percent-two-zero, and only a second pass recovers 'a b'.
It creeps in through layers that each look reasonable. A framework percent-encodes query parameters automatically, then a template escapes the URL again before printing it. A developer encodes a value 'to be safe' when it was already encoded upstream. A link is built by concatenating an encoded string into another URL that is itself encoded at the end. Each layer is individually correct; the composition is the bug. The symptom is always the same shape — encoded text appearing where a decoded value should be, a page showing %20, or a search endpoint that literally cannot find the words you typed.
The repair rule is to encode once and late: at the single point where a value leaves your code and enters a URL, using the component-level encoder, and never before. Any string that already contains %20, %2F or %3A should be treated as finished output, not as input to protect again. When you inherit a value that looks over-encoded, this page's Decode tab is the diagnostic: decode it, look at what comes out, and if the result still contains percent escapes, it was encoded twice — decode again and then go fix the producer rather than stacking another decode downstream.
When decoding fails, and what it means
Encoding can succeed on any input, but decoding can genuinely fail, and the failures divide into two families. The first is structural: a percent sign that is not followed by two hexadecimal digits. A string like 100% ends with a percent that has nothing after it, and %2 or %zz start escapes they cannot finish. Strict decoders refuse these rather than guessing, because treating a bare percent as literal text would make the outcome depend on the decoder's mood instead of the data. This page reports the problem in plain language instead of throwing.
The second family is byte-level: escapes that resolve to bytes which are not valid UTF-8. Percent sequences encode bytes, and bytes only spell characters when they follow UTF-8's rules — a lone %C3 promises a two-byte character and then never delivers the second byte, and %FF is not a legal UTF-8 start at all. Sequences like these usually mean the string was encoded from a different character set, Latin-1 being the usual suspect, and the honest answer is that the text cannot be recovered as written; it needs the character set it was actually encoded from.
One quiet trap sits between the families, and it is the plus again. Decode the form-encoded a+b with RFC 3986 mode and nothing fails — the tool returns a+b, technically correct and practically wrong, because the plus you are looking at was a space all along. Nothing flags it, because nothing is malformed. That is why mode choice is a question about the string's origin rather than its contents: malformed input announces itself, while a mode mismatch stays silent until a human notices the spaces that should have been pluses.
Frequently asked questions
Why does '+' turn into a space in one URL and not another?
What's the difference between encodeURI and encodeURIComponent?
Why do I keep seeing %2520, and how do I undo it?
Why does one emoji become a dozen characters of hex?
Which mode should I pick?
Does anything I paste leave my browser?
Related tools
Last updated: October 9, 2026