URL Encoding Explained: %20, +, Query Strings, and Common API Bugs
A URL is not one string with one encoding rule. It has a scheme, host, path, query string, fragment, and sometimes user-controlled values inside those parts. Bugs appear when code treats all of them as interchangeable t
A URL is not one string with one encoding rule.
It has a scheme, host, path, query string, fragment, and sometimes user-controlled values inside those parts. Bugs appear when code treats all of them as interchangeable text.
The familiar examples are spaces encoded as %20 or +, a literal plus sign that turns into a space, and an API request whose query parameters break when a user enters & or #.
The fix starts with a simple question: what exactly are you encoding?
Percent-encoding protects structure
Some characters have structural meaning in a URL:
? starts a query string
& separates query parameters
= separates a parameter name from its value
# starts a fragment
/ separates path segments
% begins a percent-encoded byte
If one of those characters belongs to user-provided data, it must be encoded before it is placed into the relevant URL component.
For example, this search term is data:
coffee & tea
If it is inserted directly into a query string, the & can be read as the start of another parameter:
https://example.com/search?q=coffee & tea
The intended value should instead be encoded:
https://example.com/search?q=coffee%20%26%20tea
Percent-encoding represents bytes with % followed by hexadecimal digits. A space can become %20, & becomes %26, and a literal plus sign becomes %2B.
MDN has a useful overview of percent-encoding in URLs, including why the same character may be encoded differently in different contexts.
Why %20 and + both mean space sometimes
This is where many API bugs begin.
For an ordinary URL component, a space is commonly represented as %20:
encodeURIComponent("summer sale")
// "summer%20sale"
HTML form encoding uses a related but different convention. In application/x-www-form-urlencoded data, a space is serialized as +:
const params = new URLSearchParams({
campaign: "summer sale",
});
params.toString();
// "campaign=summer+sale"
Both strings can represent the same value in the right context.
The important distinction is that a literal plus sign is not a space. If the original value is C++ guide, a form-style query must encode the plus signs as %2B:
const params = new URLSearchParams({
q: "C++ guide",
});
params.toString();
// "q=C%2B%2B+guide"
When that query string is parsed, + becomes a space and %2B becomes a literal plus sign.
URLSearchParams follows the application/x-www-form-urlencoded rules. MDN documents that its string parser decodes + as a space, and its serializer writes spaces as +. Read the URLSearchParams reference.
encodeURI and encodeURIComponent solve different problems
JavaScript provides two similarly named functions. They should not be swapped casually.
encodeURI() assumes that the input is already a complete URI. It preserves URL punctuation such as :, /, ?, &, =, and #.
encodeURIComponent() is for one component or value. It encodes a wider set of characters, including &, =, and #.
This is unsafe when value comes from a user:
const value = "coffee & tea";
const url = `https://example.com/search?q=${encodeURI(value)}`;
console.log(url);
// https://example.com/search?q=coffee%20&%20tea
The & remains structural. The server may interpret part of the value as another query parameter.
For an individual value, use encodeURIComponent():
const value = "coffee & tea";
const url = `https://example.com/search?q=${encodeURIComponent(value)}`;
console.log(url);
// https://example.com/search?q=coffee%20%26%20tea
For a full URL with query parameters, the URL and URLSearchParams APIs are usually clearer:
const url = new URL("https://example.com/search");
url.searchParams.set("q", "coffee & tea");
url.searchParams.set("source", "docs");
console.log(url.toString());
// https://example.com/search?q=coffee+%26+tea&source=docs
The serialized space appears as + because searchParams uses form-style query encoding. The value remains correct.
Encode values, not an entire URL by accident
A common mistake is passing a complete URL through encodeURIComponent():
encodeURIComponent("https://example.com/search?q=coffee");
// "https%3A%2F%2Fexample.com%2Fsearch%3Fq%3Dcoffee"
That output is useful only when the complete URL itself must become one value, such as a redirect parameter:
const redirect = "https://example.com/account?tab=billing";
const url = `https://auth.example.com/login?redirect=${encodeURIComponent(redirect)}`;
It is not a replacement for constructing a normal URL.
The reverse mistake is decoding a full URL or query string blindly. A decoded &, =, or # may regain structural meaning in the wrong layer. Decode the component you expect, then validate what it means.
Avoid double encoding
Double encoding happens when already encoded text is treated as raw text and encoded again.
encodeURIComponent("hello%20world");
// "hello%2520world"
% becomes %25, so %20 becomes %2520.
This is sometimes intentional. It is often a sign that one layer of the application does not know whether it receives raw or encoded data.
A practical rule helps:
- Store and pass decoded values inside your application.
- Encode only when constructing the URL or request.
- Decode only when parsing a URL or request.
- Document any boundary that intentionally expects encoded text.
The same rule applies to URLSearchParams. Give it raw keys and values:
const params = new URLSearchParams();
params.set("q", "coffee & tea");
Do not pre-encode the value before calling set().
URL encoding is not validation or security
Encoding preserves URL structure. It does not prove that a URL is allowed, trustworthy, or safe to request.
For example, percent-encoding does not:
- allowlist redirect destinations;
- prevent server-side request forgery;
- validate an API token;
- escape HTML for display in a page;
- make a query parameter safe for SQL;
- confirm that a URL uses HTTPS.
Those are separate checks with separate rules.
If your application accepts a user-supplied URL, parse it with new URL(), validate the protocol and host against the rules for that feature, and avoid using string replacement as a parser.
Test the cases that tend to break
When testing URL handling, do not stop at a simple word.
Use values such as:
summer sale
C++ guide
coffee & tea
a=b
100% ready
https://example.com/a?x=1
你好
These cases reveal different failures:
- spaces test
%20versus+; -
+tests whether a literal plus survives form parsing; -
&and=test query-string structure; -
%tests malformed or double encoding; - complete URLs test component boundaries;
- non-ASCII text tests UTF-8 encoding.
I built the ToolExo URL Encoder and Decoder for this exact distinction. It keeps URL component encoding and form query-value encoding separate, shows %20 versus +, encodes a literal plus as %2B in form mode, and decodes only one layer at a time.
The tool is for a single component or query value. It deliberately does not rewrite a complete URL, host name, or protocol. Those parts have their own structure and should be parsed as URLs, not treated as text.
A URL bug often looks small in a log. It can be one missing %26, one accidental +, or one value encoded twice. Keeping the URL structure separate from its data values makes those bugs much easier to prevent.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.