Skip to content
SEO & Content

Changing a URL is easy. The redirect decides whether the move works

A redirect answers in place of the page you moved: which code to use, where it's written, and how to check the server isn't getting it wrong

It takes two minutes to set up a redirect. You check in the browser and the page you wanted shows up. That’s usually where the testing stops.

The old URL, though, is still the one everyone else knows. It lives in links published on sites you don’t control, in emails sent months ago, in bookmarks, in feeds, in PDFs someone downloaded, in search engine indexes. You can change the address on your site whenever you want: those references will keep pointing to the old one, and there’s almost nothing you can do about most of them.

You decide what responds when that old address gets called. The person who clicks a three-year-old link and the crawler that comes back tonight both end up wherever you sent them.

What a redirect is

A redirect is the response a server sends to say that the content requested at a given URL now lives at a different address. Instead of the page, the server sends a 3xx status code, which states whether the change is permanent or temporary, along with the header Location, where the destination address is written. There’s no actual document involved: at most the server attaches a couple of courtesy lines that nobody reads, because the client heads straight for the destination.

The recipient is a client, which in HTTP terms means anything that makes a request: the browser of someone browsing the web, but also Googlebot, the crawlers that gather content for generative engines, a program downloading a feed, or a partner’s integration querying your endpoint. The client reads the code, takes the address written in Location and requests that instead: it’s the next call that returns the page. Whoever is browsing doesn’t notice, because the browser handles this on its own and only shows the final destination.

That’s why a broken redirect doesn’t announce itself the way other errors do. A page with a typo, a broken link, a form that fails to submit: someone opens it, sees it, tells you. A response isn’t read by anyone – it’s executed, and executed the same way whether it’s right or wrong, because as far as HTTP is concerned the transaction succeeded either way: the server answered, the client complied. There’s nothing anyone has a reason to flag, and you find out only if you go looking.

Attiva una prova di SEOZoom e scopri cosa dice di te l'AI

Even your own browser won’t help you, because it’s doing exactly what it’s supposed to: the headers arrive before the document, the browser processes them before rendering anything, and by the time the page appears, the code has already done its job. You open the old address, see the right destination, and move on. What you didn’t see is which code got you there, and through how many hops.

What a redirect is not

Internal rewrites, canonical tags, meta refresh and JavaScript-based moves all get called redirects in everyday conversation, but none of these four actually works like a 3xx response. What sets them apart is when they act, because that determines who receives them.

An internal rewrite doesn’t move anything: the server delivers a different file under the URL that was requested, and whoever made the request gets no new information. A canonical tag leaves both pages reachable with a 200 response and tells search engines which version to prefer, without forcing anyone to change the address they use. A meta refresh and a script-based move both act after the document has already been delivered, so they only reach whoever interprets or executes that document.

The real test, then, is whether the old address stops responding. With a rewrite, a canonical tag, or a move declared inside the document, the old URL keeps returning a 200 to anyone who requests it: the page is still served from that address, and anyone who doesn’t parse the HTML has no way of knowing you meant to send them elsewhere. A 3xx response kicks in before the document, and it’s the only one that works the same way for every client.

What actually arrives for whoever calls the address
Mechanism What arrives for whoever calls Who receives it What it’s for
3xx redirects a response instead of the document, with the destination whoever makes the call move a URL
Internal rewrite (RewriteRule without [R]) the page, under the URL that was requested stays inside the server deliver a different file while keeping the address
rel="canonical" the page, with an indication in the document, and both versions returning 200 search engines, as a preference indicate which version to consolidate
meta refresh and JavaScript the whole page, with the move happening after delivery whoever parses the HTML or runs the code edge cases, when the server can’t respond beforehand

The last option costs you more today than it used to. Crawlers for generative engines download the document but don’t execute the code – the mechanics, with the numbers, are covered in our guide to technical SEO – and a redirect written in JavaScript simply doesn’t exist for them: the old URL stays the valid one, and the new page is never considered published.

What types of redirect exist

The code you choose makes a claim about that address’s future, and it keeps making that claim to every caller, even after you’ve changed your mind. The available codes all belong to the 3xx family, and they say different things.

301 states that the address isn’t coming back. 302 states the opposite: the detour is temporary and the original stays valid. Google treats a 301 as a strong canonicalization signal – meaning an indication that the new URL takes the old one’s place – and a 302 as a weak one.

307 and 308 make the same two claims, temporary and permanent, while adding a guarantee about how the request reaches its destination. For Google, 308 is equivalent to 301 and 307 to 302. The exception is 303, which points to a resource different from the one requested and makes no claim about duration.

The five 3xx codes and what each one declares
Code What it declares Method preserved Cached without expiration For Google
301 permanent move no yes strong signal
302 temporary move no no weak signal
303 different resource, no duration no no –
307 temporary move yes no same as 302
308 permanent move yes yes same as 301

The same family also includes 304, which doesn’t redirect anyone: it tells the client that the copy it already has is still valid.

Difference between 301 and 302

This choice comes up every time an address changes, and it usually gets settled on whatever seems like the safest code: when in doubt, 301, because it sounds definitive. That’s the wrong reasoning, because it confuses certainty about the past – I moved that page – with certainty about the future, which is the only thing the code actually communicates.

Permanent and temporary, however, are not symmetrical, and the asymmetry doesn’t concern the day you write the rule, but the one when you have to correct it. RFC 9110 establishes that 301 and 308 can be stored in cache even without an explicit expiration, while 302, 303, and 307 cannot. A 301 that you published by mistake therefore keeps sending to where you had said even after you’ve corrected it, because clients that had already received it don’t go back to asking your server anything: the correction simply doesn’t reach them. It’s up to them to clear what they have stored, and when that happens, they decide.

A mistaken 302, on the other hand, gets fixed the moment you correct it, and the change applies starting with the next call. When you genuinely don’t know that address’s future, the cautious choice is the temporary code – the opposite of what instinct tells you to do.

307 and 308, the codes that preserve the method

RFC 9110 notes, for both 301 and 302, that for historical reasons a client may change the request method from POST to GET. It’s a holdover from how browsers behaved thirty years ago, and it stayed in the spec because changing it would have broken everything built on top of it since.

On your site this means: if you move the address a form submits data to, and you do it with a 301, the form arrives at its destination without the data it contained. The person who filled it out sees the confirmation page, because the confirmation gets served regardless, but you receive nothing. There’s no error anywhere and every response involved is technically correct, so no technical check flags it. You only notice the breakage in the submissions that never arrive.

307 and 308 preserve the method, which is the only reason they exist. On an endpoint that receives data, these are the codes to use, and they’re also the two codes nobody looks at, because forms don’t show up in a site’s URL map.

The links pointing to the old address

An address that’s been online for years has picked up links from other sites, and those links stay where they are: you can’t rewrite them and you can’t fix them. The question that follows immediately is whether they still count for anything now that anyone following them passes through a redirect.

Google’s answer is yes: permanent redirects don’t cause PageRank loss, and a 301 is the strong signal that tells Google the new URL has taken the old one’s place. Anyone arriving from that link ends up where they would have ended up anyway, just with one extra hop.

That statement doesn’t stand on its own, though. It holds on the engine’s timeline, not yours, because consolidation happens when the old address gets recrawled and the move gets registered. It holds if the destination is relevant, because many URLs sent to a page that has nothing to do with them get treated as soft 404s, and nothing consolidates from a soft 404. And it holds only for permanent redirects: with a 302 the signal stays weak, and Google can keep treating the old address as canonical, which is exactly what that code is telling it to do.

Writing the right rule isn’t enough on its own, then. The code tells the engine what happened to the address; the relevance of the destination decides whether there’s anything to consolidate. A wrong code slows the transfer down; a wrong destination cancels it out.

Outside of Google, this mechanism has no stated equivalent. A redirect only affects calls that come after it, and whatever a generative system has already collected sits in an archive that isn’t yours. There’s no documented way for a 301 to retroactively rewrite what was captured earlier. On the old address, you’re deciding what tomorrow’s visitors will find, not what last year’s readers already saw.

Those automated readers, by the way, aren’t all the same reader. OpenAI distinguishes OAI-SearchBot, which is used to appear in ChatGPT’s search responses, from GPTBot, which collects content for model training, and governs them with robots.txt independent controls. These are two visits with different purposes and their own timing, to the point that the documentation states it takes about twenty-four hours to pick up a change to the file. Same address, same redirect, but consequences that play out on different timelines.

When you need a redirect

A URL you published stops leading to the content it promised, and the reason it stopped decides the work ahead of you. If what’s moving is the infrastructure – the domain, the protocol, the shape of the paths – the destinations already exist and the job is matching them one by one. It’s long, it’s mechanical, and you can hand it off to someone who knows how to write a rule. If what’s moving is the content, every destination has to be chosen individually, and that part can’t be delegated, because it requires knowing what the person landing on that page was actually looking for.

  • You change domains. The entire site moves and the pages stay the same, so every old path has an exact match on the new domain.
  • You switch to HTTPS. Until the old version stops responding, the site is reachable on two protocols with the same content, and search engines face two copies to choose between.
  • You change CMS or URL structure. The new platform rewrites the shape of the paths, and the pages you had indexed now respond to addresses that didn’t exist before.
  • You merge two sites. After a merger or acquisition one of the two domains gets retired, and whatever it had built up only reaches the surviving domain if you redirect it.
  • You retire a page. A product listing out of the catalog, two articles saying the same thing that become one, a guide replaced by an updated version: here there’s no exact match and the choice is yours.
  • You consolidate variants of the same content. With and without www, with and without a trailing slash, with parameters tacked on that leave the page identical: without a single version, links and signals spread across different addresses showing the same thing.
  • You put a page into maintenance. The move lasts only as long as the work does, and declaring it temporary rather than permanent changes which code to use.

Then there are the addresses nobody thinks of as part of the site, and that still respond: a campaign’s short URLs, a feed endpoint, the address a partner sends calls to. These get missed easily during migrations, because they’re not in the menu and nobody notices them until they stop working.

Alternatives to a redirect

Redirecting makes sense only if there’s content covering the same need as whatever you’re removing. Reachable doesn’t mean relevant: you can send everything somewhere, get a 200 response for every old address, and still have moved people onto something they weren’t looking for. The right answer depends on what actually happened to that page, and the options are few.

The right response depending on what happened to the page
What happened What to respond What happens if you get it wrong
A page exists that answers the same need 301 to that page –
The content is gone and has no replacement 404 o 410 redirecting forces someone looking for one thing toward another
Both pages need to stay reachable rel="canonical" or noindex a redirect pulls one of them out of circulation
The page will be back in a few hours 503 with Retry-After a placeholder page served with 200 tells the crawler that content is the final version
The entire site is unreachable 503 across the board with 5xx and 429 crawling slows down, and indexed URLs stick around for a while before getting dropped

The difference between 404 and 410 is smaller than people make it out to be, but not where you’d expect. 410 has a reputation for speeding up removal from the index, and it doesn’t: Google treats all 4xx errors the same way except for 429, and in both cases the page comes out when it gets recrawled, exactly as with a 404.

That “when” is the part that decides, and it’s not up to you. Recrawling an address doesn’t follow a schedule. It follows how often that address gets requested and linked to: pages people keep calling come back under the crawler’s eye within days, while pages nobody calls anymore can sit around for a long time, and those are exactly the ones nobody’s looking for. If the problem is pulling content out of the results quickly – a wrong price, a document published by mistake – neither 404 nor 410 is the right tool. You need Search Console’s removals tool, which blocks that address from showing up in results for about six months while deindexing runs its course. It’s one of the routes we laid out in the guide on how to remove content from Google.

The effects of choosing between 404 and 410 show up elsewhere, and on a large site they’re concrete. A 410 states that removal was intentional; a 404 leaves open the possibility of a glitch. For monitoring, that means you can archive the 410s, while every 404 needs a look to figure out whether it’s a removal or a glitch. On a catalog where hundreds of listings come and go every month, that’s the difference between a report someone actually reads and a list of errors nobody opens anymore.

A 503, on the other hand, runs on a single variable: duration. It states that the page isn’t available right now and that it’s worth checking back, and with Retry-After you can say when. As long as the downtime lasts hours or days, it’s the right response and it costs you nothing. If it lasts weeks, the consequence changes, because with 5xx and 429 crawling slows down and indexed URLs get kept around for a while before being dropped. A 503 that nobody removes ends up where a 404 ends up, just more slowly and without anyone deciding it that way.

The workaround people reach for instead is worse: a courtesy page served with a 200 status, the “we’ll be back soon” kind. That response states that the final content at that address is the apology message, so it can get indexed in place of the real page, and when the site comes back, you’re left with the wrong version in the index to fix.

Canonical and noindex matter when removing isn’t what you want. If pages need to stay reachable, both of them, such as variants, filters, print versions, a redirect would pull one of them out of circulation even for people who use it. The rel="canonical" signals which version to consolidate while leaving the other accessible, and it’s a preference the search engine can choose not to follow. The noindex does something else entirely: the page stays reachable for people but drops out of the results. Putting both on the same page means declaring that this page is a secondary copy of another one while also saying it shouldn’t appear at all. Those are conflicting instructions, and you don’t get to decide which one wins.

Finally, there’s one condition that applies no matter what destination you choose: it has to be a page that ends up in the index. AI Overview and AI Mode draw on Search’s index, and to show up there a page has to be indexed and eligible to be shown with a snippet. If you send old URLs to a destination that never reaches the index, you’re removing them from both surfaces at once.

The risk of sending everything to the home page

Any map, no matter how carefully built, leaves a remainder: the addresses nobody managed to match to anything. For that remainder there’s a single rule that catches them all and sends them to the homepage, and at first glance it looks like the sensible choice, one rule instead of hundreds of decisions, and no 404 to explain to anyone.

Google’s migration guide says that redirecting many old URLs to a single, unrelated destination can be treated as a soft 404. So those pages drop out of the index the way they would have anyway. The homepage rule didn’t prevent anything, it just took away your way of noticing, because in your reports they no longer show up as addresses returning an error.

Then there’s a cost no report measures. Someone arriving from a three-year-old link lands on the homepage with no explanation, in a part of the site that has nothing to do with what they were looking for, and you lose that visitor right there. A well-built 404 page, with internal search and links to the main sections, at least tells them what happened and gives them a way to keep going.

The catch-all rule is useful – it’s the safety net under the map, and without it, anything you didn’t account for is left exposed – but its destination should be a 404, not the homepage. And when there’s no exact successor but there is a matching category, the category is a legitimate destination, not a fallback, because it holds the same kind of content and gives visitors a place to start over. That’s the order you build a map in for any migration: first the exact successor, then the level above it, and only at the end, the removal response.

Where to set up a redirect

As with any piece of technical SEO, you can declare the same redirect in more than one place, and the one that answers is always the first layer that intercepts the request. If there’s a CDN in front, it stops there and nothing reaches the origin; if it does reach the origin and gets passed to the application, WordPress can respond before your .htaccess is even considered. When a redirect isn’t doing what you think, the first thing to figure out is who answered, because a fix only takes effect at the layer that’s actually responding. Comparing the response you get against the origin’s logs tells you right away: if the request doesn’t show up in the logs, something in front is answering.

  • Apache. For simple cases, the mod_alias directives are enough, and they don’t need regular expressions.
    Redirect 301 /old-page/ https://www.example.com/new-page/
    RedirectMatch 301 ^/blog/(.*)$ https://www.example.com/articles/$1

    When you need conditions, you move to mod_rewrite, and the [R] flag is what turns an internal rewrite into an actual move: without it, the server delivers a different file while leaving the address where it was.

    RewriteEngine On
    RewriteRule ^products/([0-9]+)/?$ /catalog/$1/ [R=301,L]

    On the query string, the default behavior is asymmetric: if the replacement doesn’t contain one, the original is kept; if it introduces one, the original gets dropped. That’s how a campaign’s tracking parameters disappear along the way, and the visits that come from them get attributed elsewhere. [QSA] appends the original to the new one, [QSD] removes it.

  • nginx. The directives live in the configuration and require a reload, so touching them means having server access, and nobody changes them by accident.
    server {
        listen 443 ssl;
        server_name www.example.com;
    
        location = /old-page/ {
            return 301 https://www.example.com/new-page/;
        }
    
        rewrite ^/blog/(.*)$ /articles/$1 permanent;
    }

    return is more efficient than rewrite and should be your choice when the match is exact. The order of the location blocks doesn’t follow the file’s order but nginx’s own precedence rules, and an exact match with = wins over the others: if your rule seems to be getting ignored, it’s almost always because a more specific one is winning instead.

  • PHP. The function that sets the header returns a 302 if you don’t declare otherwise, and that’s how you end up publishing a permanent move while declaring it temporary.
    header('Location: https://www.example.com/new-page/', true, 301);
    exit;

    It has to be called before the server sends any output – even a stray space outside the PHP tags counts – otherwise you get a warning and the page is served normally; if output has already started, ob_start() at the top of the script holds it back. And it needs to be followed by exit, because otherwise the script keeps running and the code below still executes, with whatever that implies if there are database writes in there.

  • WordPress. Both available functions default to 302 and neither one stops execution, so exit is needed here too.
    add_action('template_redirect', function () {
        if (is_page('old-page')) {
            wp_safe_redirect(home_url('/new-page/'), 301);
            exit;
        }
    });

    The difference is in the destination. wp_safe_redirect() validates it and, if it points to a disallowed host, sends the user to wp-admin instead of off-site; the list of allowed hosts can be extended with the allowed_redirect_hosts filter. wp_redirect() validates nothing, and you should avoid it anywhere the destination comes from a parameter: an endpoint that redirects wherever the URL tells it to is an open redirect, and it’s used to send visitors off to a site that isn’t yours while keeping your domain in the link. If you manage redirects through a plugin, keep in mind that it becomes one more layer in the chain of responders, and the first place to check when a rule written at the server level seems to have no effect.

  • CDN and reverse proxy. The layer in front responds before the origin does, and it produces effects you’ll only notice if you go looking for them. Calls redirected there never show up in your server logs, so if that’s where you check, you’ll conclude nobody is requesting that URL anymore. And Location, which per RFC 9110 can hold either an absolute address or a relative reference resolved against the requested URL, can send you to the wrong host when a relative reference sits behind a proxy that changes the host. If your HTTP-to-HTTPS switch is set up at the origin while TLS termination happens at the edge, the origin always sees an HTTP request and redirects to HTTPS endlessly: that’s what X-Forwarded-Proto is for, and it’s why that loop only shows up in production. And then there’s a redirect you’ll see in the browser that you never wrote: with HSTS active, the browser converts HTTP calls to HTTPS on its own before they even go out, and logs it as a 307 Internal Redirect, which doesn’t come from your server and you’ll never find it in your configuration.

    The layer in front can also generate redirects on its own, for just one category of callers. Cloudflare’s Redirects for AI Training applies your rel="canonical" tags as 301s toward verified AI training crawlers, when the canonical points to another URL on the same host, and leaves the page untouched for browsers, search engines, and assistants. It’s a redirect that exists for some but not others, and it won’t show up in your server configuration.

When .htaccess isn’t read

A correct rule can produce no effect at all, and the most common reason lies elsewhere. The default value of AllowOverride is None, and with that value, .htaccess files get ignored entirely: what you wrote isn’t wrong, nobody’s reading it. AllowOverrideList lets you enable individual directives without opening everything up.

Inheritance makes a correct rule disappear in another way. When it merges sections from the same context, mod_rewrite by default replaces the rules instead of adding to them, so a .htaccess in a subdirectory can wipe out the one above it; RewriteOptions Inherit, InheritBefore e InheritDown change this behavior.

If instead you used a directive that isn’t allowed in that context, the server returns a 500 for the entire directory, and the log shows a line naming the rejected directive. You can’t get that name from the error page itself; it only shows up there.

From HTTP to HTTPS and from www to non-www

The protocol and the host prefix don’t affect a single page but every page, and that’s exactly why they’re almost always decided at different times, by different people. The typical result is a chain: http://www that sends to https://www, which sends to https://, with two hops where one would do. Every call that comes in through the old form pays for it – the one written into links published years ago and bookmarked by visitors who’ve known the site since before.

RewriteEngine On
RewriteCond %{HTTPS} off [OR]
RewriteCond %{HTTP_HOST} ^www\. [NC]
RewriteRule ^(.*)$ https://esempio.it/$1 [R=301,L]

A single rule that resolves to the final form in one hop beats two correct rules written separately, and when you find them already stacked, rewrite them together instead of adding a third. The check is immediate: request the farthest-back form, http://www, and count the hops that bounce back.

Different destinations by language and country

A redirect that changes destination based on IP orAccept-Language seems to work because you’re testing it from where you are. Googlebot crawls mostly from US addresses, so of your entire multilingual site it sees just one version, and that’s the only one that can end up in the index.

Google’s recommended alternative is adaptive pages, which serve different content on the same URL, or separate addresses per language connected with hreflang. Either way, every version stays reachable by anyone, which is the condition for it to get crawled.

Redirects during a migration

What makes a migration difficult isn’t the number of addresses but their inconsistency: a thousand URLs with the same pattern are covered by a single rule, while fifty with no pattern mean fifty decisions made by hand.

So the work doesn’t start with the rules, it starts with the map: the list of old addresses paired with the destination chosen for each one. That part can’t be delegated, and it’s also what makes everything else mechanical, because once the map exists, writing the rules becomes a matter of syntax. Whatever doesn’t make it into the map gets no response, and you won’t notice, because you’re not the one requesting that address.

Where the old URLs live

No single source contains the complete list, and each one has a blind spot that depends on how it’s built.

The sitemap is the most convenient and the most incomplete: it collects what the site declares today, meaning what the CMS considers published, and knows nothing about addresses you retired years ago or ones the CMS never generates – PDFs, hand-built landing pages, paths from an old platform. Server logs hold the opposite: not what you declare, but what was actually requested, including addresses you didn’t remember having. You’ll also find calls from generative crawlers there, which work off outdated inventories: the URLs they keep requesting are old by definition, and that’s a list no other source gives you. Their limitation is the time window, since logs are usually rotated after a few weeks, so a URL requested twice a year won’t show up.

Search Console shows what Google has seen, and it’s the source that also tells you who links to what, but its reports are sampled and the example lists are capped. Backlinks give you the addresses other sites have published, which matter most because they’re the ones that will keep getting requested for years, and they’re also the ones that may no longer exist in your CMS. A site crawl, finally, finds what’s reachable by navigation: it leaves out orphan pages, the ones no internal link reaches and that you’d never think of on your own.

The sources overlap because their blind spots don’t line up, and the order you use them in changes the result. Start with backlinks and logs, which give you live addresses even forgotten ones; add the crawl, which gives you the current structure; finish with sitemap and Search Console for cross-checking. What comes out is still an approximation, which is why a fallback rule that catches unplanned cases is always worth having, even when you think you’ve mapped everything.

Which addresses take precedence

A map with a few thousand rows doesn’t get published or verified in one afternoon, and almost no rollout is instant: you go block by block, or branch by branch of the site. That means some addresses have their rule from day one, while others get theirs weeks later, and in the meantime anyone calling the latter hits a 404. The order, then, isn’t a matter of convenience: it decides which addresses go unprotected for the longest stretch.

What determines it isn’t the search volume of the associated keywords, which tells you how much a topic is searched, not how much that page is worth to you. It’s determined by links from other sites, which you can’t reconstruct and which will keep coming in for years; presence in search results, because those addresses are already circulating and get picked up by search engines; and calls in the logs, which tell you who’s using that URL right now, including integrations and feeds no other tool can see.

Visibility here should be taken for what it is: an estimate of presence in search results, not a measure of traffic, since no ranking guarantees a number of clicks. It’s useful for ranking your URLs against each other, not for predicting how much traffic each will get.

The three readings live in different places, and none of them work by keyword, because the question isn’t which topics drive searches but which addresses are still earning their keep. For the Search Console part, that means looking at aggregated data by URL instead of by query: that’s the Pages performance view, which for each address keeps clicks, impressions, average position and keyword count together, letting you sort the list before you start. You read inbound links through backlink analysis, and server calls through the logs: cross-reference all three, and wherever they agree, that’s your first block to roll out.

Addresses left out of that block shouldn’t just sit there waiting their turn: that’s exactly the job of the catch-all rule, which routes them to the category page or to a curated 404 instead of leaving them with no response at all.

What to declare to Search Console

While you move, Google wants two sitemaps in Search Console: one for the old URLs and one for the new ones. At the start the new sitemap has zero indexed pages and the old one has many. Watching those numbers flip tells you where the migration stands better than anything else you could look at. Once the flip is complete, the move is done and you remove the old sitemap.

The address change setting is something else, and it only applies in one case: when the site moves to a different domain. That’s when you declare it explicitly and Google treats it as such. If only the URL structure changes and the domain stays the same, there’s nothing to declare. What matters then are the rules you wrote.

The limits of a regex written in a hurry

A regular expression that rewrites an entire section of the site works as long as the relationship between old and new URLs is truly regular, and it breaks on the edge cases you weren’t thinking about when you wrote it. A rule like ^/blog/(.*)$ to /articoli/$1 looks like it covers everything, but it also catches /blog/ with no slug, which ends up at /articoli/ even when that page doesn’t exist. Just write it without the trailing slash, ^/blog(.*)$, and it drags along /blogger/too. And it misses paths with accented or encoded characters, where %C3%A8 e è aren’t the same string.

Then there are parameters, which a regex built on the path doesn’t see, and capitalization, which on a case-sensitive file system turns the same content into two different addresses. Before you ship the rule, test it against a hand-picked list of edge cases, one for each of these patterns, not against the three URLs you used to write it, since those work by construction.

Then there’s the part a rule can’t handle. A regex knows how to turn one path into another path, not whether the destination page talks about the same thing. On a uniform section the transformation is enough. Where content has been reorganized, every match has to be decided by hand, and the rule only covers what’s left over.

Checking redirects, from a single URL to the whole site

The browser confirms the destination and hides everything else: what code the server actually responded with, how many hops the request made, what a non-browser client receives.

Recovering that information is a job that changes with scale. For a single URL you can read the full response line by line, nothing to interpret. Across an entire site you can’t read everything in full: you build a list and look for anomalies in it, and at that point the quality of the check depends on which tool built the list. What’s missing from both approaches is the trickiest part: what that address returns when someone other than you asks.

Reading headers with curl and DevTools

From the command line, the option that follows redirects shows the full sequence of responses.

curl -sIL https://www.esempio.it/vecchia-pagina/

Each block is a response: the code is on the first line and the header Location just below, and the number of blocks tells you how many steps there were. Without -L That behavior, for that matter, varies from program to program, which is why a single check isn’t enough. The browser follows automatically up to a maximum number of hops, beyond which it stops and reports an error.

doesn’t follow redirects unless you tell it to. An integration written by a partner might stop at the first 3xx and return that response as is, because whoever wrote it cared about the content, not the redirect: for that program, your address stopped returning usable data the moment you moved it. The curloption adds another variable, because it sends a -I request, which some servers treat differently from a HEAD . If the result looks off, it’s worth repeating with GETand comparing. -sL -o /dev/null -w "%{http_code}\n" and compare them.

In the browser you do the same work in the Network panel of the developer tools, with one thing to fix first: without Preserve log turned on, navigating clears the list, and the intermediate hop disappears before you can see it.

Redirect chains and loops

Nobody writes a redirect chain on purpose: it builds up one rule at a time, and each rule was correct when it was added. They pile up precisely because each one, on its own, bothers nobody: the extra hop costs a few milliseconds, shows up in no report, breaks nothing. The domain changes, then the structure changes, then two pages get merged, and a URL from three years ago reaches its destination through three responses written by different people at different times.

Google’s crawlers follow up to ten hops in a single crawl session, so a short chain won’t cost you indexing. The migration guide still recommends staying under three hops ideally, and under five in any case, and the reason has more to do with you than with Google: every hop is a point where something can break down the road, and the longer the chain, the less likely it is that whoever touches it next knows why each hop is there.

Keep in mind that Search Console’s URL Inspection tool doesn’t follow redirects: it tells you the address redirects, not where it ends up.

A loop is when the chain closes in on itself. The browser stops on its own after a certain number of attempts and shows a too-many-redirects error. The typical case is two rules written at different layers, one on the server and one in the application, that keep forwarding the request to each other.

Crawling the whole site

On a whole site, commands aren’t enough, and the right tool depends on what you’re looking for: the redirects already live on your site, the ones you’re about to publish, and the ones Google has actually encountered.

A link-based crawl finds the redirects you already have. SEOZoom’s SEO Spider crawls the site and sorts the URLs it finds into the Response Codes group, where the 3xx view shows the code and destination for each address. Those 3xx entries show up because the spider follows internal links, so each one is a link inside your own site that still points to the old address. These are the links Google asks you to update, the one part of the chain you have direct control over, and fixing them removes one hop from every visit. What this view doesn’t do is rebuild the full chain: it tells you where the starting URL leads, not where the destination leads next, and to find that out you have to query the destination in turn.

A pre-launch list calls for a different tool, one that accepts a list of URLs without crawling the site. Screaming Frog works in list mode, and the setting that makes this check useful is Always follow redirects, under the advanced tab. In list mode crawl depth is zero, meaning only the loaded URLs get called, but with that option on the tool follows each chain all the way to its final response, whether that’s a 2xx, 4xx, or 5xx. The All Redirects report then returns every address on your list along with its destination, the code each one returned, the number of hops, the chain type (HTTP redirect, JavaScript, or meta refresh meta refresh) and a column flagging whether it ends in a loop. This is the check that matters most before launch, because it compares what you wrote against what actually happens to each of the addresses you care about. For a handful of URLs, web-based redirect checkers are enough, and browser extensions that show the chain as you browse are handy for quick spot checks.

What Google has actually encountered, finally, shows up nowhere else: only Google can tell you that. In the Search Console Pages report, old addresses land under “Page with redirect,” and that list doesn’t describe your site, it describes what the search engine saw of it, with a sample of up to a thousand example URLs meant for spot-checking, not counting. An address that doesn’t appear there isn’t necessarily fine: it might just be one Google hasn’t revisited yet, or one it never saw at all.

Any crawl, finally, runs as a single caller: conditional redirects slip past it, and you reproduce them with curl’s -A option to change the user agent, or from the browser’s developer tools panel.

The most common mistakes

Almost every problem traces back to one of a handful of causes, and you isolate it by elimination, in an order worth sticking to. Start with who issued the response, because until you know that you’re fixing things at random: the code tells you curl, the layer responsible shows up in the origin’s logs, and if the request doesn’t appear there, something in front of it answered instead.

Next, check whether the rule was never reached or was reached and did something else, because these are failures with different causes. In the first case the old address returns a 200 as if nothing happened, and the problem is upstream: the order of rules in the file, a .htaccess that never gets read, a layer answering before it should. In the second case a 3xx does fire, but the code, the destination, or the number of hops don’t match what you wrote, and the cause lies in that rule or in a second rule further down the chain.

From symptom to cause, and how to confirm it
What you see First hypothesis How to confirm it
Nothing happens, the old page responds normally the rule never gets reached: file order, .htaccess ignored, a layer responding before it check AllowOverride, the order in the file, and the origin logs
The redirect fires but the status code is different from what you wrote an undeclared default in PHP or WordPress, or a second redirect downstream run curl -sIL and read every hop, not just the first one
You’re still landing on the old address even after the fix client cache from an earlier permanent redirect try it in an incognito window or with curl, which doesn’t have that cache
The destination is correct but there are three hops rules stacked on top of each other over time count the blocks in the curl -sIL response
The redirect leads to a page that returns 404 the destination changed after the rule was written crawl the site filtered on 3xx status codes, and check the destinations
It works for you but not for Google condition based on IP, language, or user agent reproduce the condition with curl -A

Malicious redirects on compromised sites

There’s a category of redirects you never wrote. Google’s spam policies classify these as deceptive and treat them as a violation, and on compromised sites they’re one of the most common symptoms: the site behaves normally for the person administering it and sends visitors arriving from search results somewhere else entirely.

The typical case triggers conditionally, for a specific user agent or for visitors coming from a specific referrer, precisely so it stays invisible to anyone checking. Search Console’s Security Issues report is the first place the warning shows up; from there you search through the theme files, the .htaccess, the database, and any externally loaded scripts, and you confirm it by reproducing the condition: if an address behaves differently depending on how you call it, and you didn’t write that difference in, you’ve found it.

After publishing

Once everything is live, the focus of the work shifts: you’re no longer checking your own rules, you’re checking whether the new address has actually replaced the old one for people you never notified. But that handover doesn’t happen on your schedule, it happens on the schedule of whoever revisits the page: Google says that for a medium-sized site it takes a few weeks for most pages to move over, and that ranking fluctuations are normal in the meantime while the site gets recrawled and reindexed. Anyone who reads the first week’s data as a final result ends up fixing rules that were never broken, and each fix adds one more hop to the ones already there.

You read progress in Search Console in reverse. Old addresses show up under “Page with redirect”, which in the Page indexing report sits among the non-indexed pages and isn’t an error but confirmation that the redirect has been seen; the sample list of URLs tops out at a thousand entries, so it tells you what’s happening, not how much is left. For that you need the two sitemaps, compared in the report between the old site’s indexed pages and the new site’s.

How long to keep redirects active

A redirect has no technical expiration date: it stays in place until someone removes it, and removing it is the one task nobody ever puts on the calendar. Google’s guidance is to keep it up for as long as possible, generally at least a year, and from a visitor’s point of view to treat it as permanent.

That same guide, though, includes a warning that changes how you read it: redirects are slow for visitors, so you should update your own links instead of letting everything pass through them. The two pieces of advice don’t contradict each other, because they’re talking about different kinds of references, and that distinction is what decides the matter: the rule stays active for references you can’t fix – links on other people’s sites, bookmarks, sent emails – while anything you control should be rewritten to point to the final destination.

A year on, then, the question isn’t whether to remove redirects but which ones, and you answer it address by address, using checks you already ran during the migration. In the origin server’s logs, that old URL is still getting hits, and from whom: if they come from people or from crawlers, the rule is doing its job. Outside the site, it’s still attracting links, and from pages that matter: if so, removing the rule turns those links into 404s with your domain in front of them. Inside the site, there’s still a link pointing to it: in that case, it’s not the rule that needs to go, it’s the link that needs rewriting.

There’s also a reason that didn’t exist a few years ago, and one you can actually measure. In the Vercel and MERJ study, 34.82% of ChatGPT’s requests and 34.16% of Claude’s end up on pages that don’t exist, and another 14.36% of ChatGPT’s requests gets spent following redirects. These crawlers work off inventories of old addresses: the share of their requests that doesn’t die in a 404 only reaches its destination because one of your rules is still active, and the day you remove it, those requests move into the first group. The same goes for AI-generated answers that cite you: the address that shows up in the citation may be the one collected before the change.

Keeping redirects active isn’t free for you either. Every rule is one more case to account for when you write a new one, and a long configuration is a configuration nobody reads again: that’s how chains happen. The difference between a system and a pile-up comes down to this – what stays are the rules someone still calls, not the ones nobody bothers to check.

Measuring the effect on performance

Comparing before and after runs into a problem created by the move itself: the addresses aren’t the same anymore, so a URL-by-URL comparison pits two lists against each other that don’t line up. What you can still compare are the sets – the keywords the site ranks for and the pages holding those rankings, before and after. That’s what Period Comparison gives you: it lines up two time ranges and separates what was gained from what was lost, with one limit worth keeping in mind: five thousand keywords per period, which on a large site means you’re reading the top slice, not the whole set. For addresses that stayed the same, Page Performance gives you the Search Console detail. Here too, read the numbers for what they are: an estimate of search presence, useful for comparing two moments on the same site, not for predicting how many clicks will show up.

If visibility drops, the move itself isn’t the cause, since permanent redirects don’t cost you PageRank, and looking there wastes time. What’s left are the addresses that stayed uncovered and destinations that were chosen poorly, which are also the only things you can still fix.

The redirect that stays

A redirect has done its job once the only references still passing through it are the ones you can’t fix. At that point, you don’t remove it because it’s no longer needed: you keep it because it’s useful to someone you don’t know, and because removing it costs them, not you.

It’s the part of the site that keeps responding when everything else has changed. It’s worth knowing what it says.

The market will not wait.
Take control now.

The only platform to hold your ground on Google and AI engines.

  • Full
    platform (trial included)
  • Strategy demo
    with an expert
  • Support
    in Italian