hreflang has an awkward property: when it breaks, nothing tells you. There is no penalty and no error message — Search Console stopped reporting hreflang problems when Google retired its International Targeting report in 2022. The annotations are simply ignored, and your markets start competing with each other. In Ahrefs’ study of 374,756 domains that use hreflang, 67% had at least one issue.
What hreflang does, and what it doesn’t
hreflang tells Google that several URLs are versions of the same page for different languages or regions, so it can show each searcher the one that fits. When it works, Google swaps the URL in the result: the British searcher gets the British page, the German searcher the German one. John Mueller of Google describes it in exactly those terms — hreflang will often “swap out the URL”.
Most confusion comes from expecting it to do more:
- It doesn’t tell Google what language a page is in. Google says it uses neither hreflang nor the HTML
langattribute for that. It reads the visible content. - It doesn’t guarantee indexing. “hreflang doesn’t guarantee indexing”, as Mueller put it in May 2025. Every version still has to be crawled and indexed on its own.
- It doesn’t fix duplicate content between languages, because translated pages aren’t duplicates in the first place. Google only treats localised versions as duplicates when the main content is left untranslated.
- It doesn’t make a page rank higher. It changes which of your URLs appears for a query, not whether the page deserves the position.
- It isn’t read the same way everywhere. Yandex documents the same link element. Bing is the exception: Microsoft’s Fabrice Canel said in 2020 that hreflang is “a far weaker signal than content-language” there.
What an annotation looks like
Every version of a page lists every version of that page, itself included. For a pricing page with British, American and general English versions and a German one, the head of all four pages carries the same five lines:
<link rel="alternate" hreflang="en-gb" href="https://example.com/uk/pricing/"> <link rel="alternate" hreflang="en-us" href="https://example.com/us/pricing/"> <link rel="alternate" hreflang="en" href="https://example.com/en/pricing/"> <link rel="alternate" hreflang="de" href="https://example.com/de/preise/"> <link rel="alternate" hreflang="x-default" href="https://example.com/en/pricing/">
The value is a language code, optionally followed by a region: de is German for everyone, en-gb is English for users in the United Kingdom. Google supports languages from ISO 639-1 and regions from ISO 3166-1 alpha-2, plus a script from ISO 15924 where it matters — zh-Hant for Traditional Chinese — and nothing else. Case makes no difference, because language tags are case-insensitive: en-GB and en-gb are the same annotation.
The codes that cause trouble are the ones that look right:
| Written | What it actually says |
|---|---|
en-UK | English, with the region ignored. UK is a reserved code, not the United Kingdom’s: that is GB. |
es-419 | Not supported. Google’s documentation names this Latin American Spanish code as one it doesn’t accept. |
es-LA | Spanish for users in Laos. |
be | Belarusian, not Belgium. The same trap turns be-fr into Belarusian for France. |
gb, mx | Nothing. A country code on its own isn’t a language, and Google doesn’t derive one from it. |
iw, in | Old codes for Hebrew and Indonesian, replaced by he and id. |
en-eu, en-ww | No usable region. EU is reserved and ignored; WW isn’t assigned at all. |
The hreflang generator on this site rejects the invalid codes and warns about valid ones that are usually a mistake, such as es-LA.
Three places to put it
Google reads hreflang from three places and treats them as equivalent:
- The HTML head, as
linkelements. The easiest to inspect, and more fragile than it looks: Google stops reading the head at the first element that doesn’t belong there, such as animgoriframe. When it meets one, its documentation says, it “assumes the end of the<head>element”. A tag manager that injects a tracking pixel above your annotations hides them from Google. - An HTTP
Linkheader on the response. The only option for files that have no head, such as PDFs. - The XML sitemap, with an
xhtml:linkchild for every version under the entry for every version. It keeps page heads short on sets with a hundred versions, and the child entries don’t count towards the sitemap’s URL limit. The cost is that nobody sees it when they look at a page, so mistakes last longer.
You can use more than one, but Google says there is “no benefit in Search” in doing so. Pick one. Two methods that coexist have to agree forever, and in practice they drift apart: two teams, two implementations, two answers to the same question.
One syntax detail: keep hreflang in its own link element. Google’s guidelines say not to combine it with other attributes, such as media, in a single tag.
The rules a set has to follow
- Every version lists every version, itself included. Google’s documentation is direct: “Each language version must list itself as well as all other language versions.” Ahrefs’ Patrick Stox calls the self-reference more of a best practice than a requirement. It costs one line, so there is no reason to find out.
- Every pair points both ways. “If two pages don’t both point to each other, the tags will be ignored” — a safeguard, so that another site can’t declare itself a version of yours. Google applies this pair by pair and “will still process the ones that point to each other”, so one missing return tag breaks one pair, not the whole set.
- URLs are absolute.
https://example.com/de/, never/de/or//example.com/de/. - Every URL is final and canonical. It answers 200, doesn’t redirect, isn’t
noindex, and names itself as canonical. More on that below. - One URL per code, with a catch-all where you split a language. If you have
en-gb,en-usanden-au, add a plainenfor English speakers everywhere else. Google recommends it, and it can be one of the regional pages. - x-default is optional. It names the page for everyone your set doesn’t match. Where to point it is covered in a separate note.
- Versions can live on different domains. A set can span
example.deandexample.fr. The rules don’t change.
hreflang and canonical tags
This is where implementations that look correct most often go wrong. Both tags describe relationships between URLs, and nothing stops them from disagreeing.
Each version is its own canonical. Google’s documentation on canonical URLs says that if you use hreflang, you should “specify a canonical page in the same language”, or the best substitute language if there isn’t one. The common mistake is the reverse: every translation canonicalised to the English page, usually by a template or a plugin. Mueller spelled out the result in 2015: don’t make one language version the canonical for another, “otherwise we won’t index the ‘en’ version”.
hreflang lists canonical URLs. In the same post, Mueller wrote that if the canonical URL isn’t part of the hreflang pairs, the hreflang markup is ignored. The post is old, and his own archive marks it as such, but Google’s current documentation draws the same connection from the other side: it warns that listing HTTP URLs in hreflang annotations, instead of the HTTPS ones, can lead Google to choose the HTTP page as canonical.
The pull works in both directions. Google’s canonicalisation documentation, last updated in July 2026, says that “for canonicalization purposes Google prefers URLs that are part of hreflang clusters”. Its example: if German pages for Germany and Switzerland point to each other and neither lists the Austrian page, Google prefers the German and Swiss pages as canonicals over the Austrian one. A version left out of a set can lose the canonical choice to one that’s in it.
Versions in the same language can still be folded together. If your en-gb and en-us pages are near-identical, Google may choose one as the canonical for both. Mueller explained what happens next in May 2025: hreflang will often still swap the URL shown, “but reporting will be on the canonical URL”. British searchers can land on the British page while Search Console reports their clicks under the American one. Google’s advice for regional variants in one language is to use canonicalisation and hreflang together; the way to give it a reason to keep both pages is to make them genuinely different — prices, currency, delivery, contact details.
hreflang doesn’t belong on the canonical link. Google doesn’t use a rel="canonical" link that carries an hreflang attribute for canonicalisation at all. Use a separate rel="alternate" link.
A real case shows how quietly the two can drift apart. In May 2025 an SEO asked Mueller on Bluesky why a client’s Belgian French pages were appearing in searches from France despite hreflang. Mueller looked into it, and the cause was the host name: the annotations referenced www. URLs, but the Belgian site was also indexed without www. hreflang pointed at addresses Google wasn’t using.
To see what Google actually chose, run each version through URL Inspection in Search Console. It shows the canonical you declared next to the one Google selected. Where they differ, your hreflang points at a URL Google isn’t using.
How often hreflang breaks
Two large studies give a sense of scale. In August 2023, Ahrefs published an analysis of issues on 374,756 domains using hreflang: 67% had at least one.
| Issue (Ahrefs, 2023) | Domains |
|---|---|
| Pages missing x-default | 56.3% |
| Pages missing self-referencing tags | 18% |
| Tags pointing at redirected or broken pages | 16.9% |
| Pages missing return tags | 15.3% |
| Tags pointing at non-canonical URLs | 8% |
| Incorrect hreflang values | 4.6% |
HTML lang attribute contradicting hreflang | 3.2% |
| More than one page for the same language | 2.5% |
| The same page for more than one language | 2.5% |
A missing x-default tops the list but breaks nothing, since the value is optional. Among the problems that stop annotations working, redirected or broken targets and missing return tags are the largest.
The second study, by Dan Taylor of SALT.agency, published in Search Engine Land in April 2023, read the HTML of 18,786 sites that use hreflang. It found conflicting hreflang directives on 31% of sites serving several languages, no self-referencing tags in 16% of hreflang clusters, and unknown language codes on 9% of multilingual sites.
What 50 brand home pages showed
Both studies are three years old, so on 16 September 2026 I ran my own check. I requested the home pages of 110 international brands with the code behind this site’s hreflang checker, then every version each home page declared — 1,579 URLs, once each — and checked that each answered 200 without redirecting, wasn’t noindex, named itself as canonical and listed the home page back.
Thirty-eight home pages couldn’t be read, most of them behind bot checks or timing out. Another 22 had no hreflang in their HTML or headers: several served a local page to a request from Poland, and some may keep their annotations in a sitemap, which this check doesn’t read. That left 50 sets, with a median of 21 annotations and a maximum of 137.
| Finding (the rows overlap) | Sets, of 50 |
|---|---|
| Passed every check | 26 |
| A version redirects permanently (301) | 8 |
| A return tag or self-reference problem | 7 |
| An invalid language or region code | 6 |
| A version canonicalised to another URL | 4 |
| A version returns an error (404 or 410) | 3 |
A version is noindex | 3 |
| Relative URLs | 1 |
| No x-default (optional, not counted as a problem) | 10 |
Eighteen of the 50 sets had at least one of those problems, on the page most brands watch most closely. Six more were clean apart from temporary redirects (302 or 307), which can depend on where a request comes from, so I didn’t count them. Nor did I count a version with no annotations in its HTML as a missing return tag, because its set might live in a sitemap.
The patterns were more instructive than the counts:
- Codes written backwards. A hardware maker’s set wrote every code country first:
py-es,uy-es,gb-en. Most of those are simply invalid. The dangerous ones are valid by accident —be-fris Belarusian for France,ar-esArabic for Spain,br-ptBreton for Portugal — and no validator flags a real code. - A home page outside its own set. A software company’s English home page listed eleven other languages but not itself, and none of the eleven listed English. Every other language paired up. English didn’t.
- A whole set pointing at the wrong English URL. A payments company’s home page listed itself for
en,en-GBand x-default. Its 42 other versions listed a/gb/page for those codes instead — and that page names the home page as its canonical. - Moves nobody told the set about. A fitness-device maker had moved thirteen versions to local domains with permanent redirects, and its home page still listed the old addresses; the local sites carried smaller sets of their own that left the home page out. A project-management tool listed every version on a
www.host that permanently redirects to the bare domain. - Invented regions.
es-SPandes-CEfor a retailer’s Balearic and Canary Islands sites,en-wwfor “worldwide English” on a set with no x-default,zh-FTon a language app’s Chinese version. hreflang has no codes for parts of a country or for the whole world, and the whole world is what x-default is for. - Codes from older software. The same fitness-device maker listed Hebrew and Indonesian twice, under
heandidand under the oldiwandin, each pointing at a different URL. Those old forms are what Java’sLocaleclass produced for years: until Java SE 17 it convertedhe,yiandidtoiw,jiandin.
The usual caveats apply. The requests came from Poland, without an Accept-Language header and with an identified user agent, and a site can answer Googlebot differently. The results describe what the pages served that day, not how Google processed them.
The eleven mistakes, and the fixes
- 1. Missing return tagsA version whose set leaves out a page that lists it — usually a market on a different platform, or a template that builds each set from its own market’s list. Fix: generate every set from one shared table of equivalent URLs.
- 2. Annotations pointing at redirectsThe residue of a migration, a domain move or an HTTPS switch. Fix: regenerate every set after any URL change, with final URLs only.
- 3. Annotations pointing at errors or noindex pagesA version was removed or blocked, and the others still list it. Fix: take it out of every set, or restore it.
- 4. Canonicals that contradict hreflangA translation canonicalised to another language, or a set listing URLs that canonicalise elsewhere. Fix: every version self-canonical, every hreflang URL a canonical one.
- 5. Invalid or misleading codes
en-UK,es-419, codes written country first, old language codes. Fix: validate against ISO 639-1 and ISO 3166-1, language first. - 6. No self-referenceFix: add it. One line per page.
- 7. Relative URLsFix: absolute URLs, with the protocol, in every annotation.
- 8. Annotations Google never readsTags placed below an image or iframe in the head, where Google has already stopped reading. Fix: keep those elements out of the head, or below the annotations.
- 9. Two methods that disagreeA head set and a sitemap set maintained by different people. Fix: keep one.
- 10. Redirecting visitors by language or locationGoogle advises against automatically redirecting users between language versions. Googlebot’s default IP addresses are in the US and it sends no
Accept-Languageheader, so a redirecting page may only ever show it one version. Fix: separate URLs, annotated, with a visible language switch. - 11. Regional versions without regional contentNothing breaks, and it costs the most. Fix: split a language by region only when the content has to differ.
It is the most expensive mistake that breaks nothing. One English version can rank in the UK, the US, Ireland and Australia at once. Splitting it into four regional copies multiplies production and maintenance, spreads links and signals across URLs Google may fold back into one canonical, and multiplies every other mistake on this list by four.
Split a language by region when the content has to be different — prices, availability, regulation, delivery — and not before.
How to audit hreflang
Checking pages one at a time finds invalid codes and relative URLs, but not missing return tags, because a return tag only exists between two pages. The method that works is the matrix:
Crawl every market
With a crawler that extracts hreflang, such as Screaming Frog or Sitebulb — from the sitemaps too, if that is where the annotations live.
Build the matrix
One row per URL, one column per version it declares. You are looking for asymmetric cells: A lists B, and B doesn’t list A.
Check every target
Each listed URL must answer 200, without a redirect and without
noindex.Cross-check canonicals
Each version names itself as canonical. URL Inspection shows whether Google agrees.
Validate every code
Against ISO 639-1 and ISO 3166-1, not against intuition.
en-UKlooks reasonable and isn’t.Watch results by country
In Search Console’s Performance report, filter by country and look at which URLs collect the impressions. If the American page gets the British impressions, the swap isn’t happening, or Google is reporting both versions under one canonical.
For sets of up to 25 pages, the hreflang checker builds the matrix for you: give it one URL and it fetches every version, and every version those declare. The hreflang generator writes a complete, validated set. On a site with thousands of URLs the matrix has to be rebuilt on a schedule, which is why I automate it.
What to expect after fixing it
Fixing hreflang rarely produces a traffic jump. What changes is where people land: the British page for British searchers, the right currency, the right delivery terms. It shows in conversion rates more than in sessions.
If someone sells you an hreflang fix as a traffic multiplier, be sceptical. It is hygiene — cheap, necessary and unspectacular. Growth in a new market comes from elsewhere, usually from having enough content there, researched for that market.
How this site does it
jorgehorst.com has an English and a Spanish section. Every page with an equivalent in the other language carries three annotations in its head — en, es and x-default, with x-default on the English URL — and lists itself. Pages without an equivalent carry none, and there is no hreflang in the sitemaps, so the head is the only place to look. This guide and its Spanish counterpart are one of twenty pairs.
Sources
- Google Search Central, Tell Google about localized versions of your page, last updated 22 December 2025.
- Google Search Central, How to specify a canonical with rel="canonical" and other methods, last updated 10 July 2026, and What is URL canonicalization, last updated 20 August 2026.
- Google Search Central, Managing multi-regional and multilingual sites, How Google crawls locale-adaptive pages and Use valid HTML to specify page metadata, all last updated 10 December 2025.
- Search Console Help, The International Targeting report is deprecated, and Search Engine Land, Google Search Console to remove International Targeting report, 24 August 2022.
- John Mueller on Bluesky, 9 May 2025, replying to a question about Belgian pages whose thread records the cause on 21 May 2025.
- John Mueller, hreflang canonical, a Google+ post of 7 November 2015, republished on his site.
- Patrick Stox, Over 67% of domains using hreflang have issues, Ahrefs, 10 August 2023.
- Dan Taylor, Study: 31% of international websites contain hreflang errors, Search Engine Land, 4 April 2023.
- Barry Schwartz, Bing says hreflang a weak signal for its search engine, Search Engine Roundtable, 11 September 2020, quoting Fabrice Canel of Microsoft.
- Yandex Webmaster, Indexing localized pages.
- IETF, RFC 5646, Tags for Identifying Languages, section 2.1.1, and Oracle, java.util.Locale in Java SE 17, on legacy language codes.
- Home pages of 110 international brands and the 1,579 URLs they declared, requested once each from Poland on 16 September 2026 with this site’s hreflang inspector: 50 readable sets.
Questions
link elements, in an HTTP Link header for files such as PDFs, or in the XML sitemap. Google treats the three as equivalent and says using more than one brings no benefit in Search, so pick one.