hreflang: a practical guide

International SEO 16 September 2026·18 min read

hreflang has an awkward property: when it breaks, nothing tells you. There is no penalty and no error message — Search Console stopped reporting hreflang problems when Google retired its International Targeting report in 2022. The annotations are simply ignored, and your markets start competing with each other. In Ahrefs’ study of 374,756 domains that use hreflang, 67% had at least one issue.

What hreflang does, and what it doesn’t

hreflang tells Google that several URLs are versions of the same page for different languages or regions, so it can show each searcher the one that fits. When it works, Google swaps the URL in the result: the British searcher gets the British page, the German searcher the German one. John Mueller of Google describes it in exactly those terms — hreflang will often “swap out the URL”.

Most confusion comes from expecting it to do more:

  • It doesn’t tell Google what language a page is in. Google says it uses neither hreflang nor the HTML lang attribute for that. It reads the visible content.
  • It doesn’t guarantee indexing. “hreflang doesn’t guarantee indexing”, as Mueller put it in May 2025. Every version still has to be crawled and indexed on its own.
  • It doesn’t fix duplicate content between languages, because translated pages aren’t duplicates in the first place. Google only treats localised versions as duplicates when the main content is left untranslated.
  • It doesn’t make a page rank higher. It changes which of your URLs appears for a query, not whether the page deserves the position.
  • It isn’t read the same way everywhere. Yandex documents the same link element. Bing is the exception: Microsoft’s Fabrice Canel said in 2020 that hreflang is “a far weaker signal than content-language” there.

What an annotation looks like

Every version of a page lists every version of that page, itself included. For a pricing page with British, American and general English versions and a German one, the head of all four pages carries the same five lines:

<link rel="alternate" hreflang="en-gb" href="https://example.com/uk/pricing/">
<link rel="alternate" hreflang="en-us" href="https://example.com/us/pricing/">
<link rel="alternate" hreflang="en" href="https://example.com/en/pricing/">
<link rel="alternate" hreflang="de" href="https://example.com/de/preise/">
<link rel="alternate" hreflang="x-default" href="https://example.com/en/pricing/">

The value is a language code, optionally followed by a region: de is German for everyone, en-gb is English for users in the United Kingdom. Google supports languages from ISO 639-1 and regions from ISO 3166-1 alpha-2, plus a script from ISO 15924 where it matters — zh-Hant for Traditional Chinese — and nothing else. Case makes no difference, because language tags are case-insensitive: en-GB and en-gb are the same annotation.

The codes that cause trouble are the ones that look right:

WrittenWhat it actually says
en-UKEnglish, with the region ignored. UK is a reserved code, not the United Kingdom’s: that is GB.
es-419Not supported. Google’s documentation names this Latin American Spanish code as one it doesn’t accept.
es-LASpanish for users in Laos.
beBelarusian, not Belgium. The same trap turns be-fr into Belarusian for France.
gb, mxNothing. A country code on its own isn’t a language, and Google doesn’t derive one from it.
iw, inOld codes for Hebrew and Indonesian, replaced by he and id.
en-eu, en-wwNo usable region. EU is reserved and ignored; WW isn’t assigned at all.

The hreflang generator on this site rejects the invalid codes and warns about valid ones that are usually a mistake, such as es-LA.

Three places to put it

Google reads hreflang from three places and treats them as equivalent:

  • The HTML head, as link elements. The easiest to inspect, and more fragile than it looks: Google stops reading the head at the first element that doesn’t belong there, such as an img or iframe. When it meets one, its documentation says, it “assumes the end of the <head> element”. A tag manager that injects a tracking pixel above your annotations hides them from Google.
  • An HTTP Link header on the response. The only option for files that have no head, such as PDFs.
  • The XML sitemap, with an xhtml:link child for every version under the entry for every version. It keeps page heads short on sets with a hundred versions, and the child entries don’t count towards the sitemap’s URL limit. The cost is that nobody sees it when they look at a page, so mistakes last longer.

You can use more than one, but Google says there is “no benefit in Search” in doing so. Pick one. Two methods that coexist have to agree forever, and in practice they drift apart: two teams, two implementations, two answers to the same question.

One syntax detail: keep hreflang in its own link element. Google’s guidelines say not to combine it with other attributes, such as media, in a single tag.

The rules a set has to follow

  • Every version lists every version, itself included. Google’s documentation is direct: “Each language version must list itself as well as all other language versions.” Ahrefs’ Patrick Stox calls the self-reference more of a best practice than a requirement. It costs one line, so there is no reason to find out.
  • Every pair points both ways. “If two pages don’t both point to each other, the tags will be ignored” — a safeguard, so that another site can’t declare itself a version of yours. Google applies this pair by pair and “will still process the ones that point to each other”, so one missing return tag breaks one pair, not the whole set.
  • URLs are absolute. https://example.com/de/, never /de/ or //example.com/de/.
  • Every URL is final and canonical. It answers 200, doesn’t redirect, isn’t noindex, and names itself as canonical. More on that below.
  • One URL per code, with a catch-all where you split a language. If you have en-gb, en-us and en-au, add a plain en for English speakers everywhere else. Google recommends it, and it can be one of the regional pages.
  • x-default is optional. It names the page for everyone your set doesn’t match. Where to point it is covered in a separate note.
  • Versions can live on different domains. A set can span example.de and example.fr. The rules don’t change.
A four-by-four matrix of the pages example.com/en/, /de/, /es/ and /fr/ against the versions each one declares. Every cell is filled except one: the French page does not declare the German page. A note explains that German and French are therefore not a pair, so Google ignores the annotation between them and Search Console reports nothing, while every other pair, such as English and German, still works, because Google processes the pairs that point to each other.
The error that matters most can’t be seen from a single page. It only appears when you compare what each page declares with what the others declare back.

hreflang and canonical tags

This is where implementations that look correct most often go wrong. Both tags describe relationships between URLs, and nothing stops them from disagreeing.

Each version is its own canonical. Google’s documentation on canonical URLs says that if you use hreflang, you should “specify a canonical page in the same language”, or the best substitute language if there isn’t one. The common mistake is the reverse: every translation canonicalised to the English page, usually by a template or a plugin. Mueller spelled out the result in 2015: don’t make one language version the canonical for another, “otherwise we won’t index the ‘en’ version”.

hreflang lists canonical URLs. In the same post, Mueller wrote that if the canonical URL isn’t part of the hreflang pairs, the hreflang markup is ignored. The post is old, and his own archive marks it as such, but Google’s current documentation draws the same connection from the other side: it warns that listing HTTP URLs in hreflang annotations, instead of the HTTPS ones, can lead Google to choose the HTTP page as canonical.

The pull works in both directions. Google’s canonicalisation documentation, last updated in July 2026, says that “for canonicalization purposes Google prefers URLs that are part of hreflang clusters”. Its example: if German pages for Germany and Switzerland point to each other and neither lists the Austrian page, Google prefers the German and Swiss pages as canonicals over the Austrian one. A version left out of a set can lose the canonical choice to one that’s in it.

Versions in the same language can still be folded together. If your en-gb and en-us pages are near-identical, Google may choose one as the canonical for both. Mueller explained what happens next in May 2025: hreflang will often still swap the URL shown, “but reporting will be on the canonical URL”. British searchers can land on the British page while Search Console reports their clicks under the American one. Google’s advice for regional variants in one language is to use canonicalisation and hreflang together; the way to give it a reason to keep both pages is to make them genuinely different — prices, currency, delivery, contact details.

hreflang doesn’t belong on the canonical link. Google doesn’t use a rel="canonical" link that carries an hreflang attribute for canonicalisation at all. Use a separate rel="alternate" link.

A real case shows how quietly the two can drift apart. In May 2025 an SEO asked Mueller on Bluesky why a client’s Belgian French pages were appearing in searches from France despite hreflang. Mueller looked into it, and the cause was the host name: the annotations referenced www. URLs, but the Belgian site was also indexed without www. hreflang pointed at addresses Google wasn’t using.

To see what Google actually chose, run each version through URL Inspection in Search Console. It shows the canonical you declared next to the one Google selected. Where they differ, your hreflang points at a URL Google isn’t using.

How often hreflang breaks

Two large studies give a sense of scale. In August 2023, Ahrefs published an analysis of issues on 374,756 domains using hreflang: 67% had at least one.

Issue (Ahrefs, 2023)Domains
Pages missing x-default56.3%
Pages missing self-referencing tags18%
Tags pointing at redirected or broken pages16.9%
Pages missing return tags15.3%
Tags pointing at non-canonical URLs8%
Incorrect hreflang values4.6%
HTML lang attribute contradicting hreflang3.2%
More than one page for the same language2.5%
The same page for more than one language2.5%

A missing x-default tops the list but breaks nothing, since the value is optional. Among the problems that stop annotations working, redirected or broken targets and missing return tags are the largest.

The second study, by Dan Taylor of SALT.agency, published in Search Engine Land in April 2023, read the HTML of 18,786 sites that use hreflang. It found conflicting hreflang directives on 31% of sites serving several languages, no self-referencing tags in 16% of hreflang clusters, and unknown language codes on 9% of multilingual sites.

What 50 brand home pages showed

Both studies are three years old, so on 16 September 2026 I ran my own check. I requested the home pages of 110 international brands with the code behind this site’s hreflang checker, then every version each home page declared — 1,579 URLs, once each — and checked that each answered 200 without redirecting, wasn’t noindex, named itself as canonical and listed the home page back.

Thirty-eight home pages couldn’t be read, most of them behind bot checks or timing out. Another 22 had no hreflang in their HTML or headers: several served a local page to a request from Poland, and some may keep their annotations in a sitemap, which this check doesn’t read. That left 50 sets, with a median of 21 annotations and a maximum of 137.

Finding (the rows overlap)Sets, of 50
Passed every check26
A version redirects permanently (301)8
A return tag or self-reference problem7
An invalid language or region code6
A version canonicalised to another URL4
A version returns an error (404 or 410)3
A version is noindex3
Relative URLs1
No x-default (optional, not counted as a problem)10

Eighteen of the 50 sets had at least one of those problems, on the page most brands watch most closely. Six more were clean apart from temporary redirects (302 or 307), which can depend on where a request comes from, so I didn’t count them. Nor did I count a version with no annotations in its HTML as a missing return tag, because its set might live in a sitemap.

The patterns were more instructive than the counts:

  • Codes written backwards. A hardware maker’s set wrote every code country first: py-es, uy-es, gb-en. Most of those are simply invalid. The dangerous ones are valid by accident — be-fr is Belarusian for France, ar-es Arabic for Spain, br-pt Breton for Portugal — and no validator flags a real code.
  • A home page outside its own set. A software company’s English home page listed eleven other languages but not itself, and none of the eleven listed English. Every other language paired up. English didn’t.
  • A whole set pointing at the wrong English URL. A payments company’s home page listed itself for en, en-GB and x-default. Its 42 other versions listed a /gb/ page for those codes instead — and that page names the home page as its canonical.
  • Moves nobody told the set about. A fitness-device maker had moved thirteen versions to local domains with permanent redirects, and its home page still listed the old addresses; the local sites carried smaller sets of their own that left the home page out. A project-management tool listed every version on a www. host that permanently redirects to the bare domain.
  • Invented regions. es-SP and es-CE for a retailer’s Balearic and Canary Islands sites, en-ww for “worldwide English” on a set with no x-default, zh-FT on a language app’s Chinese version. hreflang has no codes for parts of a country or for the whole world, and the whole world is what x-default is for.
  • Codes from older software. The same fitness-device maker listed Hebrew and Indonesian twice, under he and id and under the old iw and in, each pointing at a different URL. Those old forms are what Java’s Locale class produced for years: until Java SE 17 it converted he, yi and id to iw, ji and in.

The usual caveats apply. The requests came from Poland, without an Accept-Language header and with an identified user agent, and a site can answer Googlebot differently. The results describe what the pages served that day, not how Google processed them.

The eleven mistakes, and the fixes

  • 1. Missing return tagsA version whose set leaves out a page that lists it — usually a market on a different platform, or a template that builds each set from its own market’s list. Fix: generate every set from one shared table of equivalent URLs.
  • 2. Annotations pointing at redirectsThe residue of a migration, a domain move or an HTTPS switch. Fix: regenerate every set after any URL change, with final URLs only.
  • 3. Annotations pointing at errors or noindex pagesA version was removed or blocked, and the others still list it. Fix: take it out of every set, or restore it.
  • 4. Canonicals that contradict hreflangA translation canonicalised to another language, or a set listing URLs that canonicalise elsewhere. Fix: every version self-canonical, every hreflang URL a canonical one.
  • 5. Invalid or misleading codesen-UK, es-419, codes written country first, old language codes. Fix: validate against ISO 639-1 and ISO 3166-1, language first.
  • 6. No self-referenceFix: add it. One line per page.
  • 7. Relative URLsFix: absolute URLs, with the protocol, in every annotation.
  • 8. Annotations Google never readsTags placed below an image or iframe in the head, where Google has already stopped reading. Fix: keep those elements out of the head, or below the annotations.
  • 9. Two methods that disagreeA head set and a sitemap set maintained by different people. Fix: keep one.
  • 10. Redirecting visitors by language or locationGoogle advises against automatically redirecting users between language versions. Googlebot’s default IP addresses are in the US and it sends no Accept-Language header, so a redirecting page may only ever show it one version. Fix: separate URLs, annotated, with a visible language switch.
  • 11. Regional versions without regional contentNothing breaks, and it costs the most. Fix: split a language by region only when the content has to differ.
About number 11, which costs money quietly

It is the most expensive mistake that breaks nothing. One English version can rank in the UK, the US, Ireland and Australia at once. Splitting it into four regional copies multiplies production and maintenance, spreads links and signals across URLs Google may fold back into one canonical, and multiplies every other mistake on this list by four.

Split a language by region when the content has to be different — prices, availability, regulation, delivery — and not before.

How to audit hreflang

Checking pages one at a time finds invalid codes and relative URLs, but not missing return tags, because a return tag only exists between two pages. The method that works is the matrix:

  1. Crawl every market

    With a crawler that extracts hreflang, such as Screaming Frog or Sitebulb — from the sitemaps too, if that is where the annotations live.

  2. Build the matrix

    One row per URL, one column per version it declares. You are looking for asymmetric cells: A lists B, and B doesn’t list A.

  3. Check every target

    Each listed URL must answer 200, without a redirect and without noindex.

  4. Cross-check canonicals

    Each version names itself as canonical. URL Inspection shows whether Google agrees.

  5. Validate every code

    Against ISO 639-1 and ISO 3166-1, not against intuition. en-UK looks reasonable and isn’t.

  6. Watch results by country

    In Search Console’s Performance report, filter by country and look at which URLs collect the impressions. If the American page gets the British impressions, the swap isn’t happening, or Google is reporting both versions under one canonical.

For sets of up to 25 pages, the hreflang checker builds the matrix for you: give it one URL and it fetches every version, and every version those declare. The hreflang generator writes a complete, validated set. On a site with thousands of URLs the matrix has to be rebuilt on a schedule, which is why I automate it.

What to expect after fixing it

Fixing hreflang rarely produces a traffic jump. What changes is where people land: the British page for British searchers, the right currency, the right delivery terms. It shows in conversion rates more than in sessions.

If someone sells you an hreflang fix as a traffic multiplier, be sceptical. It is hygiene — cheap, necessary and unspectacular. Growth in a new market comes from elsewhere, usually from having enough content there, researched for that market.

How this site does it

jorgehorst.com has an English and a Spanish section. Every page with an equivalent in the other language carries three annotations in its head — en, es and x-default, with x-default on the English URL — and lists itself. Pages without an equivalent carry none, and there is no hreflang in the sitemaps, so the head is the only place to look. This guide and its Spanish counterpart are one of twenty pairs.

Sources

Questions

An annotation that tells Google which URLs are language or regional versions of the same page, so it can show each searcher the version that fits. It doesn’t set a page’s language, guarantee indexing or improve rankings. It decides which of your URLs appears.
In the HTML head as link elements, in an HTTP Link header for files such as PDFs, or in the XML sitemap. Google treats the three as equivalent and says using more than one brings no benefit in Search, so pick one.
Yes. Each version names itself as canonical, and hreflang lists those canonical URLs. Never canonicalise a translation to a page in another language: Google may not index the version you canonicalised away.
Google ignores the annotation between the two pages that don’t point to each other. The pairs that do point to each other keep working. Search Console won’t tell you: it stopped reporting hreflang errors when the International Targeting report was retired in 2022.
Only weakly. Microsoft’s Fabrice Canel said in 2020 that hreflang is a far weaker signal than content-language for Bing. Yandex documents support for hreflang link elements.