Skip to content

Be entity-first in SEO localization

featured image: screenshot from Google Cloud NLP

SEO and localization often consist in selecting optimal target keywords based on competitor rankings in local search markets, as well as a checklist of optimizations that go under the banner of “international SEO”. These are important steps to take, to be sure. But there are more opportunities for search optimization which should be considered in SEO localization. In particular, optimized entity localization is a must for executing SEO localization strategies which can compete in target markets.

In fact, keyword localization without entity localization is a critical but insufficient step effective SEO localization. It suggests to Google the solution for a surface-level need (“please rank me for this keyword”), but, especially for longer form content, but does not go far enough to actually produce the optimized content.

In this post, I will describe a rankings-based method for entity optimization in localization which leverages competitor or SERP data. In a later post, I will show you a similar method, but leveraging Wikipedia instead of SERP data. While the rankings-based method is more useful for reproducing competitor rankings, the Wikipedia method (or “knowledge base method”) can be used to audit an existing glossary or set of terms without reference to ranking pages. While the former is more useful for explicit SEO purposes; the latter can be useful for reliable glossary localization at scale, with likely benefits for SEO as well.

In both posts I argue that because of the rise of semantic search, the glossary is now the most important part of SEO localization beyond keyword localization. As the site of your most regularly used business-critical terminology, industry terms and phrases of art, product names, and more, the localization glossary is a repository for important semantic data.

Understanding Entities in Search

An “entity” simply describes any “thing”. In the context of search, an entity refers to a semantic, or meaningful object. These objects are connected in human-curated knowledge graphs like WikiData, or automatically produced knowledge graphs, like the Google Knowledge Graph, and through these connections, take on meaning.

Screenshot of a knowledge graph by Ontotext
Knowledge Graph basics, by Ontotext

The most visible example of an entity in search is the “Knowledge Panel”. We often encounter cases where Google returns a Knowledge Panel for a given search, generally when your query is more or less an entity without much modification (like “Cologne Dome” vs “how do I get to the Cologne Dome”?) or when a question is asked for which the entity is itself the answer or information about the entity is the answer.

Screenshot of a Knowledge Panel featuring the Cologne Cathedral showing in the Google Search results
Google’s Knowledge Panel is returned for millions of searches of places, people, and other named things.

In reality, this is a rather limited exposure of entities in search. Knowledge Panels represent a small subset of results from the “Topic Layer” of Google’s Knowledge Graph. The Google Knowledge Graph itself contains many billion entities, connected by many more “facts”.

Entities also work on a much more abundant level, as objects that can be detected within a sentence by natural language understanding tools. Such tools identify instances of these entities, and attempt to understand the role they play in a text and the importance of different entities to the “aboutness” of a given text. This understanding in turn has increasingly come to supplement classic keyword-based search, which in. the past ranked pages on co-occurrence of certain words among specific sets of documents (among other ranking factors).

Entity optimization based on competitor rankings

Assume you need to optimize localized contents in the fintech sector. Assembling keyword gap data should be fairly straightforward, provided you know who you are competing with in the target market. In this case, let’s say you are competing with N26. To understand what sort of traffic N26 brings in for the content they publish, you check (non-brand only) Semrush data for their blog:

Screenshot of keywords for N26's blog in Semrush
Keyword data is the first layer of a market or competitor gap analysis in local markets

What that data tells you is which keywords your competitor is currently ranking for in the target language, and what you can attempt to rank for in that language. But understanding how they are ranking means understanding a bit more about their language choices more broadly, and specifically, which entities they mention in their text and how they mention them. 

To get a sense for how entities work in this context, well beyond the keyword you will use to target a page, compare to the case of a randomly chosen article from the N26 blog:

Screenshot of a blog post by N26

Run a section of this text through Google Cloud NLP, a natural language understanding tool, and it recognizes some of the following entities in the text, scoring salience as a measure of the entity’s importance to the text.

Screenshot of an analysis of a blog post by N26 in Google Cloud NLP
NLU tools return insights like salience, sentiment, nearest category, etc.

I highlighted some of these entities in purple: these are examples of terms that your PM, linguist, or internal language stakeholder may find relevant enough and abundant enough throughout your content to warrant inclusion in your glossary.

How to do this yourself

Make sure you have a sense for which competitor pages you want to look at. While there are tools that you can use which will process many URLs at once, a sampling of pages should do. The reason is that most of the entities you will want to eventually optimize will probably appear many times over a set of documents. If you do have the resources to consider shaping more changes to your glossary and want to aim for maximal coverage of entities across competitor pages, I suggest using something like the relevance of a detected entity to a page, but more on that below. 

The important thing therefore is just to extract a set of representative pages across a few different categories or page types (for example, sampling from evergreen content marketing, product marketing, industry trends, etc.).

You can use any of the following tools to extract entities from a single or several URLs in target language. Note that not all tools have total coverage in all languages (Meaningcloud for example is missing DE):

Here, I am using entitieschecker to extract entities from the n26 blog post referenced above.

Screenshot of results from the tool Entitieschecker

Et voilà, it returns a csv-exportable list of entities detected on the page, along with their corresponding pages.

How to use this export?

I suggest looking at relevance scores for entities across a number of pages, and grouping these pages together on the basis of keyword group, content theme, content type etc.. The reason this is valuable is that these results are indications of certain entities which are part of the fabric of the text which ranks for keywords you would like to target.

Which terms you actually choose for inclusion will depend on the state of your glossary, but I suggest being generous with producing new entries: after all, glossaries are there to control your vocabulary for important terms, and gaining consistency and standardization for important in-language terms is basically all upside.

Entity Optimization as an SEO solution

While the process I have outlined is not quite the same as full-text optimization, it’s a strong and for many cases much more achievable alternative. Given the long and complex series of handoffs that make up localization workflows, incorporating a content brief or expecting a linguist to work in the editor of a content optimization tool is generally not viable. Entity optimization for your glossary, however, allows you to move beyond keyword optimization while producing content which ensures that the topical relevancy inherent in your source content is vouchsafed into target content.

Curious about Glossary Optimization? Get in Touch!

You agree to receive email communication from us by submitting this form and understand that your contact information will be stored with us.