Skip to content

Translation Memory

A translation memory (TM) is a database that is used in localization to store translated (or “target”) text and phrases, along with their corresponding original (or “source”) text. Translation memories are typically used in conjunction with computer-assisted translation (CAT) tools, which are software programs that aid linguists in text translation by providing machine translation options, presenting past translations from the TM and offering options to insert and edit these TM strings, and providing collaborative translation environments. 

Translation memories typically store source and target segments in XML versions known as Translation Memory Exchange (TMX) files. They are intended to be highly interoperable with other tools, and especially for ingestion into web applications.

TMs and String segmentation 

String segmentation is a process used in translation management systems (TMSs) to break down sentences or phrases into smaller segments, called “strings,” for translation. This is done in order to make the translation process more efficient and cost-effective, and is the primary use case for translation memories in enterprise settings.

After string segmentation, fuzzy matching is used to determine the degree to which a source string syntactically matches an early source string. Depending on the “match tier”, or the degree to which a source string being translated matches another source string stored in the TM, each translated word will be discounted a corresponding percentage for each billable word upon billing of the localization vendor.

When writers use computer-assisted translation (CAT) tools, they often encounter segments of text that have already been translated. These segments are known as “matches” and can be either exact or fuzzy matches. Exact matches are segments of text that have been previously translated and are identical to the source text. Fuzzy matches, on the other hand, are segments of text that have been previously translated, but may have some minor differences with the source text.

String segmentation becomes important in terms of cost savings when writers work with CAT tools because it allows them to leverage previously translated segments, which can save time and money. For example, if a writer is working on a document that has many exact and fuzzy matches, they can use the string segmentation feature in their TMS to quickly identify and insert the matching segments, rather than having to translate the entire document from scratch. This can significantly reduce the amount of time and effort required for the translation, which can lead to cost savings for the client.

Screenshot of table showing example of billing rates per word according to fuzzy match tier of untranslated source strings to source strings stored in a translation memory
Example of billing per word by TM fuzzy match tier

Translation Memories and SEO

Similar to glossaries, a TM’s primary advantage for organic search is its ability to provide a repository of linguistic choices which can be reused in the future. This can have the effect of producing a “controlled vocabulary”, which may be leveraged in support of wider content governance practices that can have a positive effect for SEO in target locales.

While linguistic alignment is not often discussed in connection with SEO and localization efforts, full-document optimization is a strategy pursued for content optimization, especially for longform content, and several tools such as MarketMuse and SurferSEO exist to serve that need. While TM doesn’t serve that need directly, by regularly producing localized text containing the sort of domain-specific vocabulary that is found in fully optimized documents, it can play a strong indirect or supporting function.