The Watchlist Search (Customers) page lets you screen any name, country, free-text or vessel value against your configured watchlists and OpenSanctions — using the exact same matching engine as the screening module, so a hit here is a hit there. This guide explains how the search matches, so you can read results with confidence.

Field type

The field type selector controls how your value is screened — it mirrors exactly how the same value is screened inside the screening module.

  • Name — screens the whole value as a name. Long names are also tried as a first + last combination (e.g. "Ahmad Ali Bahar" also tries "Ahmad Bahar"). Honours the Exact/Fuzzy toggle and searches every source, including OpenSanctions.
  • Country — matches the whole value against country records only. Always exact — the Exact/Fuzzy toggle is disabled, because fuzziness would confuse similar country names/codes (e.g. Syria vs Somalia). OpenSanctions is not queried.
  • Free text — for narrative values such as a SWIFT :79: message body. Rather than screening each word on its own, EFI screens overlapping windows of consecutive words — currently two at a time (configurable). "payment to mahan air" is searched as "payment to", "to mahan" and "mahan air", so a multi-word entity like Mahan Air is found without the flood of false hits that isolated common words used to produce. Labelled buyer lines get special handling too — see Free-text: sliding-window matching below. Honours the Exact/Fuzzy toggle and searches every source, including OpenSanctions.
  • Vessel — screens the value as a ship against OpenSanctions Vessel entities. It runs two searches: the value exactly as entered, and an IMO-prefixed variant (e.g. 9274446 also tries IMO9274446), so an IMO number matches whether or not you include the IMO prefix. Matches on the vessel's name and its IMO / registration number; hits carry a blue vessel badge and, on the record detail, show the matched IMO number. Honours the Exact/Fuzzy toggle and searches OpenSanctions.

Free-text: sliding-window matching

Free text — chiefly a SWIFT :79: narrative — is matched differently from names. Prose is full of ordinary words, so screening each word on its own floods you with false hits (a stray PORTS matching a person named poros, a lone 5 matching a company name). Instead, EFI screens the text as overlapping windows of consecutive words.

How windows are built. The value is first split into words on punctuation (so END-USER becomes END and USER, and BUYER:MAHAN becomes BUYER and MAHAN), then every run of N adjacent words is screened as a phrase. N defaults to 2 and is configurable. "transshipment via mahan air" is searched as:

"transshipment via" · "via mahan" · "mahan air"

Because each window is a phrase, a hit needs every word in the window to appear in the candidate name — so a two-word window matches a genuine two-word entity (Mahan Air) but not a single common word. The whole message is not sent as one value, and words are not screened on their own.

Worth knowing. Since a window needs all of its words in the matched name, a single-word sanctioned name sitting in prose (a lone ROSNEFT, a one-word vessel) is not caught by two-word windows on its own — longer names such as Mahan Air or National Iranian Oil Company are unaffected. The window size can be lowered to 1 to restore word-by-word screening.

Labelled buyer lines. Trade-finance narratives often name the party after a label, e.g. END BUYER:ROSNEFT. EFI recognises an END BUYER label followed by any separator — :, /, -, ., ,, +, (, ? and the like — and screens the value that follows, up to the end of that line, on its own. So END BUYER/ROSNEFT OIL COMPANY, END BUYER-HIZBALLAH and END BUYER(SYRIAN PETROLEUM COMPANY) all resolve to the party name, and even a single-word buyer like ROSNEFT is matched. (A plain space — END BUYER DETAILS — is not a separator, so ordinary prose isn't mistaken for a value.) When this happens on a screening request, you'll see it in the request log: free-text: extracted "end buyer" label value "ROSNEFT" — screening it on its own.

Only names that are really there. Each free-text hit is then re-checked against the original text and kept only if the matched entity name actually appears in it (a distinctive single-token match such as Putin is always kept). This drops fragment matches — a query word that lined up with only part of a longer sanctioned name — so a free-text result reflects a name genuinely present in the text.

Stop words, numbers & single characters

Prose is full of ordinary words — articles, prepositions, conjunctions, pronouns, auxiliaries — plus payment jargon (REF, PAYMENT, IBAN), currency codes (USD, EUR) and bare numbers. On their own these carry no name and only ever fuzzy-match a sanctioned entity by accident (a lone AS, IT, THE, or 1/2/3). To stop that noise, EFI drops any free-text window whose every word is a "noise" token.

A token counts as noise when it is:

  • a stop word from the list below (matched case-insensitively, accents folded — so ÇA = ca);
  • a bare number — any token of only digits (1, 42, 2024, 000);
  • a single character (a, x, 5).

A window is only dropped when all of its words are noise. A window that keeps even one real token is untouched, so this removes false positives without ever hiding a genuine match. For example, with two-word windows:

  • "as it", "payment ref", "1 2"dropped (every word is noise).
  • "of america", "bank tejarat", "the putin"kept (america, tejarat, putin are real tokens).

This applies only to the Free text field type. Name, Country, Vessel and BIC/identifier values are screened whole and are never filtered this way — so a country code such as AS (American Samoa) or a short legitimate name is unaffected.

This filter changes only what is sent to the matching engine; it does not change how a hit is scored or reverse-checked.

The complete stop-word list

The list is maintained in the screening service at screening/src/efi_screening/assets/screening-stop-words.txt (the source of truth). Numbers and single characters are handled by the rules above and are not enumerated in the file. It currently holds 751 words:

English a, about, above, after, again, against, all, am, an, and, aren, as, at, be, because, been, before, being, below, between, both, but, by, can, cannot, could, did, do, does, doing, don, down, during, each, few, for, from, further, had, has, have, having, he, her, here, hers, herself, him, himself, his, how, i, if, in, into, is, it, its, itself, just, me, more, most, my, myself, no, nor, not, now, of, off, on, once, only, or, other, our, ours, ourselves, out, over, own, same, she, should, so, some, such, than, that, the, their, theirs, them, themselves, then, there, these, they, this, those, through, to, too, under, until, up, very, was, we, were, what, when, where, which, while, who, whom, why, will, with, would, you, your, yours, yourself, yourselves

German aber, alle, als, also, am, auch, auf, aus, bei, bin, bis, bist, da, damit, dann, das, dass, dein, dem, den, der, des, die, doch, dort, du, durch, ein, eine, einem, einen, einer, eines, er, es, euer, für, hab, habe, haben, hat, hier, ich, ihr, im, in, ist, ja, kann, kein, mein, mit, nach, nein, nicht, noch, nun, nur, ob, oder, sein, sich, sie, sind, so, über, um, und, uns, unter, vom, von, vor, war, wenn, werden, wie, wir, wird, zu, zum, zur

French au, aux, avec, ce, ces, dans, de, des, du, elle, en, et, eux, il, je, la, le, les, leur, lui, ma, mais, me, même, mes, moi, mon, ne, nos, notre, nous, on, ou, par, pas, pour, qu, que, qui, sa, se, ses, son, sur, ta, te, tes, toi, ton, tu, un, une, vos, votre, vous, ça

Spanish al, algo, ante, como, con, contra, cual, de, del, desde, donde, el, ella, ellos, en, entre, era, eres, es, esa, ese, eso, esta, este, esto, ha, han, hasta, hay, la, las, le, lo, los, mas, me, mi, mis, mucho, muy, nada, ni, no, nos, o, os, otra, otro, para, pero, poco, por, porque, que, quien, se, sea, si, sin, sobre, son, su, sus, te, tu, tus, un, una, uno, unos, vosotros, ya, yo

Italian agli, ai, al, alla, alle, anche, che, chi, ci, col, come, con, contro, dal, degli, dei, del, della, delle, di, dove, e, ed, gli, il, io, la, le, lo, ma, mi, nei, nel, nella, no, noi, non, o, per, più, quale, quello, questo, se, sia, solo, sono, su, sui, sul, ti, tra, tu, tuo, tutti, un, una, uno, voi

Portuguese ao, aos, as, até, com, como, da, das, de, do, dos, e, ela, eles, em, entre, essa, esse, esta, este, eu, foi, isso, já, lhe, mais, mas, me, mesmo, meu, minha, muito, na, nas, nem, no, nos, nós, num, o, os, ou, para, pela, pelo, por, qual, que, quem, se, sem, seu, sua, são, só, te, tem, teu, tu, um, uma, você, vocês

Dutch aan, al, als, bij, dan, dat, de, der, deze, die, dit, doch, door, een, en, er, ge, had, heb, het, hij, hoe, ik, in, is, je, kan, me, men, met, mij, na, naar, niet, nog, nu, of, om, onder, ons, ook, op, over, te, tegen, toch, toen, tot, uit, uw, van, veel, voor, was, wat, werd, wezen, wie, wil, zal, ze, zei, zich, zij, zijn, zo, zonder, zou

Russian (Cyrillic) а, без, более, бы, был, была, были, было, быть, в, вам, вас, весь, во, вот, все, всего, всех, вы, где, да, даже, для, до, его, ее, ей, ему, если, есть, еще, же, за, здесь, и, из, или, им, их, к, как, ко, когда, кто, ли, либо, меня, мне, много, может, мы, на, над, надо, наш, не, него, нее, нет, ни, них, но, ну, о, об, он, она, они, оно, от, по, под, при, с, со, так, также, такой, там, те, тем, то, того, тоже, той, только, том, ты, у, уже, хотя, чего, чей, чем, что, чтобы, чуть, эта, эти, это, я

Russian (Latin transliteration) — transaction messages are often ASCII bez, byl, dlya, ego, esli, est, eto, iz, ili, kak, kogda, kto, mne, nad, net, oni, ona, pod, pri, tak, tam, tot, uzhe, chto

Turkish / Kyrgyz / Central-Asian (Bakai corridor) bir, bu, da, de, den, dan, icin, ile, mi, mu, ne, o, ve, ya, жана, менен, үчүн

Transaction-message / payment noise account, acc, amount, bank, bic, charge, charges, credit, customer, date, debit, detail, details, fee, fees, for, iban, info, information, inv, invoice, msg, no, note, num, number, order, pay, payable, payer, payee, payment, purpose, ref, reference, salary, sepa, swift, tax, transaction, transfer, value

Currency codes (ISO 4217) usd, eur, gbp, chf, jpy, cny, rub, kgs, kzt, uzs, try, aed, inr

Month names & abbreviations january, february, march, april, may, june, july, august, september, october, november, december, jan, feb, mar, apr, jun, jul, aug, sep, sept, oct, nov, dec

Tokenization & multiple words

A record's value is stored as a single keyword (the full string, not analyzed). Per-word matching isn't done by the index — instead EFI expands your query into individual words (see Splitting values on punctuation and the first+last expansion above) and, after the search, applies a token-level filter that requires every word of your query to line up with a word of the candidate name. The net effect is an "all words must appear" match:

  • Single word"Ahmad" matches any indexed name containing the word ahmad: Ahmad Bahar, Ahmad Ali Ahmad, etc.
  • Multiple words"Ahmad Bahar" requires both words to appear in the same record. Order doesn't matter, but every word must hit.
  • Whole-string lookups against pre-computed hashes (md5, blake2b, sha3) of the raw value are tried alongside for exact matches.

Splitting values on punctuation

Before a value is searched, EFI breaks it into words on every non-alphanumeric character and searches each word in addition to the whole value. This is what lets a name that is glued to another word by punctuation still be found: a Name search for BUYER:ROSNEFT — no space after the colon — is searched as BUYER, ROSNEFT, and BUYER:ROSNEFT, so the sanctioned entity ROSNEFT is matched instead of missed. This runs for the Name, Country and Vessel field types on both engines (the built-in watchlist and OpenSanctions). Free text is handled by the sliding-window rules above instead — it is split into words but screened as phrases, not as isolated words (a glued buyer in a :79: narrative is caught by the END BUYER: rule). Values with no internal punctuation — plain names, country codes, BICs — are searched unchanged, so it only ever adds coverage.

What counts as a separator. The rule is simply: anything that is not a letter or a digit. Runs of letters/digits (including accented letters, e.g. Müller) are kept as words; every other character splits the value — whitespace, the underscore _, and every punctuation mark or symbol. The characters that occur in SWIFT and ISO 20022 (pacs) messages, and therefore split here, are:

Group Separator characters
SWIFT x character set (used by :70: / :72: / :79: narrative, names, addresses) / - ? : ( ) . , ' +
Extended z set & pacs XML text (after entity decoding) & = ! " % * ; < > @ _ #
Whitespace space, tab, carriage return, line feed

Because the rule is "split on anything that isn't a letter or digit", any character not listed above is a separator too — the table shows the ones you will actually encounter in financial messages.

A word is only ever added — the original value is always searched as well — so this never weakens an exact or whole-name match; it only exposes words that punctuation would otherwise have hidden.

Exact vs Fuzzy

Fuzziness uses edit distance per word — the number of single-character insertions, deletions or substitutions needed to turn one word into another. The tolerance scales with word length. Applies to Name, Free text and Vessel; Country is always exact.

  • Exact — no edits allowed. Requires the word to exist verbatim (after lowercasing) in the indexed name.
  • Fuzzy — Elasticsearch is configured with fuzziness: auto:5,8, the same length scale the post-search word filter uses: 0 edits for words under 5 characters, 1 edit for 5-8, 2 edits for 9 or more. So "putin" (5 chars, 1 edit) catches putln and puttin, and "poroshenko" (10 chars, 2 edits) catches poroshencko, while a 4-letter word must match exactly. The budget is tighter than it looks: "Mohamad" and Mohammed differ by two edits and, at 7 characters, do not match. (Elasticsearch's own fuzzy match also treats a transposition as a single edit; the word filter that re-checks each hit counts it as two, so it is marginally stricter.)

Alphabets on the lists

Sanctioned entities are not all recorded in the Latin alphabet. The official sources publish names in the script of origin as well as in romanized form. A Russian entity may be listed in Cyrillic and in Latin, an Iranian one in Arabic and in Latin, a Chinese one in Han characters and in Latin, and Greek, Hebrew and Korean entries appear the same way. Coverage is uneven: not every record carries every script, and the romanized spelling chosen by one authority often differs from another's.

Matching against the built-in index is script-preserving. The value is compared in the alphabet it arrives in, with no conversion between scripts. Searching Путин matches the Cyrillic spellings held on the lists; searching Putin matches the Latin ones. Neither finds the other.

Two consequences worth planning around:

  • On official lists, whether a search hits depends on which alphabets that particular record carries. If your traffic is romanized ASCII, as most SWIFT messages are, you are matching against the romanized spellings on the list, and a variant romanization may fall outside the fuzzy budget above.
  • On your own custom watchlists, you control which alphabets are present. Add every spelling you expect to see in traffic as a separate entry, Путин as well as Putin, rather than relying on the matcher to bridge them.

Match score

The % match badge reflects OpenSanctions' own match confidence (0-100%). The built-in watchlist match is a filter, not a ranked search — a record either satisfies the exact/fuzzy word match or it doesn't — so watchlist hits carry no relevance score, and the badge is shown only for results that come from OpenSanctions, where a real confidence figure is available.

Because that figure is OpenSanctions' raw score, it is not rescaled so the top hit is always 100% — a genuine but imperfect match can sit below 100%. Treat the percentage as how confident we are it's the same entity, not as did the text match letter-for-letter.

Where it is shown, the badge is colored by tier so you can scan results quickly:

Tier Meaning
≥ 85% strong match — almost certainly the same entity, including with typos.
50% – 84% partial match — typically a name fragment or a fuzzy hit; worth reviewing.
< 50% weak match — only one rare word in common, or a multi-edit fuzzy hit.

Vessels and IMO numbers. This matters most for vessels: when you search by IMO number, OpenSanctions blends the (perfect) match on the number with how closely your text matches the vessel's registered name — so a vessel can show slightly under 100% (say 95%) even when the IMO number matches exactly.