---
title: "Fuzzy matching: methods, thresholds, and tuning"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Fuzzy matching: methods, thresholds, and tuning}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(eval = FALSE)
```

This vignette shows how `taxify()` corrects misspelled names and how to tune
that correction. When a name finds no exact match, `taxify()` computes a string
distance between it and the backbone candidates in the same genus, and returns
the closest candidate whose distance falls below a threshold. Genus blocking
keeps the search fast on backbones with millions of rows. String distance counts
character edits, so a fuzzy match says two strings are spelled alike and nothing
about whether they name the same organism: it catches typos, transliteration
errors and OCR artefacts, and it cannot resolve a taxonomic disagreement or turn
a common name into a Latin binomial. A fuzzy match at distance 0.05 almost
certainly corrects a minor typo; one at 0.18 might correct a larger OCR error or
might have landed on the wrong species, and the `fuzzy_dist` column in the
output is there to tell the two apart. Fuzzy matching runs only on names that
survived cleaning and every exact pass without a hit, so it never overrides an
exact match, even when a closer candidate exists under another spelling.

1. **Match** the names with `taxify()`, which falls back to fuzzy matching for
   names no exact pass resolves.
2. **Pick** the distance algorithm with `fuzzy_method`.
3. **Set** the tolerance with `fuzzy_threshold`, as a fraction of name length or
   as a raw edit count.
4. **Review** the questionable matches by their `fuzzy_dist`.
5. **Skip** fuzzy matching for curated lists with `fuzzy = FALSE`.

## Example

```{r, eval = TRUE, message = FALSE, warning = FALSE}
library(taxify)
```

### Names that need no fuzzy matching

A curated species list, with or without authorship strings, typically resolves
entirely by exact match. Fuzzy matching runs but finds nothing to do.

```{r clean-names}
clean_names <- c(
  "Quercus robur",
  "Pinus sylvestris",
  "Betula pendula",
  "Fagus sylvatica",
  "Acer pseudoplatanus"
)
result <- taxify(clean_names)

# All rows have match_type == "exact"
table(result$match_type)
# exact
#     5

# fuzzy_dist is NA for all rows
all(is.na(result$fuzzy_dist))
# TRUE
```

The cleaning pipeline strips authorship before matching, so adding it does not
change the picture:

```{r clean-names-authorship}
with_authors <- c(
  "Quercus robur L.",
  "Pinus sylvestris L.",
  "Betula pendula Roth",
  "Fagus sylvatica L.",
  "Acer pseudoplatanus L."
)
result <- taxify(with_authors)
table(result$match_type)
# exact
#     5
```

For curated data like this, `fuzzy = FALSE` skips the step entirely and saves a
small amount of time on large lists.

### Typos and OCR errors

Species lists transcribed from handwritten field notes or extracted from scanned
PDFs by OCR arrive with the errors fuzzy matching is designed to rescue:

```{r ocr-degraded}
messy_names <- c(
  "Qurecus robur",         # transposition: ur -> ru
  "Taraxacum officianle",  # transposition: al -> la
  "Plantago lanceoalata",  # transposition: la -> al
  "Trifolium repnes",      # transposition: en -> ne
  "Dactylis gloemrata",    # transposition: me -> em
  "Lolium perrene",        # two edits from perenne
  "Achillea millefolum",   # deletion: i missing
  "Ranunculus acris"       # correct (should exact-match)
)
result <- taxify(messy_names)

# Check what matched and how
result[, c("input_name", "accepted_name", "match_type", "fuzzy_dist")]
```

The transposition errors (*Qurecus*, *officianle*, *lanceoalata*) each cost 1
edit under Damerau-Levenshtein, producing `fuzzy_dist` values around 0.05-0.08
for these 13-20 character names. The deletion in *millefolum* (missing *i*) also
costs 1 edit, and *Ranunculus acris* exact-matches with `fuzzy_dist = NA`.

For *Taraxacum officianle* (20 characters) the intended target, *Taraxacum
officinale*, differs by a transposition of *a* and *l* at positions 18-19: 1
edit, a normalized distance of `1 / 20 = 0.05`, inside the 0.2 threshold and
inside a conservative 0.1 too. *Lolium perrene* (14 characters) puts the doubled
consonant on the wrong letter compared with *Lolium perenne*, which costs 2
edits, `2 / 14 = 0.143`. Every one of these falls within the default threshold of
0.2, so for data with this error profile the default settings work as they are.
The single-edit corrections cluster at 0.05-0.08; *Lolium perrene* sits higher,
in the range where a glance is worthwhile.

### Choosing a distance method

The `fuzzy_method` argument selects one of three string distance algorithms. All
three are computed at the C level inside vectra's `fuzzy_join()`, which runs the
genus-blocked comparisons in parallel via OpenMP.

Damerau-Levenshtein (`fuzzy_method = "dl"`, the default) counts four edit
operations, each costing 1:

- insertion: add a character (*Querus* to *Quercus*)
- deletion: remove a character (*Quercuss* to *Quercus*)
- substitution: replace one character with another (*Quarcus* to *Quercus*)
- transposition: swap two adjacent characters (*Qurecus* to *Quercus*)

Transpositions are among the most common typos in hand-entered data, and
counting one as a single edit, where plain Levenshtein counts a deletion plus an
insertion, gives tighter distances for real-world errors.

Levenshtein (`fuzzy_method = "levenshtein"`) supports only insertion, deletion
and substitution, so *Qurecus* to *Quercus* costs 2 edits (delete the *r*,
insert *r* at the right position). It is stricter than Damerau-Levenshtein for
transposition errors and identical for everything else, so the same threshold
rejects candidates Damerau-Levenshtein would accept. It suits input from a
controlled source (a database export, a curated checklist) where transpositions
are rare. For OCR or hand-typed data, Damerau-Levenshtein is almost always the
better choice.

Jaro-Winkler (`fuzzy_method = "jw"`) computes a similarity score between 0
(completely different) and 1 (identical), which taxify converts to a distance as
`1 - similarity`. The algorithm gives extra weight to characters that match at
the beginning of the string. For taxonomic names that suits the observation that
the genus is the most informative part and that prefix errors are rarer than
epithet errors. Because the scale is 0 to 1 by definition, only fractional
thresholds are supported: an integer threshold such as `fuzzy_threshold = 2`
with `fuzzy_method = "jw"` raises an error immediately. Jaro-Winkler can help
with very short names (3-5 characters), where a single edit produces a large
normalized Damerau-Levenshtein distance, and with datasets whose errors
concentrate in the epithet. For general-purpose matching, Damerau-Levenshtein
remains the safer default.

The same input list can produce different results under each method, most
visibly when the errors include transpositions:

```{r compare-methods}
test_names <- c(
  "Qurecus robur",        # transposition in genus
  "Achillea milefolium",  # deletion (l dropped)
  "Plantago lanceoalata", # transposition in epithet
  "Betula pednula",       # transposition in epithet
  "Fagus sylvatcia"       # transposition in epithet
)

dl_result  <- taxify(test_names, fuzzy_method = "dl")
lev_result <- taxify(test_names, fuzzy_method = "levenshtein")
jw_result  <- taxify(test_names, fuzzy_method = "jw")

# Compare fuzzy_dist across methods
comparison <- data.frame(
  input = test_names,
  dl_dist  = dl_result$fuzzy_dist,
  lev_dist = lev_result$fuzzy_dist,
  jw_dist  = jw_result$fuzzy_dist,
  dl_match  = dl_result$match_type,
  lev_match = lev_result$match_type,
  jw_match  = jw_result$match_type
)
comparison
```

For a transposition like *Qurecus* to *Quercus*, Damerau-Levenshtein reports 1
edit (distance 0.08 on a 13-character name) and Levenshtein 2 edits (0.15).
Both fall within the default 0.2 threshold, so both methods match it, but the
Levenshtein distance is nearly double. A deletion like *milefolium* to
*millefolium* involves no transposition, and both methods report the same
distance.

Jaro-Winkler distances tend to be smaller overall because the algorithm rewards
matching prefixes: a name that shares its entire genus with the candidate starts
from a high base similarity. At the same numeric threshold Jaro-Winkler is
therefore more permissive; 0.2 is quite loose, and 0.1 is a more comparable
starting point. The table below computes the three distances, with vectra's
string distance engine, for one example of each error type:

```{r method-table, eval = TRUE, echo = FALSE, message = FALSE, warning = FALSE}
pairs <- data.frame(
  error  = c("Single transposition", "Single deletion",
             "Single substitution", "Two transpositions"),
  input  = c("Qurecus robur", "Achillea millefolum",
             "Quarcus robur", "Qeurcus robru"),
  target = c("Quercus robur", "Achillea millefolium",
             "Quercus robur", "Quercus robur")
)
pair_dist <- function(method) {
  q <- tempfile(fileext = ".vtr")
  b <- tempfile(fileext = ".vtr")
  on.exit(unlink(c(q, b)))
  vectra::write_vtr(data.frame(qid = seq_len(nrow(pairs)), input = pairs$input), q)
  vectra::write_vtr(data.frame(bid = seq_len(nrow(pairs)), target = pairs$target), b)
  m <- vectra::collect(vectra::fuzzy_join(
    vectra::tbl(q), vectra::tbl(b),
    by = c(input = "target"), method = method, max_dist = 1
  ))
  m <- m[m$qid == m$bid, ]
  round(m$fuzzy_dist[match(seq_len(nrow(pairs)), m$qid)], 2)
}
knitr::kable(
  data.frame(pairs$error, pairs$input, pairs$target,
             pair_dist("dl"), pair_dist("levenshtein"), pair_dist("jw")),
  col.names = c("Error type", "Input", "Target", "DL dist", "Lev dist", "JW dist")
)
```

The Levenshtein column is always equal to or larger than the Damerau-Levenshtein
column, because Levenshtein charges double for transpositions. Jaro-Winkler is
consistently the smallest, because the shared genus prefix dominates the
similarity. This is why the same threshold value behaves differently across
methods and needs recalibrating on a switch.

### Setting the threshold

`fuzzy_threshold` sets how different two strings can be before the match is
rejected, in one of two modes depending on its value.

A fractional threshold (0 < threshold < 1) caps the normalized distance. The
default of `0.2` means the normalized distance must not exceed 0.2, where

```
normalized_distance = raw_edits / max(nchar(input), nchar(candidate))
```

The cap scales with name length. A 5-character name (*Abies*) gets at most 1
edit at threshold 0.2, because `1 / 5 = 0.2`; a 12-character name gets at most 2,
because `2 / 12 = 0.167 < 0.2` but `3 / 12 = 0.25 > 0.2`; a 20-character name
gets up to 4. The table computes the largest number of edits each threshold
allows for a few representative names:

```{r threshold-table, eval = TRUE}
nm  <- c("Poa annua", "Quercus robur", "Taraxacum officinale",
         "Achillea millefolium", "Brachypodium sylvaticum")
len <- nchar(nm)
knitr::kable(
  data.frame(nm, len, floor(len * 0.2), floor(len * 0.1), floor(len * 0.3)),
  col.names = c("Input name", "Length", "Max edits at 0.2",
                "Max edits at 0.1", "Max edits at 0.3")
)
```

The "max edits" columns are `floor(length * threshold)`. The comparison itself
uses the floating-point ratio, so a 9-character name with 2 edits gives
`2 / 9 = 0.222`, which exceeds 0.2 and is rejected.

An integer threshold (1, 2, 3, ...) caps the raw edit count regardless of name
length: `fuzzy_threshold = 2L` means at most 2 edits, whether the name is 5
characters or 25. This mode suits data whose error pattern is known. If the
input comes from an OCR pipeline that occasionally drops or doubles a single
character, `fuzzy_threshold = 1L` captures those errors without over-matching on
longer names. Integer thresholds are not supported for Jaro-Winkler, which does
not count discrete edits.

```{r integer-threshold}
# Allow exactly 1 edit, regardless of name length
result <- taxify(
  c("Qurecus robur", "Achillea milefolium", "Poa anua"),
  fuzzy_threshold = 1L
)
# "Qurecus robur" matches (1 transposition)
# "Achillea milefolium" matches (1 deletion: ll -> l)
# "Poa anua" matches (1 deletion: nn -> n)
```

A loose threshold can match names to the wrong species. This is the main risk of
fuzzy matching, and it is highest for short names and names in species-rich
genera:

```{r threshold-too-loose}
# Poa is a large genus with many similar epithets
poa_names <- c(
  "Poa anua",       # intended: Poa annua (1 edit)
  "Poa pratenss",   # intended: Poa pratensis (1 edit)
  "Poa trialis"     # intended: Poa trivialis (2 edits)
)

# With a loose threshold, some may match the wrong species
loose <- taxify(poa_names, fuzzy_threshold = 0.4)
loose[, c("input_name", "accepted_name", "fuzzy_dist")]
```

At threshold 0.4, *Poa trialis* (11 characters) is allowed up to 4 edits. That
reaches *Poa trivialis*, the intended target at 2 edits, and potentially other
*Poa* species that happen to be closer in string distance. With 500+ *Poa*
species in the backbone, the risk of a false match is real. Tightening the
threshold removes it:

```{r threshold-tightened}
tight <- taxify(poa_names, fuzzy_threshold = 0.15)
tight[, c("input_name", "accepted_name", "match_type", "fuzzy_dist")]
# "Poa anua" still matches (1/9 = 0.11 < 0.15)
# "Poa pratenss" still matches (1/13 = 0.08 < 0.15)
# "Poa trialis" fails (2/13 = 0.154 > 0.15), safer to leave unmatched
```

A name that fails fuzzy matching gets `match_type = "none"`. An unmatched name
can be reviewed by hand; a wrong match gives no sign of itself and propagates
into downstream analyses.

The genus *Poa* has over 500 accepted species in the backbone, many with
epithets a few characters apart (*pratensis* and *palustris*). The shorter the
name, the fewer edits it takes to reach the threshold and the more candidate
species fall within range. Carex (2,000+ species), Astragalus (3,000+) and
Euphorbia (2,000+) raise the same problem, and for species-rich genera a
threshold of 0.1-0.15 is almost always the better setting.

### Reviewing matches with `fuzzy_dist`

Every row of the output has a `fuzzy_dist` column. It is `NA` for exact matches
(including case-insensitive and Latin-normalized ones) and holds the normalized
distance, between 0 and 1 with lower meaning closer, for fuzzy matches. A simple
filter separates high-confidence matches from questionable ones:

```{r fuzzy-dist-filter}
result <- taxify(my_species_list)

# High-confidence fuzzy matches (likely just typos)
good_fuzzy <- result[result$match_type == "fuzzy" &
                     result$fuzzy_dist < 0.1, ]

# Questionable fuzzy matches (review manually)
check_fuzzy <- result[result$match_type == "fuzzy" &
                      result$fuzzy_dist >= 0.1, ]
```

A `fuzzy_dist` below 0.1 on a name of 10+ characters means at most 1 edit, and
such matches are almost always correct. Between 0.1 and 0.2 means 1-3 edits
depending on name length and warrants a glance, and anything above 0.15 on a
short name (under 10 characters) deserves scrutiny. Sorting by `fuzzy_dist` in
descending order puts the most suspect matches at the top:

```{r sort-by-dist}
fuzzy_rows <- result[result$match_type == "fuzzy", ]
fuzzy_rows <- fuzzy_rows[order(-fuzzy_rows$fuzzy_dist), ]
head(fuzzy_rows[, c("input_name", "accepted_name", "fuzzy_dist")], 20)
```

Most datasets show a bimodal distribution of `fuzzy_dist`: a peak near 0.05-0.08
(single typos on medium-length names) and a sparse tail above 0.12 (multiple
errors, or short names with one error). The tail is where false matches hide.
As a rule of thumb, if more than 5% of fuzzy matches have `fuzzy_dist` above
0.15, the threshold is probably too loose for the dataset; either tighten it, or
keep it and flag every match above 0.12 for manual review. Reviewing a few dozen
names costs little next to carrying a wrong species identity through a trait
analysis or a distribution model.

### A two-pass workflow for messy data

For datasets with unknown error rates (historical collections, aggregated
multi-source lists), two passes avoid the all-or-nothing choice between a tight
and a loose threshold. The first pass runs with a tight threshold to get the
high-confidence matches; the second runs the names still unmatched with a looser
threshold, and its additional fuzzy matches go to manual review.

```{r two-pass}
# Pass 1: conservative
pass1 <- taxify(my_names, fuzzy_threshold = 0.1)
unmatched <- pass1$input_name[pass1$match_type == "none"]

# Pass 2: permissive, for manual review
pass2 <- taxify(unmatched, fuzzy_threshold = 0.25)
needs_review <- pass2[pass2$match_type == "fuzzy", ]
needs_review[, c("input_name", "accepted_name", "fuzzy_dist")]
```

The bulk of the data is matched at high confidence, and only the residual names
get the looser treatment, with a person checking the result.

## What runs before fuzzy matching

`taxify()` matches in a fixed sequence. Name cleaning comes first. Then come the
exact passes: case-sensitive, case-insensitive, and Latin orthographic
normalization. Next, an abbreviated genus (*Q. robur*) is resolved from the genus
initial plus the epithet. Fuzzy matching follows, and only after it has failed
does an infraspecific name that no backbone carries fall back to its species
(`match_type = "rank_fallback"`), so a misspelled infraspecific epithet reaches
the fuzzy stage before its species is used.

The cleaning pipeline strips qualifiers (`cf.`, `aff.`, `s.l.`, `s.str.`),
removes authorship strings (`L.`, `(Aiton) Sm.`), drops brackets and trailing
numbers, collapses whitespace, and lowercases everything except the genus.
Backbone names are already clean, so this step brings user input into the same
format, and many names that look as if they need fuzzy matching resolve by exact
match once the noise is gone:

```{r cleaning-before-matching}
# All three resolve to the same clean form: "Quercus robur"
result <- taxify(c(
  "Quercus robur L.",
  "Quercus robur (L.) Sm.",
  "  Quercus  robur  "
))
# match_type will be "exact" for all three (no fuzzy needed)
```

Latin orthographic normalization is a separate exact pass. Alternations like
*ae/i* (*hirtaeformis* and *hirtiformis*), *ph/f*, *rh/r*, *th/t*, and *ii/i* at
word endings are normalized before comparison, and these matches appear as
`exact_ci` in the output. Hybrid markers (the multiplication sign or a
standalone "x") are detected and stripped during cleaning: `Quercus × hispanica`
is matched as `Quercus hispanica`, with the `is_hybrid` column set to `TRUE`. By
the time fuzzy matching runs, the remaining names have genuine character-level
errors.

## Misspelled genera

Fuzzy matching is genus-blocked: taxify extracts the genus from the input name
and compares only against backbone entries with the same genus. This avoids
comparing every input against millions of candidates, and it keeps a misspelled
epithet from matching a name in a completely different genus.

A misspelled genus therefore finds no candidates in the genus-blocked pass. For
WFO, COL, COL XR and GBIF, taxify runs a second, prefix-blocked fuzzy pass on the
names still unmatched, blocking on the first two characters of the name instead
of the full genus. Most genus typos preserve those characters (*Qurecus* still
starts with *Qu*, *Betual* with *Be*), so the prefix block catches them while
still pruning the search space. A typo in the first two characters of the genus
is not caught, and the other backbones run the genus-blocked pass only.

A misspelled genus adds edits on top of any epithet error, so such names carry a
higher `fuzzy_dist`. *Qeurcus robru* (a transposition in the genus and another in
the epithet) costs 2 edits, a normalized distance of `2 / 13 = 0.154`: inside the
default threshold, but in the range where manual review is advisable.

## Practical guidance

### When to disable fuzzy matching

For curated checklists, validated databases, or any input already run through a
name-resolution service, fuzzy matching adds risk without benefit. Disabling it
also skips the fuzzy join step; on a list of 100,000 names the difference can be
several seconds.

```{r disable-fuzzy}
result <- taxify(curated_list, fuzzy = FALSE)
```

### When to tighten the threshold

Tighten below the default 0.2 when the input names are short (many two-word
names under 12 characters), when the genera are species-rich (Carex, Poa,
Astragalus, Euphorbia), or when false matches would be costly (conservation
assessments, regulatory lists). A threshold of 0.1 still catches single-character
typos on names of 10+ characters and rejects matches that need 2+ edits on
shorter names.

```{r tight-threshold}
result <- taxify(short_grass_list, fuzzy_threshold = 0.1)
```

### When to loosen the threshold

Loosen above 0.2 when the input comes from OCR on degraded documents, when names
have been transliterated across character encodings, or when completeness
matters more than precision (an initial screening pass where unmatched names are
expensive to follow up). A threshold of 0.25-0.3 is reasonable for OCR data;
going above 0.3 is rarely justified.

```{r loose-threshold}
result <- taxify(ocr_names, fuzzy_threshold = 0.25)
# Then filter questionable matches:
suspect <- result[result$fuzzy_dist > 0.15, ]
```

### When to switch methods

Damerau-Levenshtein (`"dl"`) is the default for general use. Levenshtein
(`"levenshtein"`) gives the stricter distance for controlled data where
transpositions are unlikely. Jaro-Winkler (`"jw"`) helps with very short names
(3-6 characters, e.g., matching at genus level) where the prefix weighting
matters, with the threshold lowered to 0.1 or below.

### Integer thresholds for a uniform error budget

When the error model is known ("our OCR pipeline drops or adds at most 1
character"), an integer threshold gives direct control: `fuzzy_threshold = 1L`
means at most 1 edit on a name of any length. A fractional threshold instead
allows a 5-character name 1 edit and a 25-character name 5.

```{r integer-threshold-uniform}
# Uniform 2-edit budget, regardless of name length
result <- taxify(my_names, fuzzy_threshold = 2L)
```

## Output columns related to fuzzy matching

| Column       | Values                                     | Meaning                                               |
|:-------------|:-------------------------------------------|:------------------------------------------------------|
| `match_type` | `"exact"`, `"exact_ci"`, `"abbrev"`, `"fuzzy"`, `"hybrid_formula"`, `"rank_fallback"`, `"basionym"`, `"none"`, `"out_of_scope"` | How the name was matched. `"exact"` is case-sensitive, `"exact_ci"` includes case-insensitive and Latin normalization matches. |
| `fuzzy_dist` | Numeric (0-1) or `NA`                      | Normalized string distance for fuzzy matches. `NA` for exact matches and unmatched names. |
| `backbone`   | `"wfo"`, `"col"`, `"gbif"`, etc.           | Which backbone provided the match. Useful in multi-backbone fallback chains. |

Exact matches are definitive, and fuzzy matches with distance below 0.05 are
near-certain corrections of minor typos. As the distance climbs toward 0.15 and
above, manual review becomes worthwhile, because the matched name may belong to
a different species.

## Where to go next

- [Constraining matches to a geographic region](https://gillescolling.com/taxify/articles/regions.html)
  for the `region` filter, which narrows fuzzy candidates to species recorded
  where the data were collected.
