---
title: "Constraining matches to a geographic region"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Constraining matches to a geographic region}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  eval = FALSE
)
```

This vignette shows how to steer `taxify()`'s spelling corrections toward the
species that occur where the data were collected. Two species can sit a single
edit apart on different continents, and a recorder in Belgium who misspells a
name meant the Belgian plant. With a `region` declared, `taxify()` prefers the
fuzzy candidates recorded in that region. Plant ranges come from the World
Checklist of Vascular Plants (WCVP, Govaerts et al.
[2021](https://doi.org/10.1038/s41597-021-00997-6)) on the botanical regions of
the World Geographical Scheme for Recording Plant Distributions (WGSRPD,
Brummitt [2001](https://github.com/tdwg/wgsrpd)); marine ranges come from WoRMS
distribution records rolled up to the Marine Ecoregions of the World (MEOW,
Spalding et al. [2007](https://doi.org/10.1641/B570707)).

1. **Declare** the region by name or code with the `region` argument of
   `taxify()`.
2. **Locate** the records by coordinates with `coords`.
3. **Narrow** to native or introduced occurrences with `range`.
4. **Look up** the accepted region names and codes with `taxify_regions()`.

## Example

```{r, eval = TRUE, message = FALSE, warning = FALSE}
library(taxify)
```

The example uses a short regional list with two misspellings:

```{r}
field_names <- c(
  "Gentiana acaulis", "Primula veris", "Pulsatilla vulgaris",
  "Gentiana acaulary", "Primula elatour"
)
```

### By region name

A region name is matched case- and accent-insensitively at any of the three
WGSRPD levels: a continent (`"Europe"`, Level 1), a sub-continental region
(`"Middle Europe"`, Level 2) or a country (`"Belgium"`, Level 3). Several
regions union.

```{r}
taxify(field_names, region = c("Belgium", "Netherlands", "Germany"))
```

The region acts on fuzzy candidates only. An exact match comes back unchanged:

```{r}
taxify("Gentiana acaulis", region = "Europe")
```

```
#>         input_name    accepted_name       family match_type fuzzy_dist backbone
#> 1 Gentiana acaulis Gentiana acaulis Gentianaceae      exact         NA     COL
```

A TDWG code is read directly, so `region = "BGM"` and `region = "Belgium"`
select the same region. An unrecognised name or code (`"GRE"` for Greece, whose
code is `GRC`) is dropped with a warning, and the call runs without that
constraint.

### By coordinates

Coordinates are mapped to their WGSRPD Level 3 region by point-in-polygon, in
the order `c(lon, lat)`:

```{r}
# Brussels
taxify(field_names, coords = c(4.35, 50.85))
```

A two-column matrix or data.frame of points, an `sf` object or a terra
`SpatVector` work too; spatial objects are reprojected to longitude/latitude.
Points and a `region` name can be combined, and their regions union.

```{r}
occ <- data.frame(
  lon = c(4.35, 5.12, 4.40),
  lat = c(50.85, 51.21, 50.50)
)
taxify(field_names, coords = occ)
```

The boundary file downloads once and is cached. The point-in-polygon test runs
natively by default; with terra or sf installed taxify uses that package, and
`options(taxify.pip_engine = "terra" | "sf" | "native")` forces the choice.

### Native, introduced, or present

By default any WCVP record counts as in-region, native or introduced. The
`range` argument narrows that:

```{r}
taxify(field_names, region = "Europe", range = "native")
taxify(field_names, region = "Europe", range = "introduced")
```

`"native"` suits work that should ignore naturalised populations: a species
present in the region only as an introduction does not satisfy it, so its
out-of-region native neighbour can win the tie. `"introduced"` selects the alien
records for invasion work. `range` has no effect without a region.

### Looking up regions

`taxify_regions()` lists the regions `region` accepts and filters them by a
search term matched against the code and all three level names. The botanical
regions ship with the package:

```{r, eval = TRUE}
taxify_regions("Belgium", scheme = "wgsrpd")
```

Each Level 1 region expands to its Level 3 codes:

```{r, eval = TRUE}
wgsrpd <- taxify_regions(scheme = "wgsrpd")
n_l3 <- as.data.frame(table(wgsrpd$level1_name), stringsAsFactors = FALSE)
knitr::kable(n_l3, col.names = c("Level 1 region", "Level 3 regions"))
```

The same codes appear in the native-range output of `add_wcvp()`.

### Marine regions

Marine names use the `marine_distribution` asset, which downloads on first use.
From then on `region` takes MEOW ecoregion, province and realm names and codes
the same way it takes botanical ones; a province or realm expands to its member
ecoregions.

```{r}
taxify(c("Carcinus maenus", "Gadus morhua"), region = "North Sea")
```

```
#>        input_name   accepted_name     family match_type backbone
#> 1 Carcinus maenus Carcinus maenas Carcinidae      fuzzy     col
#> 2    Gadus morhua    Gadus morhua    Gadidae      exact     col
```

```{r}
head(taxify_regions("Temperate Northern Atlantic", scheme = "meow"))
```

A point at sea maps to the MEOW ecoregion containing it and is unioned with the
botanical lookup, so one `coords` argument serves a list of plants and marine
animals. MEOW covers coastal and shelf waters; a point over a deep ocean basin
belongs to no ecoregion and leaves those names unconstrained.

## How the filter decides

For each input name, `taxify()` looks its fuzzy candidates up in the range
source that owns the resolved region codes and drops an out-of-region
candidate when another candidate survives. Three rules apply:

- Exact and case-folded matches are never filtered.
- A candidate with no range data is kept. Vascular plants and marine taxa are
  covered; a name outside both passes through unchanged, so a mixed list can
  carry a region safely.
- When every candidate for a name is out of region, all are kept.

Marine ranges inherit the grain of the WoRMS locality they were recorded
against, which runs from a single bay to an ocean basin. The median species
spans 4 ecoregions and a quarter span exactly one; about 0.3% span more than
half the ocean and are in region wherever it is asked.

The constraint is most useful on regional field lists with misspellings, where
the intended correction and a wrong one are a single edit apart. The related
check in `inspect()` works after matching: it flags matched names that WCVP does
not record in the declared region, using the same `region`, `coords` and `range`
arguments.

## Where to go next

- [Inspecting a name list](https://gillescolling.com/taxify/articles/inspecting-names.html)
  for the geographic outlier check that uses the same region inputs.

- [Fuzzy matching](https://gillescolling.com/taxify/articles/fuzzy-matching.html)
  for the candidate generation the region filter refines.

- [Enrichments](https://gillescolling.com/taxify/articles/enrichments.html)
  for `add_wcvp()`, which attaches native range on the same TDWG codes.

## References

Brummitt RK (2001). *World Geographical Scheme for Recording Plant
Distributions*, Edition 2. <https://github.com/tdwg/wgsrpd>

Govaerts R, Nic Lughadha E, Black N, Turner R, Paton A (2021). The World
Checklist of Vascular Plants, a continuously updated resource for exploring
global plant diversity. *Scientific Data* 8: 215.
<https://doi.org/10.1038/s41597-021-00997-6>

Spalding MD, Fox HE, Allen GR, Davidson N, Ferdana ZA, Finlayson M, Halpern BS,
Jorge MA, Lombana A, Lourie SA, Martin KD, McManus E, Molnar J, Recchia CA,
Robertson J (2007). Marine Ecoregions of the World: a bioregionalization of
coastal and shelf areas. *BioScience* 57: 573-583.
<https://doi.org/10.1641/B570707>
