---
title: "Choosing and combining backbones"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Choosing and combining backbones}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(eval = FALSE)
```

This vignette shows how to pick a taxonomic backbone for a name list, how to
chain several backbones so that each name is resolved by the source best placed
to judge it, and how to check afterwards which source resolved which name.
taxify matches names against locally stored Darwin Core backbone databases.
<!-- manifest:backbone-count -->19<!-- /manifest:backbone-count --> backbones are available, each compiled from a different
authoritative source. The backbone we choose determines which names can be
matched, which taxonomic opinion governs synonym resolution, and which extra
metadata columns are available downstream.

1. **Choose** the backbones that cover the taxonomic scope of the list (see
   [The backbones](#the-backbones)); `list_backbones()` reports each one's
   current version, name count, and whether it is installed.
2. **Download** them ahead of time, or pin a version, with `taxify_download()`.
3. **Match** against a single backbone with the `backbone` argument of
   `taxify()`.
4. **Chain** several backbones in fallback order by passing a vector to
   `backbone`.
5. **Audit** which backbone matched each name through the `backbone` and
   `backbone_version` columns.
6. **Extend** the result with backbone-specific columns from `add_wfo_info()`,
   `add_col_info()`, and `add_gbif_info()`.
7. **Diagnose** names outside a backbone's scope with `lookup_genus()` and
   `taxify_register_coverage()`.

## Example

```{r, eval = TRUE, message = FALSE, warning = FALSE}
library(taxify)
```

### Downloading backbones

taxify downloads a backbone on first use. When `taxify(names, backbone =
"wfo")` finds no local WFO backbone, it fetches the pre-built `.vtr` file from
the taxifydb release it is published under, writes it to `taxify_data_dir()`,
and caches the path for the rest of the R session. Later calls, in the same
session or in future ones, reuse the local copy without network access.
`taxify_download()` fetches backbones ahead of time, which is useful on a
shared server or in a Docker image:

```{r download-single}
# Download one backbone
taxify_download("wfo")

# Download several at once
taxify_download(c("wfo", "col", "worms"))
```

Pre-built `.vtr` files are published as GitHub Releases on
[taxifydb](https://github.com/gcol33/taxifydb), one tag per backbone
(`gbif-2026.06`, `algaebase-2026.06`), and range from 8 MB (AviList) to 2.0 GB
(COL). They are compiled from the raw Darwin Core sources with precomputed
matching keys, embedded synonym resolution, and genus-level indexes, so they
can be queried as soon as the download completes.

A specific backbone version can be pinned:

```{r download-pinned}
taxify_download("wfo", version = "2024.01")
```

Pinned versions are stored in their own directory
(`taxify_data_dir()/wfo/2024.01/`) and are never overwritten by future
updates. The "latest" slot (`taxify_data_dir()/wfo/latest/`) is overwritten
whenever a newer version becomes available. A project that locks a backbone
version produces identical results regardless of when the analysis is re-run.

### Matching against one backbone

The simplest case matches plant names against WFO:

```{r single-wfo}
plants <- c(
  "Quercus robur",
  "Quercus petraea",
  "Pinus sylvestris",
  "Acer pseudoplatanus",
  "Betula pendula",
  "Fagus sylvatica",
  "Picea abies"
)

result <- taxify(plants, backbone = "wfo")
result[, c("input_name", "accepted_name", "family", "match_type", "backbone")]
```

Every row in the output has 27 columns regardless of which backbone produced
it: `input_name`, `matched_name`, `accepted_name`, `taxon_id`, `accepted_id`,
`rank`, `family`, `genus`, `epithet`, `authorship`, `accepted_authorship`,
`is_synonym`, `taxonomic_status`, `is_hybrid`, `match_type`, `fuzzy_dist`,
`n_ids`, `accepted_ids`, `backbone`, `backbone_version`,
`kingdom_group`, `taxon_group`, `life_form`, `qualifier`,
`qualifier_position`, `aggregate_fallback`, and `hybrid_type`. The `backbone`
column records `"wfo"` for matched rows and `NA` for unmatched ones. The
`backbone_version` column records the backbone name, version, and download
date (e.g., `"wfo:2024-12 (2026-04-01)"`), so we can cite the exact data
snapshot used.

When a name matches a synonym, taxify resolves it to the accepted name. The
`matched_name` column shows the string that matched in the backbone (which may
be a synonym), `accepted_name` shows the current accepted name after
resolution, and `is_synonym` is `TRUE` for resolved synonyms and `FALSE` for
direct matches.

Fuzzy matching is on by default with a normalized Damerau-Levenshtein
threshold of 0.2, roughly one edit per five characters, which catches typos
such as transposed letters or missing diacritics. A lower threshold is
stricter:

```{r single-strict}
result <- taxify(plants, backbone = "wfo", fuzzy_threshold = 0.1)
```

`fuzzy = FALSE` turns fuzzy matching off and accepts only exact and
case-insensitive matches:

```{r single-no-fuzzy}
result <- taxify(plants, backbone = "wfo", fuzzy = FALSE)
```

The `fuzzy_method` argument selects the distance metric: the default `"dl"`
(Damerau-Levenshtein, which counts a transposition as one edit),
`"levenshtein"` (standard Levenshtein, no transposition handling), or `"jw"`
(Jaro-Winkler, which weights agreement in the first characters of the two
strings).

### Chaining backbones

A wetland monitoring dataset might contain vascular plants, invertebrates,
amphibians, algae, and fungi, and no single backbone covers all of these
equally well. Passing a vector of backbone names builds a fallback chain:
names are matched against each backbone in order, and a name matched by an
earlier backbone is removed from the pool and never re-matched by a later one.

```{r multi-basic}
mixed <- c(
  "Quercus robur",       # plant
  "Panthera leo",        # animal
  "Amanita muscaria",    # fungus
  "Salmo trutta",        # fish
  "Escherichia coli"     # bacterium
)

result <- taxify(mixed, backbone = c("wfo", "col", "gbif"))
result[, c("input_name", "accepted_name", "match_type", "backbone")]
```

The console output during matching shows the chain in action:

```
Matching 5 names against 3 backbones: wfo -> col -> gbif
  [wfo] Matching 5 names...
  [col] Matching 4 remaining names...
  [gbif] Matching 1 remaining names...
```

`"Quercus robur"` matches in WFO and leaves the pool. The remaining four names
go to COL, and anything still unmatched after COL (perhaps an obscure
bacterial name) goes to GBIF. The process continues until all names have been
tried against all backbones or all names have matched. If earlier backbones
have matched every name, later ones are skipped with a message:

```
  [gbif] Skipped (all names matched)
```

The order of the vector decides which taxonomic opinion wins for each name.
If `"Quercus robur"` exists in both WFO and COL, putting WFO first means WFO's
accepted name, family assignment, and synonym resolution are used; putting COL
first gives COL's. For names that exist in several backbones, the first
backbone in the chain always wins. With GBIF first, everything would match
there (GBIF has ~6.4M names, the largest of any backbone) and the curated
opinions of WFO, COL, or WoRMS would never be consulted. For a plant-heavy
list with some non-plant taxa, `c("wfo", "col")` or `c("wfo", "col", "gbif")`
gives WFO's curated plant taxonomy for the plants while COL or GBIF picks up
the rest.

Fuzzy matching runs independently within each backbone. A name that fails
exact matching in WFO is fuzzy-matched against WFO, and only if that also
fails does it move to the next backbone, where it is exact-matched and then
fuzzy-matched again. A misspelled plant name therefore gets its best chance
in WFO, the plant-specialist backbone, before falling through to COL or GBIF.

### Plants: WFO with a COL fallback

WFO focuses on accepted vascular plant and bryophyte names. Names that appear
only in older literature, belong to genera not yet integrated into WFO, or are
nomenclaturally orphaned (no clear accepted name) may be absent. COL inherits
WFO's plant taxonomy as one of its sector databases and supplements it with
names from other sources, including historical synonyms and cultivar names.

```{r plants-wfo-only}
plants <- c(
  "Quercus robur",
  "Quercus petraea",
  "Pinus sylvestris",
  "Acer pseudoplatanus",
  "Coffea arabica",
  "Welwitschia mirabilis",
  "Lepidodendron aculeatum",   # extinct lycopsid
  "Nothofagus cunninghamii",
  "Dracaena draco"
)

# WFO alone
wfo_result <- taxify(plants, backbone = "wfo")
table(wfo_result$match_type)
```

If any names come back as `"none"`, COL can be added as a fallback:

```{r plants-wfo-col}
# WFO first, COL as fallback
both_result <- taxify(plants, backbone = c("wfo", "col"))
table(both_result$match_type)
both_result[, c("input_name", "accepted_name", "backbone")]
```

The `backbone` column now shows `"wfo"` for names matched by WFO and `"col"`
for names only COL could resolve. A methods section can then state "plant
names were resolved against WFO 2024-12, with unmatched names resolved against
COL 2025." In large vegetation plot datasets most names resolve in WFO, and
the handful of edge cases (cultivars, historical names, genera recently moved
between families) that fall through to COL would otherwise need manual
resolution.

### Mixed kingdoms: COL, GBIF and WoRMS

A monitoring dataset from a coastal estuary might contain vascular plants,
invertebrates, fish, and marine algae. COL has broad expert-curated coverage
across kingdoms, GBIF fills gaps with its larger name pool, including names
from national checklists not yet incorporated into COL, and WoRMS is a final
backstop for marine invertebrate synonyms, which can be slow to propagate to
generalist databases. The chain runs from highest curation to broadest
coverage, and most of these names resolve in COL.

```{r mixed-kingdom}
estuary_species <- c(
  "Zostera marina",              # seagrass (plant)
  "Salicornia europaea",         # glasswort (plant)
  "Carcinus maenas",             # shore crab
  "Mytilus edulis",              # blue mussel
  "Platichthys flesus",          # European flounder
  "Nereis diversicolor",         # ragworm
  "Fucus vesiculosus",           # bladderwrack (brown alga)
  "Littorina littorea",          # common periwinkle
  "Arenicola marina",            # lugworm
  "Cerastoderma edule"           # common cockle
)

result <- taxify(estuary_species, backbone = c("col", "gbif", "worms"))
result[, c("input_name", "accepted_name", "family", "backbone")]
```

For a predominantly marine list with only a few terrestrial taxa, leading with
WoRMS makes its marine taxonomy take precedence:

```{r marine-first}
result <- taxify(estuary_species, backbone = c("worms", "col"))
```

### Fungi: Species Fungorum with a COL fallback

Species Fungorum Plus is curated specifically for fungi, with ~315k names
including anamorphs, teleomorphs, and the pleomorphic naming changes of the
2011 Melbourne Code. Its synonym coverage for fungal genera is better than in
generalist databases, where fungal taxonomy is often a secondary concern.

```{r fungi}
fungi <- c(
  "Amanita muscaria",
  "Boletus edulis",
  "Cantharellus cibarius",
  "Tuber melanosporum",
  "Saccharomyces cerevisiae",
  "Aspergillus niger",
  "Penicillium chrysogenum",
  "Agaricus bisporus",
  "Trametes versicolor",
  "Cordyceps militaris"
)

result <- taxify(fungi, backbone = c("fungorum", "col"))
result[, c("input_name", "accepted_name", "is_synonym", "backbone")]
```

Species Fungorum resolves the standard names, and COL picks up any obscure,
recently described, or historically orphaned species that fall through. For a
list of fungi and plants, a three-backbone chain gives each group its
specialist, with COL catching what falls through both:

```{r fungi-plants-mixed}
mixed <- c(
  "Quercus robur",              # plant
  "Amanita muscaria",           # fungus
  "Lactarius deliciosus",       # fungus
  "Pinus sylvestris",           # plant
  "Russula emetica"             # fungus
)

result <- taxify(mixed, backbone = c("wfo", "fungorum", "col"))
```

Because the genus `Amanita` is not in WFO's coverage table, `"Amanita
muscaria"` is marked out of scope for WFO immediately and passed to the next
backbone without fuzzy matching (see [The genus register](#the-genus-register)).

### Algae: AlgaeBase

AlgaeBase covers micro- and macroalgae, cyanobacteria, and some protists. Its
curation is strong for freshwater and marine microalgae, where generalist
databases often have thin coverage and outdated synonymy.

```{r algae}
algae <- c(
  "Chlamydomonas reinhardtii",
  "Chlorella vulgaris",
  "Ulva lactuca",
  "Fucus vesiculosus",
  "Sargassum muticum"
)

result <- taxify(algae, backbone = c("algaebase", "col"))
result[, c("input_name", "accepted_name", "backbone")]
```

AlgaeBase is licensed CC BY-NC, and taxify prints a license notice during its
download. For commercial applications, COL or WoRMS can serve instead, with
less specialized algal coverage.

### Molecular ecology: NCBI

Species lists from metabarcoding or eDNA studies are often linked to NCBI
accessions. The NCBI backbone aligns taxify's accepted names with the taxonomy
used in GenBank and BOLD.

```{r ncbi-molecular}
edna_hits <- c(
  "Salmo trutta",
  "Phoxinus phoxinus",
  "Anguilla anguilla",
  "Cottus gobio",
  "Lampetra planeri",
  "Chironomus riparius",     # midge (insect)
  "Potamopyrgus antipodarum" # New Zealand mud snail
)

result <- taxify(edna_hits, backbone = c("ncbi", "col"))
result[, c("input_name", "accepted_name", "taxon_id", "backbone")]
```

The `taxon_id` values of NCBI-matched rows are NCBI tax_ids, which link
directly to GenBank records and NCBI taxonomy pages. COL covers names not
found in NCBI, such as taxa without sequenced representatives.

### Auditing the backbone column

`backbone` is a plain character column. In a single-backbone call every
matched row shows the same name; in a chain it records which backbone produced
each match; unmatched rows have `backbone = NA`. It can count how many names
each backbone resolved:

```{r backbone-tally}
result <- taxify(species_list, backbone = c("wfo", "col", "gbif"))
table(result$backbone, useNA = "ifany")
```

or filter to the rows matched by one backbone:

```{r backbone-filter}
wfo_matches <- result[result$backbone == "wfo" & !is.na(result$backbone), ]
col_matches <- result[result$backbone == "col" & !is.na(result$backbone), ]
```

In a list expected to be purely plants, many names matched by COL instead of
WFO point to names outside WFO's scope, perhaps algae classified as plants in
older literature, or animal-associated organisms such as plant parasites.

For a methods section, the unique `backbone_version` strings give the
provenance of the whole result:

```{r backbone-versions}
unique(result$backbone_version[!is.na(result$backbone_version)])
# e.g., c("wfo:2024-12 (2026-04-01)", "col:2025 (2026-04-01)")
```

Each string identifies both the taxonomic source and the snapshot used, so the
result can be reproduced even if a backbone releases a new version between the
analysis and a reviewer's check.

### Backbone-specific extras

Three backbones have enrichment functions that join extra backbone-specific
columns to a taxify result. Each fills only the rows matched by its own
backbone; rows from other backbones get `NA` in the new columns.

`add_wfo_info()` adds `scientificNameID`, `parentNameUsageID`,
`namePublishedIn`, `higherClassification`, `taxonRemarks`, and
`infraspecificEpithet`. `namePublishedIn` is useful for citing original
descriptions, and `higherClassification` holds the full taxonomic hierarchy as
a semicolon-separated string.

```{r add-wfo}
result <- taxify(plants, backbone = "wfo") |>
  add_wfo_info()

result[, c("input_name", "accepted_name", "namePublishedIn")]
```

`add_col_info()` adds COL classification columns (`kingdom`, `phylum`,
`col_class`, `order`), nomenclatural metadata (`notho`, `nomenclaturalCode`,
`nomenclaturalStatus`, `namePublishedIn`), `infraspecificEpithet`, and
SpeciesProfile flags (`is_extinct`, `is_marine`, `is_freshwater`,
`is_terrestrial`). The `class` column is renamed to `col_class` to avoid
conflict with R's `class()` function. The SpeciesProfile flags come from a
separate file in the COL DwC-A archive and can, for instance, exclude extinct
species from a contemporary biodiversity analysis or separate marine from
terrestrial taxa in an estuarine dataset.

```{r add-col}
result <- taxify(species_list, backbone = "col") |>
  add_col_info()

# Check which species are marine
result[result$is_marine == TRUE & !is.na(result$is_marine),
       c("input_name", "accepted_name", "kingdom", "is_marine")]
```

`add_gbif_info()` adds `notho_type` (hybrid type), `nom_status`
(nomenclatural status), `bracket_authorship` (basionym author),
`bracket_year`, `gbif_year`, `name_published_in`, `origin` (how the name
entered the GBIF backbone), and `infra_specific_epithet`. `origin` values such
as `"SOURCE"`, `"DENORMED_CLASSIFICATION"`, or `"VERBATIM_ACCEPTED"` record how
GBIF ingested the name.

```{r add-gbif}
result <- taxify(species_list, backbone = "gbif") |>
  add_gbif_info()

result[, c("input_name", "accepted_name", "origin", "nom_status")]
```

A multi-backbone result can pass through all three; each touches only rows
from its own backbone:

```{r multi-extras}
result <- taxify(species_list, backbone = c("wfo", "col", "gbif")) |>
  add_wfo_info() |>
  add_col_info() |>
  add_gbif_info()
```

The result is a wide data.frame with the union of all extra columns, where
each row has only its own backbone's columns populated and the rest `NA`. For
most workflows the base 27 columns are enough; the extras matter when an
analysis needs nomenclatural details, habitat flags, or publication references
that the standard output does not include.

### Diagnosing out-of-scope names

`lookup_genus()` returns the genus register row for a single genus. The
register is loaded into memory on the first call and cached for the session.

```{r lookup-genus}
lookup_genus("Quercus")
#   genus   kingdom phylum class   order   family   life_form
# 1 Quercus Plantae ...    ...     Fagales Fagaceae vascular plant
```

```{r lookup-genus-animal}
lookup_genus("Panthera")
#   genus    kingdom  phylum   class    order     family  life_form
# 1 Panthera Animalia Chordata Mammalia Carnivora Felidae animal
```

`taxify_register_coverage()` shows which backbones contain a genus and at what
version. If the genus does not appear for the requested backbone, a name in it
is out of that backbone's scope.

```{r register-coverage}
taxify_register_coverage("Quercus")
#     genus   backbone version date_added
# 1 Quercus  col     2025    2026-04-01
# 2 Quercus  gbif    current 2026-04-01
# 3 Quercus  wfo     2024-12 2026-04-01
```

A genus covered by all three backbones can be matched by any of them, while a
genus covered only by GBIF (perhaps a recently described bacterial genus) will
not match against WFO or COL.

When an unmatched name's genus is in the register but covered by none of the
requested backbones, taxify sets `match_type = "out_of_scope"` instead of
`"none"`. An out-of-scope result means the name likely exists in a different
backbone, where a plain `"none"` leaves open whether it is a misspelling or an
invalid name.

```{r out-of-scope}
# Trying to match a marine invertebrate against WFO (plants only)
result <- taxify("Carcinus maenas", backbone = "wfo")
result$match_type
# [1] "out_of_scope"

result$life_form
# [1] "animal"
```

The `life_form` column places the genus among animals, so WFO is the wrong
backbone for this name. In a pipeline, the `"out_of_scope"` rows can be
filtered and re-run against a broader backbone; including the right backbones
in the chain from the start avoids the second pass. The `print()` method of a
taxify result tallies out-of-scope names by `life_form`.

## The backbones

The table summarizes all <!-- manifest:backbone-count -->19<!-- /manifest:backbone-count -->. "Approx. names" is the total
number of name strings in the compiled backbone (accepted names plus
synonyms); the species count is lower because each accepted species may have
several synonym entries pointing to it.

| Backend | Full name | Scope | Approx. names | Source format |
|:--------|:----------|:------|:--------------|:--------------|
| `wfo` | World Flora Online | Vascular plants, bryophytes | ~1.6M | Zenodo ZIP (classification.txt) |
| `col` | Catalogue of Life | All kingdoms | ~5.3M | ChecklistBank DwC-A (Taxon.tsv) |
| `colxr` | Catalogue of Life Extended Release | All kingdoms | ~7.9M | ChecklistBank DwC-A export of the XR dataset |
| `gbif` | GBIF Backbone Taxonomy | All kingdoms | ~6.4M | GBIF simple.txt.gz (30 positional cols) |
| `itis` | Integrated Taxonomic Information System | All kingdoms, US focus | ~990k | SQLite dump from itis.gov |
| `ncbi` | NCBI Taxonomy | All life incl. viruses | ~2.7M | Pipe-delimited .dmp files (taxdump) |
| `ott` | Open Tree of Life | All life (synthetic) | ~3.7M | Pipe-delimited taxonomy.tsv + synonyms.tsv |
| `worms` | World Register of Marine Species | Marine and brackish | ~1.6M | ChecklistBank DwC-A |
| `euromed` | Euro+Med PlantBase | European/Mediterranean plants | ~147k | Semicolon-delimited CSV |
| `fungorum` | Species Fungorum Plus | Fungi | ~315k | ChecklistBank DwC-A |
| `algaebase` | AlgaeBase | Algae and cyanobacteria | ~170k | ChecklistBank DwC-A (CC BY-NC) |
| `fishbase` | FishBase | Fishes | ~100k | rfishbase (load_taxa + synonyms) |
| `sealifebase` | SeaLifeBase | Non-fish marine and aquatic | ~134k | rfishbase (load_taxa + synonyms) |
| `reptiledb` | Reptile Database | Reptiles | ~50k | reptarium taxa.csv + synonym/checklist XLSX |
| `lcvp` | Leipzig Catalogue of Vascular Plants | Vascular plants | ~1.3M | idiv-biodiversity tab_lcvp.rda (R data package) |
| `wcvp` | World Checklist of Vascular Plants (Kew) | Vascular plants | ~1.4M | Kew wcvp.zip (wcvp_names.csv) |
| `mdd` | Mammal Diversity Database | Mammals | ~62k | MDD.zip of CSVs (species + synonyms) |
| `avilist` | AviList (Global Avian Checklist) | Birds | ~41k | AviList extended .xlsx |
| `lpsn` | List of Prokaryotic names with Standing in Nomenclature | Bacteria and archaea | ~45k | ChecklistBank ColDP (NameUsage.tsv) |

WFO (Borsch et al. [2020](https://doi.org/10.1002/tax.12373)) is the standard
reference for plant taxonomy, maintained by the World Flora Online consortium
and updated regularly. The backbone includes all taxonomic ranks from kingdom
down to form, with full synonym resolution and authorship.

COL (Banki et al. [2024](https://doi.org/10.48580/d4t2)) and GBIF (GBIF
Secretariat [2024](https://doi.org/10.15468/39omei)) both cover all kingdoms
with different curation strategies. COL is an expert-curated checklist
assembled from over 160 sector databases, each maintained by a taxonomic
authority for its group. GBIF's backbone is assembled algorithmically from COL,
ITIS, and dozens of other sources, which gives it broader raw coverage (~6.4M
names against COL's ~5.3M) and occasional inconsistencies where source
databases disagree. In practice COL tends to give cleaner synonym resolution
and GBIF tends to match more names.

ITIS ([2025](https://doi.org/10.5066/F7KH0KBK)) was originally developed for
North American fauna and remains particularly strong on freshwater
invertebrates, insects, and US-listed species; its coverage of non-American
taxa is uneven. It is distributed as a SQLite dump, so building the backbone
from source requires the RSQLite package, a dependency the pre-built `.vtr`
avoids.

NCBI Taxonomy is the reference taxonomy for sequence-linked work. Every
GenBank, RefSeq, and BOLD sequence is linked to an NCBI tax_id, which makes
this backbone essential for molecular ecology and metagenomics, and it is the
only backbone that covers bacteria, archaea, and viruses in meaningful depth.
NCBI Taxonomy stores no authorship data, so `authorship` is always `NA` for
NCBI-matched rows.

OTT (Open Tree of Life) is a synthetic taxonomy that merges NCBI, GBIF,
WoRMS, IRMNG, and several other sources into a single tree. It has the
broadest coverage of any single source and cross-references all of its
constituent databases through the `sourceinfo` field. Synthetic taxonomies can
carry conflicts and inconsistencies at the edges, where source databases
disagree about the placement of a taxon.

WoRMS is the authoritative source for marine species, curated by a network of
over 300 taxonomic editors and covering marine, brackish, and some freshwater
species. Beyond basic taxonomy, the WoRMS backbone stores habitat flags
(marine, brackish, freshwater, terrestrial) and extinction status, some of
which are accessible through the COL SpeciesProfile.

Euro+Med PlantBase ([2026](https://europlusmed.org)) is the taxonomic
reference for the flora of Europe, the Mediterranean, and the Caucasus. It
covers all native and introduced vascular plants in its geographic scope
(~49k accepted names, ~83k synonyms). The backbone is built from the 2020 bulk
download, updated by a PESI API delta refresh (April 2026) that resolved 1,014
reclassifications and synonym changes cross-referenced against WFO and POWO.
It suits European vegetation surveys and datasets aligned with the European
Vegetation Archive (EVA). Its data is licensed CC BY-SA 3.0.

Species Fungorum Plus is the specialist reference for fungal taxonomy, with
~315k names curated by the Royal Botanic Gardens, Kew. It covers Ascomycota,
Basidiomycota, and other fungal phyla, including anamorphs and teleomorphs,
and for purely mycological datasets it gives better synonym resolution than
generalist databases.

AlgaeBase covers micro- and macroalgae, cyanobacteria, and some protists. It
is the only backbone licensed CC BY-NC (non-commercial use only); all other
backbones are open-access.

FishBase covers fishes and SeaLifeBase the non-fish marine and aquatic groups
(molluscs, crustaceans, marine mammals, and the rest). Both are compiled from
the rfishbase package rather than a file download, taking accepted species
from `load_taxa()` and synonym links from `synonyms()`. They carry
species-rank rows only, so their genera enter the genus register derived from
those species. The data is CC BY-NC 3.0.

The Reptile Database (Uetz et al. [2026](http://www.reptile-database.org)) is
the taxonomic reference for reptiles, with ~50k names covering snakes,
lizards, turtles, crocodilians, and the tuatara. It is built from the
reptarium bulk exports: the current accepted-species list, the periodic
synonym snapshot, and the family-to-order checklist. The source carries no
kingdom or class field, so the backbone stamps Animalia / Chordata / Reptilia
on every row. The data is CC BY 4.0.

The Mammal Diversity Database is the American Society of Mammalogists'
reference mammal taxonomy, distributed as a zip of CSVs holding the accepted
species with their full higher classification and every name ever applied to
them. The data is MIT-licensed.

AviList is the global bird checklist that merged the long-standing IOC,
Clements, and BirdLife split, shipped as one Excel workbook covering order
through subspecies. It publishes no synonym table, so the backbone derives
homotypic synonyms from the `Protonym` column: where a species has since moved
genus, its original combination is a synonym of the current name
(`Parus caeruleus` -> `Cyanistes caeruleus`). The data is CC BY 4.0.

LPSN is the nomenclatural authority for prokaryotes, hosted by DSMZ, and it
records whether a bacterial or archaeal name is validly published, which NCBI,
GBIF, and OTT leave open. LPSN's own download route sits behind a free DSMZ
account, so the backbone is built from the open ColDP mirror on GBIF
ChecklistBank. The data is CC BY-SA 4.0.

## Output differences between backbones

All <!-- manifest:backbone-count -->19<!-- /manifest:backbone-count --> backbones produce the same 27-column output schema, so
downstream code does not need to know which backbone produced a match. The
content of those columns varies in several ways.

**Authorship.** WFO's `scientificName` is already canonical (no authorship
appended), so the `authorship` column comes from a separate
`scientificNameAuthorship` field. COL and WoRMS store the full
`scientificName` with authorship included; taxify strips it at build time to
produce the canonical name used for matching and stores the stripped
authorship separately. NCBI and OTT have no authorship data, so `authorship`
is always `NA` for those backbones. GBIF and ITIS provide authorship, Euro+Med
provides it from its `AuthorString` field, and Species Fungorum and AlgaeBase
from their DwC-A archives.

**Taxon IDs.** Each backbone uses a different identifier system. WFO IDs look
like `"wfo-0000000123"`. COL IDs are opaque alphanumeric strings like
`"4LHBG"`. GBIF uses integer keys (`"2878688"`), ITIS TSN integers
(`"183671"`), NCBI its Taxonomy IDs (`"9606"`), and OTT its own IDs
(`"770315"`).
WoRMS uses AphiaIDs extracted from LSIDs: at build time taxify strips the
`urn:lsid:marinespecies.org:taxname:` prefix and stores just the numeric ID.
Euro+Med uses `TaxonUsageID` integers from the PlantBase export, and Species
Fungorum and AlgaeBase use ChecklistBank dataset-specific IDs. All IDs are
stored as character strings in `taxon_id` and `accepted_id`, but their format
is backbone-specific and meaningful only within that backbone's ecosystem: a
`taxon_id` from WFO cannot be looked up in the COL database, and vice versa.

**Classification depth.** The base output always includes `family` and
`genus`. WFO provides these directly from its classification file. COL stores
the full Linnaean hierarchy (kingdom through order) in the Taxon.tsv, though
the extra columns require `add_col_info()`. GBIF provides family through a
denormalized `family_key` self-join at build time. ITIS, NCBI, and OTT resolve
family and genus by parent-hierarchy walks during backbone compilation, which
traverse up to 25 levels of the taxonomic tree. WoRMS has denormalized
classification columns in its DwC-A, and Euro+Med resolves family and genus by
a hierarchy walk on `IsChildTaxonOfID`. The genus register fills in the higher
classification fields (`kingdom_group`, `taxon_group`, `life_form`) for all
backbones.

**Synonym handling.** WFO and COL use the Darwin Core field
`acceptedNameUsageID` to point from a synonym row to its accepted name. GBIF
encodes synonyms through `parent_key` pointing to the accepted taxon. NCBI
represents synonyms as alternative name strings for the same `tax_id`; at
build time taxify emits these as separate rows with synthetic IDs of the form
`"123456_syn_1"`, `"123456_syn_2"`, etc. OTT uses a separate `synonyms.tsv`
file with explicit synonym-to-accepted mappings. All of these representations
are normalized at build time into the same `is_synonym` + `accepted_name` +
`accepted_id` schema.

**Synonym chains.** Some backbones contain chained synonyms, where synonym A
points to synonym B, which points to accepted name C. taxify resolves these
chains at build time (up to 10 hops), so `accepted_name` always points to the
terminal accepted name, at no cost to query-time performance.

## The genus register

The genus register is a unified index of the genera across every supported
backbone. It holds 503,262 genera, each with its family, higher classification
(kingdom through order, where available), and a `life_form` label (e.g.,
`"vascular plant"`, `"animal"`, `"fungus"`). Where two backbones disagree
about which family a genus belongs to, the classification is resolved by
priority: WoRMS, COL Extended Release, COL, WCVP, Reptile Database, MDD,
AviList, LPSN, GBIF, Euro+Med, LCVP, ITIS, NCBI, OTT, WFO, FishBase,
SeaLifeBase, Species Fungorum, AlgaeBase. If COL and WFO disagree, COL's
assignment wins.

taxifydb builds the register over that fixed backbone set and publishes it as
a versioned asset, the same way it publishes each backbone, so it arrives by
download. taxify resolves it on first use: local disk, then the manifest
download, then a local build through `taxify_build_register()` (equivalently
`taxify_download("register")`) if taxifydb is installed. A register built from
whichever backbones a machine happened to have installed made the same code
report different `kingdom_group` and `life_form` labels on two machines, which
the published asset prevents. A `backend_coverage` asset ships alongside it,
one row per genus and backbone, and is what `taxify_register_coverage()`
reads.

The register serves two purposes in matching. It supplies the `life_form`,
`kingdom_group`, and `taxon_group` columns for every matched name, regardless
of which backbone matched it, so results can be stratified by broad taxonomic
group without looking up each family. It also enables out-of-scope detection:
before fuzzy matching begins, taxify checks whether an unmatched name's genus
is known to the register but absent from the coverage table of every
requested backbone, and if so marks the name `"out_of_scope"` immediately.
This skips fuzzy matching against a backbone that could never produce a match
and gives a more informative signal than a plain `"none"`.

## Choosing backbones

The right backbone depends on the taxonomic scope of the data. The general
rule is specialist backbones first and generalist backbones second: lead with
the backbone whose taxonomic opinion we trust most for the dominant taxon
group, and add broader backbones as fallbacks for the remainder. The
`backbone` column then records which taxonomic opinion was applied to each
name.

- For pure vascular plant lists, use `backbone = "wfo"`. If some names fall
  through (horticultural cultivars, nomenclaturally complex genera, or names
  from older floras that use outdated synonymy), add COL:
  `backbone = c("wfo", "col")`.
- For European vegetation data, use `backbone = c("euromed", "wfo")`. Euro+Med
  PlantBase is the taxonomic reference used by EVA and covers all native and
  introduced vascular plants of Europe, the Mediterranean, and the Caucasus.
  Leading with it makes European synonym resolution follow Euro+Med's opinion,
  with WFO as fallback for non-European taxa or names outside its scope.
  Euro+Med data is CC BY-SA 3.0.
- For pure marine or aquatic lists, use `backbone = "worms"`, which is curated
  by domain experts and includes habitat and extinction flags. For estuarine
  or transitional lists with some terrestrial taxa, add COL:
  `backbone = c("worms", "col")`.
- For pure fungal lists, use `backbone = "fungorum"`, with COL as fallback for
  obscure or recently described species: `backbone = c("fungorum", "col")`.
- For pure algal lists, use `backbone = "algaebase"`, with COL or WoRMS as
  fallback: `backbone = c("algaebase", "col")`. AlgaeBase is CC BY-NC.
- For mixed-kingdom ecological datasets, use `backbone = c("col", "gbif")`, or
  lead with a specialist backbone for the dominant taxon group: `c("wfo",
  "col")` for a plant-dominated dataset with some animals and fungi,
  `c("worms", "col", "gbif")` for a marine biodiversity survey, and
  `c("wfo", "fungorum", "col")` for a forest inventory that includes trees,
  fungi, insects, and epiphytes.
- For molecular and sequence-linked work, use `backbone = "ncbi"`, the
  reference for GenBank, BOLD, and other sequence databases, which covers the
  bacteria, archaea, and viruses other backbones lack. For mixed molecular and
  ecological work: `backbone = c("ncbi", "col")`.
- For phylogenetic studies, use `backbone = "ott"`, the backbone of the Open
  Tree of Life. Its cross-references to NCBI, GBIF, WoRMS, and IRMNG bridge
  different identifier systems.
- For maximum coverage, use `backbone = c("col", "gbif")`. COL provides
  expert-curated taxonomy for ~5.3M names and GBIF's backbone adds ~6.4M names
  from additional sources; together they cover virtually all described species
  with a nomenclatural record. This combination is a reasonable default when
  the taxonomic composition of the dataset is unknown.

## Performance

Backbone size affects download time and, to a lesser extent, matching speed:
WFO (~1.6M names) matches faster than GBIF (~6.4M names) for the same query.
taxify uses index-accelerated genus-blocked joins at the C level (via vectra),
so a list of 5,000 names resolves against GBIF in under a second on modern
hardware. The difference only becomes noticeable at scale (100k+ names) or
with heavy fuzzy matching against a large backbone.

In a multi-backbone chain, putting the most likely backbone first saves time,
because names matched by the first backbone skip all later ones. If 90% of a
list is plants, `c("wfo", "col")` is faster than `c("col", "wfo")`: WFO is
smaller and resolves most names on the first pass, and COL, which is larger,
only processes the remaining 10%.

Fuzzy matching is the most expensive step. It runs a genus-blocked fuzzy join
with multi-threaded string distance computation. For names with misspelled
genera, where genus blocking cannot help, backbones configured for it fall
back to a 2-character prefix block that catches most genus-level typos while
keeping the search space manageable.

## Versions and reproducibility

taxify checks for backbone updates once per R session. The first `taxify()`
call in a session fetches the manifest from GitHub, compares each requested
backbone's local version against the latest release, and downloads a new
version only if one exists. If the network is unavailable, taxify falls back
to the bundled manifest and uses whatever local copy is on disk. The version
check and any update are logged to the console with the old and new version
numbers, and the backbone version never changes mid-session. The manifest,
which maps backbone names to their download URLs and latest versions, is
cached per session and can be refreshed with `taxify_refresh_manifest()`.

For a published analysis, the `backbone_version` strings belong in the methods
section or supplementary material. A project that needs exact reproducibility
can pin every backbone with `taxify_download(version = )`, as in
[Downloading backbones](#downloading-backbones), and never use the "latest"
slot, which keeps tracking new releases independently. A project that prefers
to stay current can rely on the default "latest" behavior and cite the
`backbone_version` strings from the output.

Backbones with large source files can also be built from source:
`taxify_build("gbif")` downloads the raw 1.5 GB `simple.txt.gz` from GBIF,
parses all 30 positional columns, denormalizes the family hierarchy through
self-joins, and compiles the result into `.vtr` format. This is slower than
downloading the pre-built file and produces the same output; it is mainly
useful for CI pipelines or for customizing the compilation step.

## References

Banki O, Roskov Y, Doring M, Ower G, Hernandez Robles DR, Plata Corredor CA,
Stjernegaard Jeppesen T, Orrell TM, Pugh D, Kostichka J, et al. (2024).
*Catalogue of Life Checklist*. <https://doi.org/10.48580/d4t2>

Borsch T, Berendsohn W, Dalcin E, Delmas M, Demissew S, Elliott A, Fritsch P,
Fuchs A, Geltman D, Guner A, et al. (2020). World Flora Online: Placing
taxonomists at the heart of a definitive and comprehensive global resource on
the world's plants. *Taxon* 69: 1311-1341.
<https://doi.org/10.1002/tax.12373>

Euro+Med (2026). *Euro+Med PlantBase: the information resource for
Euro-Mediterranean plant diversity*. <https://europlusmed.org>, accessed
2026-07-11.

GBIF Secretariat (2024). *GBIF Backbone Taxonomy*.
<https://doi.org/10.15468/39omei>

ITIS (2025). *Integrated Taxonomic Information System*.
<https://doi.org/10.5066/F7KH0KBK>

Uetz P, Freed P, Aguilar R, Reyes F, Kudera J, Hosek J (2026). *The Reptile
Database*. <http://www.reptile-database.org>
