Interface for dimensionality reduction methods
The dr() function apply the function, fun,
and (if any) parameters to analyze the input ‘data’.
The data can be original data matrix (or data frame), or distance matrix (or distance object), depending on the requirement of the DR method (specify by the ‘fun’ parameter).
Methods supported by the dr() function can be listed via
the available_methods():
## The `dr()` function works for the following methods that
## require data matrix (or data frame) as input:
## + stats::prcomp()
## + Rtsne::Rtsne()
## + uwot::umap()
## + uwot::tumap()
## + uwot::lvish()
## The `dr()` function works for the following methods that
## require distance matrix (or distance object) as input:
## + stats::cmdscale()
## + MASS::sammon()
## + vegan::metaMDS()
## + ape::pcoa()
## + smacof::mds()
## + vegan::wcmdscale()
## + ecodist::pco()
## + labdsv::pco()
## + ade4::dudi.pco()
##
## Note: any other function that returns a numeric matrix with
## at least two columns is also accepted (see `?dr_extract`).
## Some methods need their own arguments, which are passed through
## `...` of `dr()`, e.g. `dr(d, ade4::dudi.pco, scannf = FALSE)` or
## `dr(x, Rtsne::Rtsne, check_duplicates = FALSE)`.
## `ecodist::pco()` returns every eigenvector, so `dr()` keeps all
## n - 1 of them; use `Dim1`/`Dim2` for plotting.
The listing is grouped by the kind of input a method expects;
available_methods("data") and
available_methods("distance") show one group each:
## The `dr()` function works for the following methods that
## require distance matrix (or distance object) as input:
## + stats::cmdscale()
## + MASS::sammon()
## + vegan::metaMDS()
## + ape::pcoa()
## + smacof::mds()
## + vegan::wcmdscale()
## + ecodist::pco()
## + labdsv::pco()
## + ade4::dudi.pco()
##
## Note: any other function that returns a numeric matrix with
## at least two columns is also accepted (see `?dr_extract`).
## Some methods need their own arguments, which are passed through
## `...` of `dr()`, e.g. `dr(d, ade4::dudi.pco, scannf = FALSE)` or
## `dr(x, Rtsne::Rtsne, check_duplicates = FALSE)`.
## `ecodist::pco()` returns every eigenvector, so `dr()` keeps all
## n - 1 of them; use `Dim1`/`Dim2` for plotting.
A distance object is passed to dr() in the same way as a
data matrix, and the result is visualized identically:
d <- dist(iris[, 1:4])
y <- dr(d, stats::cmdscale)
autoplot(y, aes(color = Species), metadata = iris[, 5, drop = FALSE]) + theme_dr()Some methods require their own arguments, which are passed through
...:
autoplot(dr(iris[, 1:4], prcomp, scale. = TRUE),
aes(color = Species), metadata = iris[, 5, drop = FALSE]) + theme_dr()The same route carries arguments that a method needs to run
non-interactively,
e.g. dr(iris[, 1:4], Rtsne::Rtsne, check_duplicates = FALSE)
or dr(d, ade4::dudi.pco, scannf = FALSE) – without
scannf = FALSE that method stops at an interactive prompt.
available_methods() repeats these reminders after the
listing.
Visualization
The tidydr package extends ggplot() to
support the output of dr() function.
Associated data (e.g., group information) can be integrated to scale
the color, shape or size of data points. The data should be provided via
the metadata parameter. It allows a vector (will be stored
in .group) or a data frame.
library(ggplot2)
## metadata as a vector
ggplot(x, aes(Dim1, Dim2), metadata=iris$Species) +
geom_point(aes(color=.group))Users can use autoplot() as a shortcut. This package
provide a theme, theme_dr(), to allow using shorten
axes.
## metadata as a data frame
autoplot(x, aes(color=Species), metadata = iris[, 5, drop=FALSE]) +
theme_dr()Comparing several methods
dr_compare() applies several methods to the same data
and returns them together. Each method is evaluated independently, so a
method that fails is reported in the summary instead of
aborting the whole call.
r <- dr_compare(iris[, 1:4],
funs = list(prcomp = stats::prcomp,
cmdscale = function(z) stats::cmdscale(dist(z))),
dim = 1:2)
r$summary## method status n k has_eigenvalue has_stress error
## 1 prcomp ok 150 4 TRUE FALSE <NA>
## 2 cmdscale ok 150 2 FALSE FALSE <NA>
The summary lists which methods ran, how many dimensions
each produced, and the error message of any method that did not.
Clustering and silhouette widths
nk() computes the average silhouette width for one or
more values of k. It uses cluster::pam() by
default, and any other clustering function can be supplied through
fun:
fun is given as a function, not a call. A method that
takes the number of clusters as its second argument, or through an
argument named k or centers, is used
directly:
A method that takes no cluster count, such as
stats::hclust(), is fitted once and cut with
stats::cutree() for each k:
The silhouette width is computed with
cluster::silhouette() for every clusterer except the
default pam(), whose own silhouette information is reused;
silinfo_widths() and
autoplot(type = "silhouette") work the same way in either
case.
Passing a k that is present in the object draws the
samples in the reduced space, coloured by cluster:
The per-sample widths of a single k are available from
silinfo_widths(). It returns a data frame, and the
per-cluster averages are attached as an attribute:
## sample cluster neighbor sil_width
## 1 8 1 2 0.8539051
## 2 1 1 2 0.8529551
## 3 50 1 2 0.8520984
## 4 18 1 2 0.8510183
## 5 40 1 2 0.8503323
## 6 41 1 2 0.8494158
## [1] 0.7981405 0.4173199 0.4511051
Setting type = "silhouette" draws the classic
per-cluster silhouette bar plot:
Method-specific information
A dr_extract() method may return a
sample_info element alongside the coordinates.
fortify() merges any element of sample_info
that holds one value per sample into the returned data frame as an
additional column, so it can be mapped in aes(). None of
the dr_extract() methods shipped with tidydr
returns one, so sample_info is NULL unless you
write a method yourself; see ?dr and
?dr_extract for the details.