Interface for dimensionality reduction methods

library(tidydr)
x <- dr(data = iris[,1:4], fun = prcomp)

The dr() function apply the function, fun, and (if any) parameters to analyze the input ‘data’.

The data can be original data matrix (or data frame), or distance matrix (or distance object), depending on the requirement of the DR method (specify by the ‘fun’ parameter).

Methods supported by the dr() function can be listed via the available_methods():

available_methods()
## The `dr()` function works for the following methods that
## require data matrix (or data frame) as input:
##   +  stats::prcomp()
##   +  Rtsne::Rtsne()
##   +  uwot::umap()
##   +  uwot::tumap()
##   +  uwot::lvish()
## The `dr()` function works for the following methods that
## require distance matrix (or distance object) as input:
##   +  stats::cmdscale()
##   +  MASS::sammon()
##   +  vegan::metaMDS()
##   +  ape::pcoa()
##   +  smacof::mds()
##   +  vegan::wcmdscale()
##   +  ecodist::pco()
##   +  labdsv::pco()
##   +  ade4::dudi.pco()
## 
## Note: any other function that returns a numeric matrix with
## at least two columns is also accepted (see `?dr_extract`).
## Some methods need their own arguments, which are passed through
## `...` of `dr()`, e.g. `dr(d, ade4::dudi.pco, scannf = FALSE)` or
## `dr(x, Rtsne::Rtsne, check_duplicates = FALSE)`.
## `ecodist::pco()` returns every eigenvector, so `dr()` keeps all
## n - 1 of them; use `Dim1`/`Dim2` for plotting.

The listing is grouped by the kind of input a method expects; available_methods("data") and available_methods("distance") show one group each:

available_methods("distance")
## The `dr()` function works for the following methods that
## require distance matrix (or distance object) as input:
##   +  stats::cmdscale()
##   +  MASS::sammon()
##   +  vegan::metaMDS()
##   +  ape::pcoa()
##   +  smacof::mds()
##   +  vegan::wcmdscale()
##   +  ecodist::pco()
##   +  labdsv::pco()
##   +  ade4::dudi.pco()
## 
## Note: any other function that returns a numeric matrix with
## at least two columns is also accepted (see `?dr_extract`).
## Some methods need their own arguments, which are passed through
## `...` of `dr()`, e.g. `dr(d, ade4::dudi.pco, scannf = FALSE)` or
## `dr(x, Rtsne::Rtsne, check_duplicates = FALSE)`.
## `ecodist::pco()` returns every eigenvector, so `dr()` keeps all
## n - 1 of them; use `Dim1`/`Dim2` for plotting.

A distance object is passed to dr() in the same way as a data matrix, and the result is visualized identically:

d <- dist(iris[, 1:4])
y <- dr(d, stats::cmdscale)
autoplot(y, aes(color = Species), metadata = iris[, 5, drop = FALSE]) + theme_dr()

Some methods require their own arguments, which are passed through ...:

autoplot(dr(iris[, 1:4], prcomp, scale. = TRUE),
         aes(color = Species), metadata = iris[, 5, drop = FALSE]) + theme_dr()

The same route carries arguments that a method needs to run non-interactively, e.g. dr(iris[, 1:4], Rtsne::Rtsne, check_duplicates = FALSE) or dr(d, ade4::dudi.pco, scannf = FALSE) – without scannf = FALSE that method stops at an interactive prompt. available_methods() repeats these reminders after the listing.

Visualization

The tidydr package extends ggplot() to support the output of dr() function.

Associated data (e.g., group information) can be integrated to scale the color, shape or size of data points. The data should be provided via the metadata parameter. It allows a vector (will be stored in .group) or a data frame.

library(ggplot2)
## metadata as a vector
ggplot(x, aes(Dim1, Dim2), metadata=iris$Species) + 
  geom_point(aes(color=.group))

Users can use autoplot() as a shortcut. This package provide a theme, theme_dr(), to allow using shorten axes.

## metadata as a data frame
autoplot(x, aes(color=Species), metadata = iris[, 5, drop=FALSE]) +
  theme_dr()

Comparing several methods

dr_compare() applies several methods to the same data and returns them together. Each method is evaluated independently, so a method that fails is reported in the summary instead of aborting the whole call.

r <- dr_compare(iris[, 1:4],
                funs = list(prcomp = stats::prcomp,
                            cmdscale = function(z) stats::cmdscale(dist(z))),
                dim = 1:2)
r$summary
##     method status   n k has_eigenvalue has_stress error
## 1   prcomp     ok 150 4           TRUE      FALSE  <NA>
## 2 cmdscale     ok 150 2          FALSE      FALSE  <NA>
autoplot(r) + theme_dr()

The summary lists which methods ran, how many dimensions each produced, and the error message of any method that did not.

Clustering and silhouette widths

nk() computes the average silhouette width for one or more values of k. It uses cluster::pam() by default, and any other clustering function can be supplied through fun:

si <- nk(iris[, 1:4], 2:4)
autoplot(si)

fun is given as a function, not a call. A method that takes the number of clusters as its second argument, or through an argument named k or centers, is used directly:

si_km <- nk(iris[, 1:4], 2:4, fun = stats::kmeans)
autoplot(si_km)

A method that takes no cluster count, such as stats::hclust(), is fitted once and cut with stats::cutree() for each k:

si_hc <- nk(iris[, 1:4], 2:4, fun = stats::hclust)
autoplot(si_hc)

The silhouette width is computed with cluster::silhouette() for every clusterer except the default pam(), whose own silhouette information is reused; silinfo_widths() and autoplot(type = "silhouette") work the same way in either case.

Passing a k that is present in the object draws the samples in the reduced space, coloured by cluster:

autoplot(si, k = 3) + theme_dr()

The per-sample widths of a single k are available from silinfo_widths(). It returns a data frame, and the per-cluster averages are attached as an attribute:

w <- silinfo_widths(si, 3)
head(w)
##   sample cluster neighbor sil_width
## 1      8       1        2 0.8539051
## 2      1       1        2 0.8529551
## 3     50       1        2 0.8520984
## 4     18       1        2 0.8510183
## 5     40       1        2 0.8503323
## 6     41       1        2 0.8494158
attr(w, "clus.avg.widths")
## [1] 0.7981405 0.4173199 0.4511051

Setting type = "silhouette" draws the classic per-cluster silhouette bar plot:

autoplot(si, k = 3, type = "silhouette") + theme_dr()

Method-specific information

A dr_extract() method may return a sample_info element alongside the coordinates. fortify() merges any element of sample_info that holds one value per sample into the returned data frame as an additional column, so it can be mapped in aes(). None of the dr_extract() methods shipped with tidydr returns one, so sample_info is NULL unless you write a method yourself; see ?dr and ?dr_extract for the details.