| Title: | Unify Dimensionality Reduction Results |
| Version: | 0.0.7 |
| Description: | Dimensionality reduction is widely used in many domains for analyzing and visualizing high-dimensional data. 'tidydr' provides uniform output and is compatible with multiple methods, including 'prcomp', 'cmdscale', 'Rtsne', 'umap' and 'metaMDS'. Any function returning a numeric matrix can also be used. The unified result can be visualized directly with 'ggplot2', and several methods can be run and compared in a single call. |
| Imports: | cluster, ggfun, ggplot2, grid, rlang, stats, utils |
| Suggests: | knitr, rmarkdown, prettydoc, SingleCellExperiment, SummarizedExperiment, ape, MASS, Rtsne, uwot, vegan, smacof, ecodist, ade4, labdsv, testthat (≥ 3.0.0) |
| VignetteBuilder: | knitr |
| ByteCompile: | true |
| License: | Artistic-2.0 |
| URL: | https://github.com/YuLab-SMU/tidydr/ |
| BugReports: | https://github.com/YuLab-SMU/tidydr/issues |
| Encoding: | UTF-8 |
| Config/testthat/edition: | 3 |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-10-09 12:05:47 UTC; wang |
| Author: | Guangchuang Yu |
| Maintainer: | Guangchuang Yu <guangchuangyu@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-09 12:30:02 UTC |
tidydr: Unify Dimensionality Reduction Results
Description
Dimensionality reduction is widely used in many domains for analyzing and visualizing high-dimensional data. 'tidydr' provides uniform output and is compatible with multiple methods, including 'prcomp', 'cmdscale', 'Rtsne', 'umap' and 'metaMDS'. Any function returning a numeric matrix can also be used. The unified result can be visualized directly with 'ggplot2', and several methods can be run and compared in a single call.
Author(s)
Maintainer: Guangchuang Yu guangchuangyu@gmail.com (ORCID) [copyright holder]
Authors:
Guangchuang Yu guangchuangyu@gmail.com (ORCID) [copyright holder]
Shuangbin Xu xshuangbin@163.com (ORCID)
See Also
Useful links:
List dimensionality reduction methods currently available
Description
This function shows available methods that worked for dr() function.
Usage
available_methods(method = "all")
Arguments
method |
one of 'data', 'distance' or 'all' (default) |
Value
A character vector of available DR methods
Author(s)
Lang Zhou and Guangchuang Yu
Examples
available_methods()
dr
Description
dimensional reduction
Usage
dr(data, fun, ...)
Arguments
data |
input data |
fun |
function to perform dimensional reduction |
... |
additional parameters passed to 'fun' |
Details
This function call the user-provided function ('fun') to perform dimensional reduction on the input data ('data')
sample_info is an optional list that a dr_extract() method may return
alongside the coordinates. It is meant to carry method-specific, per-sample
information (e.g. local density, k-nearest-neighbour indices or cluster
labels). Every element of sample_info that holds one value per sample – a
vector or factor of length nrow(data) – is merged into the data.frame
returned by fortify(), using the element name as the column name, so that
it can be mapped in autoplot(). Elements that do not hold one value per
sample, and elements whose name is already taken by a coordinate or a
metadata column, are not merged and a warning is issued. None of the
dr_extract() methods shipped with the package returns a sample_info
element, so sample_info is NULL unless a custom dr_extract() method
supplies one.
Value
a DrResult object, which contains 'data' (original data), 'drdata' (coordination after dimensionality reduction), eigenvalue (standard deviation explained by each dimension), stress (evaluate the effect of dimensionality reduction) and 'sample_info' (method-specific additional information, see Details)
Author(s)
Guangchuang Yu
Examples
x = dr(iris[,1:4], prcomp)
autoplot(x, aes(color=.group), metadata=iris$Species)
dr_compare
Description
Compare several dimensionality reduction methods
Usage
dr_compare(
data,
funs = list(prcomp = stats::prcomp),
dim = 1:2,
metadata = NULL,
...
)
Arguments
data |
input data. It follows the same contract as |
funs |
a named list of methods to compare. Each element is either a
function, or a list with components |
dim |
the two dimensions to extract, as a numeric vector of length 2.
Defaults to |
metadata |
optional sample-level metadata. It is forwarded to
|
... |
additional arguments passed to every method in |
Details
dr_compare() applies several dimensionality reduction (DR) methods to the
same input data and collects their results, so that the methods can be
compared side by side. Every method is evaluated independently: a method that
fails, or that returns fewer dimensions than requested, does not interrupt
the others. The outcome of every method is reported in the summary
component, which is the main purpose of this function.
A method that fails is never given placeholder coordinates. The failure is
recorded in summary$error, and the corresponding element of results is
the condition (error) object, so it can be inspected further.
Value
A DrCompare object, a list with the following components:
resultsa named list with one element per method. For a method that ran successfully it is a
DrResult(seedr()); for a method that failed it is the condition (error) object.summarya
data.framewith one row per method and the columnsmethod,status,n,k,has_eigenvalue,has_stressanderror.statusis"ok"if the method ran and returned at leastmax(dim)dimensions,"insufficient_dims"if it ran but returned too few dimensions, and"failed"if it raised an error.nandkare the number of samples and the number of extracted dimensions,has_eigenvalueandhas_stressindicate whether the method produced those quantities, anderrorholds the error message of a failed method (NAotherwise).dim,metadata,funsthe corresponding inputs, kept so that
autoplot()can rebuild the plot.
Author(s)
Guangchuang Yu
See Also
Examples
x <- dr_compare(iris[, 1:4], list(pca = stats::prcomp))
x$summary
autoplot(x)
dr_extract
Description
dr_extract generic
Usage
dr_extract(result)
Arguments
result |
DrResult object |
Details
dr_extract() is an S3 generic. Methods are provided for the result classes
of the supported dimensionality reduction functions. In addition, a fallback
method is provided for plain numeric matrices: any function passed to dr()
that returns a numeric matrix with at least two columns is accepted, and its
columns are used as the reduced coordinates.
The list returned by a method must contain drdata, the reduced coordinates:
either a data.frame, or a numeric matrix, with one row per sample and at least
two columns (one per retained dimension). A matrix is converted to a
data.frame, so that fortify() – a 'ggplot2' generic – always returns a
data.frame. The columns are renamed Dim1, Dim2, ... by dr(), and a
drdata that is neither a data.frame nor a numeric matrix is rejected.
eigenvalue, stress and sample_info are optional. sample_info is a list
of method-specific, per-sample vectors (e.g. local density or cluster labels),
which fortify() merges into its result as extra columns; see dr() for the
details.
Value
a list that contains components to construct a 'DrResult' object.
Author(s)
Guangchuang Yu
element_line2
Description
element_line2 for drawing shorten axis lines
Usage
element_line2(
colour = NULL,
size = NULL,
linetype = NULL,
lineend = NULL,
color = NULL,
arrow = NULL,
inherit.blank = FALSE,
id,
xlength = 0.3,
ylength = 0.3,
...
)
Arguments
colour |
line colour |
size |
line size in pts |
linetype |
line type |
lineend |
line end style (round, butt, square) |
color |
aliase to colour |
arrow |
arrow specification, as created by 'grid::arrow()' |
inherit.blank |
whether inherit 'element_blank' |
id |
1 or 2, 1 for axis.line.x.bottom and 2 for axis.line.y.left, only these two axes supported |
xlength |
length of x axis |
ylength |
length of y axis |
... |
additional parameters |
Value
element_line2 object, which is a tailored element_line object
Author(s)
Guangchuang Yu
nk
Description
Choose best K (number of clusters)
Usage
nk(data, k, fun = pam, ...)
Arguments
data |
input data (a matrix, data frame, or 'dist' object) |
k |
a vector of candidate number of clusters |
fun |
a function to perform the clustering. It must accept the number of
clusters as its second argument, or through an argument named Whether |
... |
additional parameters passed to Note: by default this function calls |
Details
This function calculate the silhouette scores of each K (number of clusters).
The output object can be used to choose the best K (via summary() or autoplot() methods)
The silhouette width is computed with cluster::silhouette() from the cluster
labels and the dissimilarity of data (stats::dist(data), unless data is
already a 'dist' object, in which case it is used as is). With the default
fun = cluster::pam the silhouette information that pam() computes itself is
reused verbatim, so the default result is unchanged.
Each element of $silinfo follows the shape produced by cluster::pam():
widths (a matrix with columns cluster, neighbor and sil_width),
clus.avg.widths (per-cluster averages) and avg.width (the overall average).
The row order of widths is the one produced by the clusterer, so it should be
aligned by sample name rather than by position.
Value
a silinfo object, which contains 'data' (original data), 'silinfo' (silhouette scores), and k (the input k vector)
Author(s)
Guangchuang Yu
Examples
x <- nk(iris[,-5], 2:8)
summary(x)
# to visualize the average silhouete score (y axis) with k (x axis)
autoplot(x)
# to visualize a PCA plot color by the choosing k
autoplot(x, k=3)
# another clustering method
x2 <- nk(iris[,-5], 2:4, fun = stats::kmeans)
summary(x2)
# a hierarchical method is cut automatically
x3 <- nk(iris[,-5], 2:4, fun = stats::hclust)
summary(x3)
Objects exported from other packages
Description
These objects are imported from other packages. Follow the links below to see their documentation.
- ggplot2
silinfo_widths
Description
Per-sample silhouette widths
Usage
silinfo_widths(x, k)
Arguments
x |
a |
k |
a single value among the candidate |
Details
Extract the fine-grained silhouette information of one k from an nk()
result: one row per sample, with the cluster the sample was assigned to, its
neighbour cluster and its silhouette width.
Value
a data.frame with one row per sample and the columns sample,
cluster, neighbor and sil_width. The rows keep the order in which the
clusterer returned them – for cluster::pam() that is by cluster and, within
a cluster, by decreasing sil_width – so the rows should be aligned by
sample rather than by position. The per-cluster average silhouette widths
are attached as the "clus.avg.widths" attribute, and are also available as
x$silinfo[[i]]$clus.avg.widths, where i is the position of k in x$k.
Author(s)
Guangchuang Yu
See Also
Examples
x <- nk(iris[,-5], 2:4)
head(silinfo_widths(x, 3))
attr(silinfo_widths(x, 3), "clus.avg.widths")
theme_dr
Description
Dimensional reduction scatter plot axis theme
Usage
theme_dr(
xlength = 0.3,
ylength = 0.3,
arrow = grid::arrow(length = unit(0.15, "inches"), type = "closed")
)
Arguments
xlength |
length of x axis |
ylength |
length of y axis |
arrow |
arrow specification, as created by 'grid::arrow()' |
Value
a theme object with shorten axes
Author(s)
Guangchuang Yu
theme_noaxis
Description
theme that remove axis
Usage
theme_noaxis(...)
Arguments
... |
additional theme setting |
Value
a theme object that disable axes
Author(s)
Guangchuang Yu