---
title: "Reporting with tidylearn"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Reporting with tidylearn}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
# gt is in Suggests and is the subject of this vignette, so skip the
# whole thing rather than fail the build when it is absent
has_gt <- requireNamespace("gt", quietly = TRUE)

knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  fig.width = 7,
  fig.height = 5,
  message = FALSE,
  warning = FALSE,
  eval = has_gt
)
```

```{r, echo = FALSE, results = "asis", eval = TRUE}
if (!has_gt) {
  cat(
    "> **Note:** the gt package is not installed, so the examples below",
    "are shown without output.\n"
  )
}
```

## Overview

Two families turn a fitted model into something publishable. `plot()` and
the `tl_plot_*()` functions return ggplot2 objects, with the exceptions
listed under [Interactive Reporting with
plotly](#interactive-reporting-with-plotly); `tl_table()` and the
`tl_table_*()` functions return `gt` tables. Both dispatch on model type, so
the same call covers a forest and a lasso fit, and both hand back an object
you can keep editing rather than printed output you cannot.

```{r setup}
library(tidylearn)
library(dplyr)
library(ggplot2)
library(gt)
```

## Plots

`plot()` picks the visualisation from the model type, and `type` narrows it
further where a model supports more than one view.

### Regression

```{r plot-regression}
model_reg <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")

# Actual vs predicted — one call
plot(model_reg, type = "actual_predicted")
```

### Classification

```{r plot-classification}
split <- tl_split(iris, prop = 0.7, stratify = "Species", seed = 42)
model_clf <- tl_model(split$train, Species ~ ., method = "forest")

plot(model_clf, type = "confusion")
```

### PCA

```{r plot-pca}
pca <- tidy_pca(USArrests, scale = TRUE)

tidy_pca_screeplot(pca)
tidy_pca_biplot(pca, label_obs = TRUE)
```

### Regularisation

```{r plot-lasso}
model_lasso <- tl_model(mtcars, mpg ~ ., method = "lasso")

tl_plot_regularization_path(model_lasso)
tl_plot_regularization_cv(model_lasso)
```

## Tables

`tl_table()` mirrors the plot interface, dispatching on model type and an
optional `type`:

```{r table-auto, eval = FALSE}
tl_table(model)                       # auto-selects the best table type
tl_table(model, type = "coefficients") # specific type
```

### Evaluation Metrics

```{r table-metrics}
tl_table_metrics(model_reg)
```

### Coefficients

For linear and logistic models, the table includes standard errors, test
statistics, and p-values, with significant terms highlighted:

```{r table-coef}
tl_table_coefficients(model_reg)
```

`conf_int = TRUE` adds a confidence interval, and `level` sets its width:

```{r table-coef-ci}
tl_table_coefficients(model_reg, conf_int = TRUE, level = 0.9)
```

These are Wald intervals, built from the standard errors in the column
beside them, so the interval and the p-value in a row always agree about
whether zero is excluded. For a logistic model, `exponentiate = TRUE`
reports odds ratios instead of log odds.

For regularised models, coefficients are sorted by magnitude and zero
coefficients are greyed out. There is no interval to add — glmnet reports
no standard errors, and `conf_int = TRUE` is an error here rather than a
column of `NA`:

```{r table-coef-lasso}
tl_table_coefficients(model_lasso)
```

The numbers behind these tables come from `tl_coefficients()`, which takes
the same arguments, returns a tibble, and does not need `gt` installed:

```{r coef-tibble}
tl_coefficients(model_reg, conf_int = TRUE)
```

### Confusion Matrix

A formatted confusion matrix with correct predictions highlighted on the
diagonal:

```{r table-confusion}
tl_table_confusion(model_clf, new_data = split$test)
```

### Feature Importance

A ranked importance table with a colour gradient:

```{r table-importance}
tl_table_importance(model_clf)
```

### PCA Variance Explained

Cumulative variance is coloured green to highlight how many components are
needed:

```{r table-variance}
pca_model <- tl_model(USArrests, method = "pca")
tl_table_variance(pca_model)
```

### PCA Loadings

A diverging red–blue colour scale highlights strong positive and negative
loadings:

```{r table-loadings}
tl_table_loadings(pca_model)
```

### Cluster Summary

Cluster sizes and mean feature values:

```{r table-clusters}
km <- tl_model(iris[, 1:4], method = "kmeans", k = 3)
tl_table_clusters(km)
```

### Model Comparison

Compare multiple models side-by-side:

```{r table-comparison}
m1 <- tl_model(split$train, Species ~ ., method = "svm")
m2 <- tl_model(split$train, Species ~ ., method = "forest")
m3 <- tl_model(split$train, Species ~ ., method = "tree")

tl_table_comparison(
  m1, m2, m3,
  new_data = split$test,
  names = c("SVM", "Random Forest", "Decision Tree")
)
```

## Interactive Reporting with plotly

Most plot functions return a ggplot2 object, and `ggplotly()` takes any of
those without special handling:

```{r plotly, eval = FALSE}
library(plotly)

ggplotly(plot(model_reg, type = "actual_predicted"))
ggplotly(tidy_pca_biplot(pca, label_obs = TRUE))
ggplotly(tl_plot_regularization_path(model_lasso))
```

These do not return a single ggplot2 object:

- `plot_dendrogram()`, which `plot()` uses for an hclust model,
  `tl_plot_tree()` and `tl_plot_nn_architecture()` draw with base graphics.
- `tl_plot_xgboost_tree()` returns a DiagrammeR widget,
  `tl_plot_deep_architecture()` draws the keras model diagram, and
  `visualize_rules(method = "paracoord")` draws with grid.
- `tl_diagnostic_dashboard()` and `plot_cluster_comparison()` return an
  arranged grid of panels. `plot(model, type = "diagnostics")` and
  `create_cluster_dashboard()` return a list of ggplot2 objects, which
  `ggplotly()` takes one at a time.
- `plot()` on a `tl_explore()` result draws its plot and returns the
  exploration result, and `tl_dashboard()` returns a Shiny app.

## Putting It Together

Fit, score, look, drill in — the four calls that make up most reporting
sections:

```{r workflow}
# Fit
model <- tl_model(split$train, Species ~ ., method = "forest")

# Evaluate
tl_table_metrics(model, new_data = split$test)

# Visualise
plot(model, type = "confusion")

# Drill into feature importance
tl_table_importance(model, top_n = 4)
```

Swap `method = "forest"` for `method = "tree"` and the reporting code above
works without modification. `tl_table_importance()` covers the tree-based
and regularised methods only, so for a method such as `"svm"` or `"nn"`,
leave out the last call.
