---
title: "Automated Machine Learning with tidylearn"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Automated Machine Learning with tidylearn}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  fig.width = 7,
  fig.height = 5,
  # tl_auto_ml() narrates its progress through message(); the narration is
  # useful at the console and noise in a document
  message = FALSE
)

# xgboost and the BLAS both parallelise. On 16 cores with a threaded BLAS
# this page used 14.9 times as much CPU time as elapsed time, and 0.97 with
# this cap; R CMD check times vignettes against a CPU-to-elapsed threshold
# as it does the tests. Setting OMP_NUM_THREADS from here would do nothing,
# because OpenMP and the BLAS read their thread counts when they load,
# before this chunk runs. The runtime API changes the pools themselves, as
# in tests/testthat.R.
if (requireNamespace("RhpcBLASctl", quietly = TRUE)) {
  RhpcBLASctl::blas_set_num_threads(2)
  RhpcBLASctl::omp_set_num_threads(2)
}
```

```{r setup}
library(tidylearn)
library(dplyr)

# Cross-validation folds, forests and the k-means behind the cluster
# features all draw random numbers. Seeding once here makes the page
# reproduce when run from the top, unless a machine is slow enough to close
# one of the time gates described below.
set.seed(42)
```

## Introduction

`tl_auto_ml()` searches across methods rather than within one. It trains
several models, engineers a couple of feature sets, cross-validates what the
budget allows, and ranks the results.

It orchestrates the wrapped packages rather than implementing anything new.
Every model on the leaderboard is an ordinary tidylearn model, so
`predict()`, `plot()` and `$fit` all work on it.

## A First Run

```{r}
result <- tl_auto_ml(
  iris, Species ~ .,
  time_budget = 30,
  cv_folds = 3
)
```

```{r}
result$leaderboard
```

Three columns: which model, what it scored, and how the score was obtained.

```{r}
result$best_model
```

```{r}
result$task
result$metric
round(as.numeric(result$runtime, units = "secs"), 1)
```

Regression is the same call with a numeric response — `task = "auto"` reads
it off the data:

```{r}
result_reg <- tl_auto_ml(
  mtcars, mpg ~ .,
  time_budget = 30,
  cv_folds = 3
)

result_reg$task
```

```{r}
result_reg$leaderboard
```

## How the Search Works

Four phases run in order, each gated on the remaining budget and on the
`use_reduction` and `use_clustering` toggles.

| Phase | What it does | Classification | Regression |
|-------|-------------|----------------|------------|
| 1. Baselines | Standard models on the raw features | tree, logistic\*, forest | tree, linear, forest |
| 2. PCA variants | PCA preprocessing, then the baselines | pca\_tree, pca\_logistic\*, pca\_forest | pca\_tree, pca\_linear, pca\_forest |
| 3. Cluster variants | Cluster assignment added as a feature | clustered\_tree, clustered\_logistic\*, clustered\_forest | clustered\_tree, clustered\_linear, clustered\_forest |
| 4. Advanced | Heavier methods | svm, xgboost | ridge, lasso |

\* Logistic regression is binary only, so it appears only when the response
has exactly two classes. On a multiclass problem such as `iris` the baselines
are tree and forest.

You can see the phases in what actually got trained:

```{r}
names(result$models)
```

```{r, warning = FALSE}
# The same search on a two-class problem picks up the logistic variants
iris_binary <- iris |>
  filter(Species != "setosa") |>
  mutate(Species = droplevels(Species))

binary_result <- tl_auto_ml(iris_binary, Species ~ ., time_budget = 30,
                            cv_folds = 3)
names(binary_result$models)
```

That call emits a run of `glm.fit: algorithm did not converge` warnings,
which this chunk hides only because there are several of them. They are worth
understanding rather than ignoring: *versicolor* and *virginica* overlap only
slightly, and a three-fold split of 100 rows will sometimes produce a fold
where the two are perfectly separable. Logistic regression has no finite
maximum likelihood estimate on separable data, so the coefficients diverge
and `glm()` says so.

The fit on the full data is fine:

```{r}
tl_model(iris_binary, Species ~ ., method = "logistic")$spec$method
```

It is the folds, not the dataset. Seeing this on your own data means the
classes are close to separable at that sample size — a reason to prefer a
regularised method, which is what `ridge` and `lasso` are doing in phase 4 of
the regression search.

### The evaluation column

Each model is fitted on the full training data, then cross-validated if the
budget allows. When it does not, the leaderboard falls back to training-set
metrics.

```{r}
table(result$leaderboard$evaluation)
```

Training metrics are optimistically biased, so a leaderboard mixing `"cv"`
and `"train"` is not ranking like with like. If you see both, raise the
budget or lower `cv_folds`.

## The Time Budget

`time_budget` is in seconds and is checked **between** model fits, not during
them. Once a model starts it runs to completion, because the wrapped packages
execute C-level code that R cannot safely interrupt. Wall-clock time can
therefore overshoot by the duration of whichever model started last.

The budget controls how much gets tried. The relationship on this data and
this machine:

```{r}
budgets <- c(2, 5, 10, 30)

sweep <- lapply(budgets, function(b) {
  t0 <- Sys.time()
  r <- tl_auto_ml(iris, Species ~ ., time_budget = b, cv_folds = 3)
  data.frame(
    budget = b,
    elapsed = round(as.numeric(difftime(Sys.time(), t0, units = "secs")), 1),
    models = nrow(r$leaderboard),
    cv_scored = sum(r$leaderboard$evaluation == "cv"),
    best = r$leaderboard$model[1]
  )
})

do.call(rbind, sweep)
```

iris is 150 rows, so everything here is quick and the budget is barely
touched. What changes across the sweep is which of these gates each budget
opens:

- The forest baseline, its PCA and cluster variants, and the advanced models
  are only attempted when `time_budget >= 30`.
- The PCA and cluster phases start only while more than 5 seconds and more
  than 10% of the budget remain, and stop adding variants once fewer than 2
  seconds or 5% are left. A budget of 5 seconds or less never reaches them.
- The advanced phase starts only while more than 40% of the budget remains.
- A model is cross-validated only while more than 30% of the budget remains
  when it is scored. After that it is scored on the training data, and the
  `evaluation` column says `"train"`.

On data where a single forest fit takes ten seconds, the last three gates
close early and the same budgets give very different results, so run the
sweep on your own data. Cross-validation is the expensive step:
`tl_cv(folds = 5)` fits five models where a plain fit costs one, so
`cv_folds` is the most effective lever.

## Controlling the Search

```{r}
# Baselines only -- no PCA, no cluster features
baseline_only <- tl_auto_ml(
  iris, Species ~ .,
  use_reduction = FALSE,
  use_clustering = FALSE,
  time_budget = 30,
  cv_folds = 3
)

names(baseline_only$models)
```

```{r}
# Keep PCA, drop clustering
no_clustering <- tl_auto_ml(
  iris, Species ~ .,
  use_clustering = FALSE,
  time_budget = 30,
  cv_folds = 3
)

names(no_clustering$models)
```

Naming a metric changes what the leaderboard optimises. Classification
accepts `accuracy`, `precision`, `recall`, `sensitivity`, `specificity`,
`f1`, `auc` and `pr_auc`; regression accepts `rmse`, `mse`, `mae`, `mape`
and `rsq`.

```{r, warning = FALSE}
by_f1 <- tl_auto_ml(iris_binary, Species ~ ., metric = "f1",
                    time_budget = 30, cv_folds = 3)

by_f1$metric
by_f1$leaderboard
```

## Using the Result

```{r}
split <- tl_split(iris, prop = 0.7, stratify = "Species", seed = 123)

automl <- tl_auto_ml(split$train, Species ~ ., time_budget = 30, cv_folds = 3)

test_preds <- predict(automl$best_model, new_data = split$test)
mean(test_preds$.pred == split$test$Species)
```

Any model on the leaderboard can be pulled out by name:

```{r}
available <- names(automl$models)
available
```

```{r}
scores <- vapply(available, function(nm) {
  preds <- predict(automl$models[[nm]], new_data = split$test)
  mean(preds$.pred == split$test$Species)
}, numeric(1))

data.frame(model = available, test_accuracy = round(scores, 3),
           row.names = NULL) |>
  arrange(desc(test_accuracy))
```

Comparing the leaderboard against held-out accuracy is worth doing: the
model that won the search is not always the one that generalises, and with a
dataset this small the difference between the top few is noise.

## AutoML Against a Single Choice

```{r}
manual <- tl_model(split$train, Species ~ ., method = "forest")
manual_acc <- mean(
  predict(manual, new_data = split$test)$.pred == split$test$Species
)

automl_acc <- mean(test_preds$.pred == split$test$Species)

data.frame(
  approach = c("forest, chosen by hand", "AutoML best"),
  test_accuracy = round(c(manual_acc, automl_acc), 3)
)
```

On iris a random forest is already the right answer, so the search does not
improve on it. That is the expected result on an easy, well-understood
dataset — AutoML earns its cost when you do not already know which method
suits the data.

## Preprocessing First

AutoML does not impute. Handle missing values before the call, and apply the
same transformation to any data you later predict on:

```{r}
processed <- tl_prepare_data(
  split$train, Species ~ .,
  scale_method = "standardize",
  remove_correlated = TRUE
)

automl_processed <- tl_auto_ml(processed$data, Species ~ .,
                               time_budget = 30, cv_folds = 3)

automl_processed$leaderboard$model[1]
```

`tl_pipeline()` handles that bookkeeping properly — see
`vignette("tuning-and-pipelines")`.

## When a Model Is Missing or Scores NA

A model that fails to fit, or errors during evaluation, is dropped with a
message explaining why and never reaches the leaderboard. Run without
`message = FALSE` to see which and why.

A model that evaluates but produces no value for the chosen metric appears
with an `NA` score. `tl_auto_ml()` refuses an unrecognised metric name before
fitting anything, so an `NA` means the metric was undefined on the rows
scored: `auc` needs rows of at least two classes, for instance, and is
undefined on a fold that holds one. Cross-validation leaves such folds out of
its average, so a cross-validated score is `NA` only when the metric was
undefined in every fold. If every score is `NA`, `tl_auto_ml()` warns and
returns the first model trained rather than pretending to have ranked them.

```{r}
sum(is.na(result$leaderboard$score))
```

## AutoML or Tuning?

They search different axes and compose in one direction:

- `tl_auto_ml()` searches **across methods** at default hyperparameters.
- `tl_tune_grid()` and `tl_tune_random()` search **within one method**.

Run AutoML first to find out which family suits the data, then tune the
winner:

```{r}
winner <- automl$best_model$spec$method
winner
```

```{r}
tuned <- tl_tune_grid(
  split$train, Species ~ .,
  method = winner,
  param_grid = tl_default_param_grid(winner, size = "small"),
  folds = 3,
  verbose = FALSE
)

attr(tuned, "tuning_results")$best_params
```

```{r}
mean(predict(tuned, new_data = split$test)$.pred == split$test$Species)
```

## Guidance

1. **Start at `time_budget = 10`** to confirm the call is well formed, then
   raise it. Below 30 seconds the forest and the advanced models are never
   tried, so a 10-second run cannot tell you which model is best.
2. **Lower `cv_folds` before lowering `time_budget`.** When fits are slow,
   the 30% gate above leaves the later models with training-set scores. Two
   folds fit two models per evaluation instead of five and still give
   out-of-sample estimates, which leaves more of the budget for
   cross-validating the rest.
3. **Preprocess first.** AutoML does not impute or scale.
4. **Score on held-out data.** The leaderboard ranks; it does not report
   final performance.
5. **Read past the first row.** When the top few scores are within noise of
   each other, pick on interpretability or fit time instead.
6. **Check the `evaluation` column** before comparing scores.

## Where to Go Next

- **Tuning and pipelines** (`vignette("tuning-and-pipelines")`) — search
  within a method, then freeze the recipe.
- **Diagnostics** (`vignette("diagnostics")`) — what to check about the model
  AutoML handed you.
- **Integration workflows** (`vignette("integration-workflows")`) — the PCA
  and clustering steps AutoML applies, driven by hand.
