---
title: "Cubist models"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Cubist models}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)
source("_threads.R")
library(tidypredict)
library(Cubist)
library(dplyr)
```


| Function                                                      |Works|
|---------------------------------------------------------------|-----|
|`tidypredict_fit()`, `tidypredict_sql()`, `parse_model()`      |  ✔  |
|`tidypredict_to_column()`                                      |  ✔  |
|`tidypredict_test()`                                           |  ✗  |
|`tidypredict_interval()`, `tidypredict_sql_interval()`         |  ✗  |
|`parsnip`                                                      |  ✗  |


## `tidypredict_` functions

```{r}
library(Cubist)
data("BostonHousing", package = "mlbench")

model <- Cubist::cubist(
  x = BostonHousing[, -14],
  y = BostonHousing$medv,
  committees = 3
)
```

- Create the R formula
    ```{r}
tidypredict_fit(model)
    ```

- SQL output example
    ```{r}
tidypredict_sql(model, dbplyr::simulate_odbc())
    ```

- Add the prediction to the original table
    ```{r}
library(dplyr)

BostonHousing %>%
  tidypredict_to_column(model) %>%
  glimpse()
    ```

We are not able to give an exact match of the original predictions [due to a minor bug](https://github.com/topepo/Cubist/issues/62) in Cubist.

## Parse model spec

Here is an example of the model spec:
```{r}
pm <- parse_model(model)
str(pm, 2)
```

```{r}
str(pm$terms[1:2])
```

## Limitations

- `tidypredict_test()` is not supported
- Prediction intervals are not supported
- Cubist uses 32-bit floats internally, which may cause prediction discrepancies at exact split boundaries. See the [float precision](float-precision.html) article for details. The same 32-bit storage puts a *relative* ceiling of roughly `1e-7` on the agreement with `predict()`, so an outcome on a large scale leaves a correspondingly large absolute difference.
- The instance-based correction that `predict()` applies when `neighbors` is greater than zero is not reproduced. It adjusts each prediction using the nearest training rows, which are not part of the fitted model, so no formula can stand in for it. Formulas from `tidypredict_fit()` match `predict(model, newdata)` with its default `neighbors = 0` only.
