---
title: "Get started with pipeflow"
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 4
description: >
  Start here if this is your first time using pipeflow.
vignette: >
  %\VignetteIndexEntry{Get started with pipeflow}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r knitr-setup, include = FALSE}
knitr::opts_chunk$set(
    comment = "#",
    prompt = FALSE,
    tidy = FALSE,
    cache = FALSE,
    collapse = TRUE
)

old <- options(width = 100L)
```

## A simple example to get started
In this example, we'll use base R's airquality dataset.

```{r show-airquality}
head(airquality)
```

Our goal is to create an analysis pipeline that performs the following steps:

* add new data column `Temp.Celsius` containing the temperature in degrees
  Celsius
* fit a linear model to the data
* plot the data and the model fit.

In the following, we'll show how to define and run the pipeline, how to inspect
the output of specific steps, and finally how to re-run the pipeline with
different parameter settings, which is one of the selling points of using
such a pipeline.

### Pipeline building

For easier understanding, we go step by step. First, we create a new pipeline
with the name "my-pipeline" and add a `data` step that provides the input
dataset.

```{r define-pipeline}
library(pipeflow)

pip <- pip_new("my-pip")

pip <- pip_add(
    pip,
    step = "data",
    fun = function(data = airquality) data
)
```

For each step, at minimum we specify the name of the step (here `step = "data`)
and a function (`fun`) that defines what is computed in that step.
Let's take a first look at the pipeline.

```{r show-initial-pipeline}
pip
```

Each step is represented by one row in the table, which currently shows four columns:

* `step` the name of the step
* `params` the names of the function parameters
* `depends` the dependencies of a step (more on that below)
* `state` the state of each step, which initially is always set to `new`

Next, we want to add a step called `prep`, which prepares our original
data by adding the new column `"Temp.Celsius"`.
This means, our `prep` step will have to use the output of the `data` step.
To refer to the output of an earlier pipeline step, we just write its name
preceded with the tilde (~) operator, which in this case means `~data`.

Since `pip_add` works "by reference", we can add the step as follows:

```{r define-data-prep-step}
pip |> pip_add(
    "prep",
    function(df = ~data) {
        df[, "Temp.Celsius"] <- (df[, "Temp"] - 32) * 5 / 9
        df
    }
)
```

To reiterate: `function(df = ~data)` basically means that once this function
gets called `df` will contain whatever was produced last by the `data` step.

So, a second step called `prep` was added and it depends on the `data`
step, which is also marked in the 2nd row in column `depends`.

```{r}
pip
```

Next, we want to add a step called `fit` that fits a linear model to the
data. The function is using the `prep`ared data and defines a
parameter `xVar`, which determines the variable that is used as
predictor in the linear model.

```{r}
pip |> pip_add(
    "fit",
    function(data = ~prep,
             xVar = "Temp.Celsius") {
        lm(paste("Ozone ~", xVar), data = data)
    }
)

pip
```

Lastly, we add a step called `plot`, which plots the data and the
linear model fit. This function references both the
`fit` and `prep` step. As in the previous step, it defines the
`xVar` parameter plus two more plot-specific parameters `xLab` and `title`.

```{r}
pip |> pip_add(
    "plot",
    function(model = ~fit,
             data = ~prep,
             xVar = "Temp.Celsius",
             xLab = "Temperature in degrees Celsius",
             title = "Linear model fit") {
        require(ggplot2, quietly = TRUE)
        coeffs <- coefficients(model)
        ggplot(data) +
            geom_point(aes(.data[[xVar]], .data[["Ozone"]])) +
            geom_abline(intercept = coeffs[1], slope = coeffs[2]) +
            labs(title = title, x = xLab)
    }
)
```

In the `depends` column of row `4:`, we see that the `plot` step depends
on both the `fit` and the `prep` step.

```{r}
pip
```

In addition to the tabular output, {pipeflow} also provides a graphical
representation that is compatible with the `visNetwork` package.
The `pip_graph()` function returns a list of arguments
that can be feed directly to `visNetwork::visNetwork()`.

```{r, eval = FALSE}
library(visNetwork)
do.call(visNetwork, args = pip_graph(pip)) |>
    visHierarchicalLayout(direction = "LR")
```

```{r, echo = FALSE}
library(visNetwork)
do.call(
    visNetwork,
    args = c(pip_graph(pip), list(height = 100, width = 600))
) |>
    visHierarchicalLayout(direction = "LR")
```

Here, the pipeline is visualized as a directed acyclic graph (DAG) where
the nodes represent the steps and the edges represent the dependencies.

### Pipeline integrity

A key feature of {pipeflow} is that the integrity of a pipeline is verified at
definition time. To see this, let's try to add another step that referencing
step that does not exist in the pipeline.

```{r try-add-bad-step, error = TRUE}
pip |> pip_add(
    "another_step",
    function(x = ~upsi) {
        x
    }
)
```

{pipeflow} immediately signals an error and the pipeline remains unchanged.

```{r}
pip
```


### Pipeline run and output

To run the pipeline, we simply call `pip_run()`,
which produces the following output:

```{r run-pipeline}
pip_run(pip)
```

Let's inspect the pipeline again.

```{r pipeline-after-run}
pip
```

```{r, echo = FALSE}
do.call(
    visNetwork,
    args = c(pip_graph(pip), list(height = 100, width = 600))
) |>
    visHierarchicalLayout(direction = "LR")
```

We can see that the `state` of all steps have been changed from `new` to `done`,
which graphically is represented by the color change from blue to green.

In addition, the output was added in a new `out` column^[
    Technically, the `out` column was there all the time but {pipeflow}
    just did not display it while it was still empty.
]. To access a single value of the pipeline, we just select the row (aka step)
and a column of the pipeline table via the `[[` operator.
For example, to inspect the `out`put of the `fit` and `plot` steps, we do:

```{r inspect-lm, message = FALSE}
pip[["fit", "out"]]
```

```{r inspect-plot, message = FALSE, warning = FALSE, fig.alt = "model-plot"}
pip[["plot", "out"]]
```


### Pipeline parameters

Even for a moderately complex analysis consisting of, say, 15 to 20 different
functions, keeping track of all the different analysis parameters can quickly
get out of hand.

As we will see, with {pipeflow} this becomes much easier, since the pipeline
itself keeps track of all parameters and their values. Let's first inspect the
parameters of the above defined pipeline using the `pip_get_params()` function.

```{r inspect-params}
pip_get_params(pip) |> str()
```

It returns a list of all *unbound* parameters (here
``r toString(names(pip_get_params(pip)))``).
By *unbound* we mean that the values of these parameters don't depend
on other steps (i.e. parameters defined with the `~` operator) and
therefore can be adjusted freely.

Also note that each parameter is only listed once, even if it's used in
multiple steps^[For example, the `xVar` parameter is used in both the `fit`
and `plot` step].
To change parameters, we simply call `pip_set_params()`:

```{r set-xVar}
pip |>
    pip_set_params(list(xVar = "Solar.R", xLab = "Solar radiation in Langleys"))

pip_get_params(pip) |> str()
```

{pipeflow} automatically propagates new parameter values to all steps that
use it. In addition, it will recognize which steps are
affected by the parameter change and mark them as `outdated` (see column
`state`).

```{r show-pipeline-with-outdated-step}
pip
```

```{r, echo = FALSE}
library(visNetwork)
do.call(
    visNetwork,
    args = c(pip_graph(pip), list(height = 100, width = 600))
) |>
    visHierarchicalLayout(direction = "LR")
```

We can see that the `fit` and `plot` steps are now in state
`outdated`, because the `xVar` and `xLab` parameters were updated.
To update the results, we just run the pipeline again.

```{r run-pipeline-again}
pip_run(pip)
```

A closer look at the run log shows that the pipeline skipped the first
two steps and ran only the steps that were outdated, which basically
can be thought of caching or mimicking the behavior of `make` in software
development.
That is, {pipeflow} always keeps track of the step states and only re-runs
those where its necessary.
This can be a huge time saver in larger pipelines^[
    Another use case is backend computation in interactive shiny applications,
    where users change parameters dynamically and want quick updates.
].

After the re-run, we can see that the output was updated accordingly now
showing the new x-variable `Solar.R`.

```{r inspect-plot-again, message = FALSE, warning = FALSE, fig.alt = "model-plot"}
pip[["plot", "out"]]
```


Let's visit some more examples of parameter changes and their effects on
the pipeline. To just change the title of the plot, only the `plot` step
needs to be rerun.

```{r set-title}
pip |> pip_set_params(list(title = "Some new title"))
pip
```

```{r inspect-plot-after-title-change, message = FALSE, warning = FALSE, fig.alt = "model-plot"}
pip_run(pip)
pip[["plot", "out"]]
```

Once we change the input data parameter from the `data` step,
since all other steps depend on it, we expect all steps to be rerun.

```{r}
small_airquality <- airquality[1:10, ]
pip |> pip_set_params(list(data = small_airquality))
pip
```

```{r inspect-plot-after-data-change, message = FALSE, warning = FALSE, fig.alt = "model-plot"}
pip_run(pip)
pip[["plot", "out"]]
```

Last but not least let's try to set parameters that don't exist
in the pipeline, which mostly happens due to accidental misspells.

```{r set-unknown-parameters, warning = TRUE}
pip |> pip_set_params(list(titel = "misspelled variable name", foo = "my foo"))
```

As you see, a warning is given to the user hinting at the respective parameter names,
which makes fixing any misspells straight-forward.

Next, let's see how to [modify the pipeline](v02-modify-pipeline.html).

```{r, include = FALSE}
options(old)
```
