---
title: "Self-modifying pipelines"
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 4
description: >
  Shows how you can setup pipelines to modify themselves at runtime, which,
  for example, allows for changing pipeline parameters based on intermediate
  results or even dynamically modify the pipeline's own structure during
  a pipeline run.

vignette: >
  %\VignetteIndexEntry{Self-modifying pipelines}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r knitr-setup, include = FALSE}
require(pipeflow)

knitr::opts_chunk$set(
    comment = "#",
    prompt = FALSE,
    tidy = FALSE,
    cache = FALSE,
    collapse = TRUE
)

old <- options(width = 200L)
```

### Internal pipeline structure

{pipeflow} aims to offer a lean and intuitive interface that enables new
users to get started quickly without having to learn a lot of new
functions. At the same time, it was designed to provide easy access
to the underlying data structures to allow advanced users to modify the
pipeline basically in any way they want.

To see this, let's briefly inspect the internal structure of a pipeline object.

```{r}
pip <- pip_new("my-pipeline") |>
    pip_add("init", \(xInit = 0) xInit) |>
    pip_add("f1", \(x = ~init) x + 1) |>
    pip_add("f2", \(x = ~f1) x + 2) |>
    pip_add("f3", \(x = ~f2) x + 3)

str(pip)
```

There is the `name` of the pipeline, which is just a character string,
as well as a `view` entry, which is undefined initially.

The most "interesting" part is the `pipenv`: an environment that holds the
pipeline's actual state. It is shared by the pipeline and all of its views,
which is why operations on a view write through to the underlying pipeline.

```{r}
ls(pip$pipenv)
```

Listing the `pipenv`^[
    By default, `ls()` hides names starting with a dot. The internal
    entries, which become visible via `ls(pip$pipenv, all.names = TRUE)`,
    are not considered in this vignette.
] reveals the `data`entry, which contains the pipeline's step table as a
`data.table` with one row per step:

```{r}
pip$pipenv$data # or pip_data(pip)
```

Many of the columns should be already familar to you and most of them
can be manipulated safely via the `[<-` and `[[<-` operators,
for example:

```{r}
pip[["f2", "step"]] <- "my_pretty_f2" # updates downstream 'depends'

pip
```

```{r}
pip[step %like% "f", "locked"] <- TRUE

pip
```

An exception are the last three columns (`depends`, `unbound`, `nodeId`):
these are usually derived indirectly from the step definition and
therefore protected against direct assignment.

```{r, error = TRUE}
pip[["init", "nodeId"]] <- 99L
```

Of course, you can even skip the `[<-` and `[[<-` operators and work
directly on the `data.table` object, but with an increased risk to
invalidate the internal consistency of the overall pipeline structure,
for example:

```{r}
dat <- pip$pipenv$data
dat[2, "step"] <- "new f1 name" # fails to update downstream 'depends'
dat[3, "nodeId"] <- 99L # breaks link to internal DAG node
assign("data", dat, envir = pip$pipenv)

pip
```

For this reason, direct manipulation is useful during debugging and
you should mostly stick to the provided operators or functions unless
you really know what you are doing.

### Changing the pipeline structure at runtime

After this little excursion let's next see how to safely modify our
pipeline structure at runtime.

```{r}
pip <- pip_new("my-pipeline") |>
    pip_add("init", \(xInit = 0) xInit) |>
    pip_add("f1", \(x = ~init) x + 1) |>
    pip_add("f2", \(x = ~f1) x + 2) |>
    pip_add("f3", \(x = ~f2) x + 3)

(pip_run(pip))
```

This pipeline just adds 1, 2, and 3 to the initial value, respectively.
Let's modify step `f2` that in turn will modify `f3` at
runtime based on the interim result passed into `f2`.

#### Modify steps

```{r}
pip |> pip_replace(
    "f2",
    \(x = ~f1) {
        if (x > 10) {
            .self$replace("f3", \(x = ~f1) x * 3)
            return(x / 2)
        }
        x + 2
    }
)
```

Basically, step `f2` now checks if the input is greater than 10,
and if so, it replaces step `f3` with a new step now referencing `f1`
that multiplies the input passed from `f1` by 3 and returns half of
the input.

To see this, let's try it with an input of 15.

```{r}
pip |>
    pip_set_params(list(xInit = 15)) |>
    pip_run()

pip
```

We see that both the output of the pipeline and the dependencies
of the last step have changed. Let's confirm by inspecting the function
of the last step.

```{r}
pip[["f3", "fun"]]
```

#### Insert and remove steps

Next, we get even more hacky and instead of just replacing,
we will go a bit further to insert and remove steps.
The pipeline definition is as follows:

```{r}
pip <- pip_new("hicky-hacky") |>
    pip_add("init", \(xInit = 0) xInit) |>
    pip_add("f1", \(x = ~init) x + 1) |>
    pip_add(
        "f2",
        \(x = ~f1) {
            if (x > 10) {
                .self |>
                    pip_add("f2a", \(x = ~f1) x + 21, after = "f1") |>
                    pip_add("f2b", \(x = ~f2a) x + 22, after = "f2a") |>
                    pip_replace("f3", \(x = ~f2b) x + 30) |>
                    pip_remove("f2")
            }
            x + 2
        }
    ) |>
    pip_add("f3", \(x = ~f2) x + 3)
```


If the input is greater than 10, we insert two new steps
`f2a` and `f2b` after `f1`, remove `f2`, and replace `f3` with a new
step that adds 30 to the input. Let's first run with the initial
value of 0 to see the original output.

```{r}
pip_run(pip)

pip
```

Next, we set the initial value to 11 to trigger the changes.

```{r}
pip |>
    pip_set_params(list(xInit = 11)) |>
    pip_run()

pip
```

While the structure has changed as expected, some steps were not
yet run. In fact, since originally step `f3`came after `f2`,
and in contrast to what the log is showing, instead of step `f3`,
actually the new step `f2b` was run^[
    See the state of `f2b` in row 4: it is set to `done`.
]
, albeit with x = NULL as input.

So to have the true results, we need to re-init the parameter
and need to re-run the pipeline.

```{r}
pip |>
    pip_set_params(list(xInit = 11)) |>
    pip_run()

pip
```

Now the output of all steps is as expected. If we want to use
{pipeflow} in production, obviously, having to re-run the pipeline
and temporarily showing a wrong log is not ideal. That is, ideally,
the pipeline run would be aborted once away after all changes were
made in `f2` and then re-run right away from the beginning.
Also, this process potentially should be repeated recursively until
the structure does not change anymore.

### Invoke restart during run

Luckily, with some minimal changes, this behaviour can be achieved
with {pipeflow}.
First, for any step where you want to restart the pipeline run, you
need to call `.self$restart()`, so we adapt the `f2` function as
follows:

```{r}
pip <- pip_new("hacky-with-restart") |>
    pip_add("init", \(xInit = 0) xInit) |>
    pip_add("f1", \(x = ~init) x + 1) |>
    pip_add(
        "f2",
        \(x = ~f1) {
            if (x > 10) {
                .self |>
                    pip_add("f2a", \(x = ~f1) x + 21, after = "f1") |>
                    pip_add("f2b", \(x = ~f2a) x + 22, after = "f2a") |>
                    pip_replace("f3", \(x = ~f2b) x + 30) |>
                    pip_remove("f2")
                .self$restart() # <-- restart the run
            }
            x + 2
        }
    ) |>
    pip_add("f3", \(x = ~f2) x + 3)
```

Second, you just run the pipeline as usual.

```{r}
pip |>
    pip_set_params(list(xInit = 11)) |>
    pip_run()
```

As you can see, the run was aborted right after step `f2`
and re-run from the start based on the new structure.
As a result, the log now is fully aligned with the performed
pipeline run.

Looking at the final pipeline overview, we see that the output matches
the expected output of the modified pipeline.

```{r}
pip
```

Of course, this was just a toy example to show some possibilities,
but I have made use of this feature already in various projects and
may present one of them in a more sophisticated example in the future.

Lastly note that since you have full access to the pipeline object,
of course, you can get even more hacky, but be aware that some
additional operations are done under the hood when steps are added
or removed. It is therefore not recommended to "manually" manipulate
the internal data.table object in terms of removing or adding rows,
or changing important columns such as `depends` or `nodeId` as
this immediately would invalidate the internal consistency of the
dependency graph.

On the other hand, changing entries in columns such as
`tags`, `time`, `state` or `output` is generally not critical.
If in doubt, just try and see what works.
