---
title: "Collect and group output"
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 4
description: >
  How to collect and group pipeline output.
vignette: >
  %\VignetteIndexEntry{Collect and group output}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r knitr-setup, include = FALSE}
knitr::opts_chunk$set(
    comment = "#",
    prompt = FALSE,
    tidy = FALSE,
    cache = FALSE,
    collapse = TRUE
)

old <- options(width = 100L)
```

{pipeflow} manages all functions and parameter dependencies for you, thereby
enabling you to create a lot of pipeline steps without losing track^[
    The linear/sequential structure of a pipeline design also helps!
].
You therefore can (and should) basically follow the principle
*"one step, one task"*, which on the other hand means that most steps will be
helpers and only a few of them contain the final output we are interested in.

This vignette shows how to conveniently collect and group those final outputs.

### Setup

Again, to keep the focus on the displayed functionality, the step functions
are kept very basic.

```{r pipeline-with-output}
library(pipeflow)

pip <- pip_new("my-pip") |>
    pip_add("data", \(x = 1:5) x) |>
    pip_add("prep", \(x = ~data) x * 2, tags = "data") |>
    pip_add("data_summary", \(x = ~prep) range(x),
        tags = c("data", "summary")
    ) |>
    pip_add("model_fit", \(x = ~prep, k = 2) x * k, tags = c("model", "fit")) |>
    pip_add("model_summary", \(x = ~model_fit) sum(x),
        tags = c("model", "summary")
    )
```

As introduced in the [previous vignette](v03b-pipeline-views.html), we use tags to
label steps, specifically, `"data"`/`"model"` to distinguish the topic,
and `"summary"`/`"fit"` for the output type.
Let's briefly run the pipeline and see what's in the `out` column.

```{r}
(pip_run(pip, lgr = NULL))
```

### Flat output collection

To collect output from a pipeline we use `pip_collect()`, which by default returns
all step outputs as a flat named list.

```{r}
pip_collect(pip)
```


### Filtered output using tags

To collect only the output of steps with a specific tag, we filter the
pipeline with `pip_view()` and then call `pip_collect()` on the resulting
view. For example, to collect just the summaries or model output, we filter
by their respective tags:

```{r}
pip_view(pip, tags = "summary") |> # or pip[tags %like% "summary"]
    pip_collect()

pip_view(pip, tags = "model") |>
    pip_collect()
```


### Grouped output

Often the output often different groups will be further combined. Let's
add some more tags to represent section titles of a statistical report.

```{r}
pip[step %in% c("data", "prep")] |> pip_tag("Introduction")
pip[step == "model_fit"] |> pip_tag("Model")
pip[step %like% "summary"] |> pip_tag("Summary")

pip
```

Naturally, we then would group the output as follows:
```{r}
report <- list(
    Introduction = pip_view(pip, tags = "Introduction") |> pip_collect(),
    Model = pip_view(pip, tags = "Model") |> pip_collect(),
    Summary = pip_view(pip, tags = "Summary") |> pip_collect()
)

str(report)
```

As this use case is so common, since version `0.4.0` the `pip_collect`
function natively supports grouping via a `by` parameter,
which allows to simplify the above call as follows:

```{r}
byTags <- pip_collect(pip, by = "tags")
report2 <- byTags[c("Introduction", "Model", "Summary")]

str(report2)
```

You now also can return the collected results as a compact table ...

```{r}
pip_collect(pip, by = "tags", as.table = TRUE)
```

... and of course `by` can be based on other variables, for example,
by all the dependencies.

```{r}
pip_collect(pip, by = "depends", as.table = TRUE)
```

For more details see `?pip_collect`.

```{r, include = FALSE}
options(old)
```
