---
title: "Statistical Testing"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Statistical Testing}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
    collapse = TRUE,
    comment = "#>",
    eval = FALSE
)
```

The **BoxPlot**, **yPlot**, and **freqPlot** modules include a **Stats** tab that adds pairwise statistical test results as bracket annotations directly on the plotly figure. The same helpers are exported so you can add bracket annotations to any custom plotly figure.

## Using the Stats tab

Enable testing from the Stats tab in any supported module app. The bundled example apps open with it already on; the box plot compares salaries between neighbouring job levels:

```{r stats-app}
library(VizModules)
plotthis_BoxPlotApp()
```

Every Stats control can also be set from `defaults`, including which comparisons to draw, as `"A vs B"` strings (either order matches):

```{r stats-defaults}
plotthis_BoxPlotApp(
    data_list = list("iris" = example_iris),
    defaults = list(
        x.data        = "Species",
        y.data        = "Sepal.Length",
        stats.enabled = TRUE,
        stat.display  = "symbol",
        stat.pairs    = c("setosa vs versicolor", "versicolor vs virginica")
    )
)
```

The Stats tab exposes these controls:

| Control | Default | Description |
|---------|---------|-------------|
| Enable Stats | OFF | Master toggle |
| Test | Wilcoxon | `wilcox.test`, `t.test`, `kruskal.test`, or `anova` |
| P-value Adjustment | holm | Any `p.adjust` method |
| Display | Adjusted P-value | `p.adj`, `p.value`, or `symbol` (`*`/`**`/`***`/`****`) |
| Significance Threshold | 0.05 | Boundary for `*` vs `ns` |
| Hide Non-Significant | ON | Suppress `ns` brackets |
| Paired Test | OFF | Paired Wilcoxon or paired t-test |
| Comparisons | (all pairs) | Restrict to specific pairs (`stat.pairs` in `defaults`) |
| Bracket Style | Capped | `capped` (ticked) or `flat` |
| Bracket Spacing / Text Offset / Bracket Inset | — | Fine layout control (fractions of y-range) |
| Per Facet Panel | ON | Test independently per facet, or across the full dataset |

### Supported tests

- **Pairwise**: Wilcoxon rank-sum, t-test (paired or unpaired) — draw a bracket per pair.
- **Omnibus**: Kruskal-Wallis, ANOVA — draw a single draggable text annotation with the overall p-value.

### Paired tests

When **Paired Test** is enabled, each group must have the same number of observations, sorted so paired samples align row-by-row within each group.

### Adjusted, faceted and rotated axes

Tests run on the values as plotted. In the `yPlot` module a **Y Adjustment** (z-score, log10, ...) is applied before testing, so a t-test on log-scaled data tests the log values, and the brackets and **Y Axis Min/Max** are in those units too. The **Y Adjustment Function** is applied first and the **Y Adjustment** then rescales the result, so log10 with z-score gives z-scored log values. Rank-based tests (Wilcoxon, Kruskal-Wallis) give the same p-values either way for an order-preserving adjustment.

Under a free y facet scale (`"free"` or `"free_y"`) each panel keeps its own range, so the **Y Axis Min/Max** are not applied and each panel's brackets sit just above that panel's data.

Brackets are stacked above the data on the y-axis, so none are drawn while the values run along the x-axis: a rotated `BoxPlot`, or a `yPlot`/`freqPlot` that includes a ridge plot. The tests still run, and their results are in the source data download.

### Downloading statistics

The **Download Summary** button (Source Data, at the bottom of the controls panel) includes a statistics CSV with a metadata header (correction method, threshold, symbol legend) alongside the plot HTML and source data.

## Using the stat helpers in a custom module

The pipeline is fully exported:

1. `compute_pairwise_stats()` — run the tests, return a data frame.
2. `create_stat_annotations()` — convert results to plotly shapes and annotations.
3. `apply_stat_annotations()` — append them to a plotly figure.

```{r stat-helpers-basic}
library(VizModules)
library(ggplot2)
library(plotly)

stats_df <- compute_pairwise_stats(
    df   = example_iris,
    x    = "Species",
    y    = "Sepal.Length",
    test = "wilcox.test",
    p.adjust.method = "holm"
)

# Build the figure with ggplot2 + ggplotly(), matching how the plot modules
# construct their figures. This matters for bracket placement: ggplotly()
# categorical axes are 1-based (the first factor level sits at x = 1), which
# is the convention create_stat_annotations() expects.
p <- ggplot(example_iris, aes(x = Species, y = Sepal.Length)) +
    geom_boxplot()
fig <- ggplotly(p)

stat_result <- create_stat_annotations(
    stats_df = stats_df,
    fig      = fig,
    df       = example_iris,
    x        = "Species",
    y        = "Sepal.Length",
    display  = "symbol"
)

apply_stat_annotations(fig, stat_result)
```

> **Axis coordinates.** `create_stat_annotations()` positions brackets using
> 1-based categorical x-coordinates, matching figures built with `ggplotly()`
> (as every VizModules plot module does). If you build the figure with a raw
> `plot_ly(type = "box")` call instead, plotly uses 0-based category positions
> and the brackets will appear shifted one category to the right — prefer
> `ggplotly()` for a `ggplot` object to stay consistent with the modules.

### Restricting comparisons

By default all pairwise combinations are tested. Pass a list of length-2 character vectors to `pairs`, or convert to/from the UI's `"A vs B"` strings with `generate_pair_strings()` / `parse_pair_strings()`:

```{r pairs}
choices <- generate_pair_strings(example_iris, x = "Species")
pairs   <- parse_pair_strings(choices[1:2])

stats_df <- compute_pairwise_stats(
    df    = example_iris,
    x     = "Species",
    y     = "Sepal.Length",
    test  = "wilcox.test",
    pairs = pairs
)
```

### Faceted plots and nested grouping

Pass `facet.by` with `per.facet = TRUE` to test independently within each panel, or `group.by` to compare group levels within each x category:

```{r faceted}
compute_pairwise_stats(
    df        = example_rnaseq,
    x         = "condition",
    y         = "log2_cpm",
    test      = "wilcox.test",
    facet.by  = "gene",
    per.facet = TRUE
)
```

Pass the same `facet.by`/`group.by` values to `create_stat_annotations()` so brackets land on the correct subplot axes. If the panels have free y scales, also pass `free.y = TRUE`, so each panel's brackets are measured against that panel's data rather than the tallest one's.

### Test and draw on the plotted values

`compute_pairwise_stats()`, `create_stat_annotations()` and `stat_bracket_y_max()` use whatever values they are given. If the plot transforms a column before drawing it (dittoViz's `var.adjustment`/`var.adj.fxn`, say), give them the transformed values, or the brackets land in a different coordinate space from the data. `adjust_column_values()` reproduces dittoViz's transforms:

```{r plotted-values}
plotted <- adjust_column_values(example_iris, y.col = "Sepal.Length", y.adj.fun = "log10")

stats_df <- compute_pairwise_stats(plotted, x = "Species", y = "Sepal.Length.adj")
fig <- dittoViz::yPlot(example_iris, "Sepal.Length", "Species",
    var.adj.fxn = log10, plots = "boxplot", do.hover = TRUE)
stat_result <- create_stat_annotations(stats_df, fig = fig, df = plotted,
    x = "Species", y = "Sepal.Length.adj")
apply_stat_annotations(fig, stat_result)
```

### `compute_pairwise_stats()` return value

| Column | Description |
|--------|-------------|
| `group1`, `group2` | Groups compared (`"all"` for omnibus) |
| `p.value`, `p.adj` | Raw and adjusted p-values |
| `p.signif` | Significance symbol (`ns`, `*`, `**`, `***`, `****`) |
| `test` | Test name |
| `facet_level`, `x_level` | Facet panel / nested x level (NA when not applicable) |
