---
title: "Variance Methods: Taylor, Bootstrap, and Jackknife"
output:
  rmarkdown::html_vignette:
    highlight: null
vignette: >
  %\VignetteIndexEntry{Variance Methods: Taylor, Bootstrap, and Jackknife}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  fig.width = 7,
  fig.height = 4,
  out.width = "100%"
)
```

tidycreel exposes three variance estimation methods through the `variance`
argument of `estimate_effort()`, `estimate_catch_rate()`, and related
functions:

| `variance =` | Description | Requires |
|---|---|---|
| `"taylor"` | Taylor linearization (default) | ≥ 2 PSUs per stratum recommended |
| `"bootstrap"` | Bootstrap resampling (500 replicates) | ≥ 2 PSUs per stratum |
| `"jackknife"` | Delete-one jackknife | ≥ 2 PSUs per stratum |

```{r load}
library(tidycreel)
```

---

## 1  Building a design for the examples

All three methods work on the same `creel_design` object. Here we use the
bundled `example_counts` and `example_interviews` datasets, which have
10 sampled weekdays and 4 sampled weekends.

```{r build-design, warning = FALSE}
data("example_counts")
data("example_interviews")

cal <- unique(example_counts[, c("date", "day_type")])
design <- creel_design(cal, date = date, strata = day_type)
design <- add_counts(design, example_counts)
design <- add_interviews(
  design, example_interviews,
  catch = catch_total,
  effort = hours_fished,
  trip_status = trip_status
)
```

---

## 2  Comparing the three variance methods

Pass `variance = "..."` to any estimator. The point estimate is identical
across methods; only the SE and CI change.

```{r compare, warning = FALSE}
eff_taylor <- estimate_effort(design, variance = "taylor")
eff_bootstrap <- estimate_effort(design, variance = "bootstrap")
eff_jackknife <- estimate_effort(design, variance = "jackknife")

# Summarise all three for comparison
rbind(
  as.data.frame(summary(eff_taylor)),
  as.data.frame(summary(eff_bootstrap)),
  as.data.frame(summary(eff_jackknife))
)
```

With 10–14 PSUs per stratum the three methods produce very similar standard
errors. This convergence is expected: bootstrap and jackknife are
asymptotically equivalent to Taylor linearization for simple probability
samples.

---

## 3  When to use each method

### Taylor linearization (`"taylor"`, default)

- Use Taylor linearization when between-day variability dominates within
  strata and the design is a straightforward stratified random sample of days.
  It is fast, uses a closed-form variance formula, and introduces no additional
  randomness. It relies on large-sample approximations, however, so it can be
  less reliable for highly skewed counts or very small strata.

### Bootstrap (`"bootstrap"`)

- Bootstrap is useful for highly skewed count distributions or complex designs,
  and when you want variance estimates that do not depend on a distributional
  assumption. It captures nonlinear variance propagation, but is stochastic:
  set a seed for reproducibility. It is also slower (500 replicates by default)
  and requires at least two PSUs per stratum.

### Jackknife (`"jackknife"`)

- Choose the jackknife when you prefer a deterministic delete-one estimate that
  is slightly more conservative than Taylor for small samples. It is often more
  stable than bootstrap with few PSUs, but still needs at least two PSUs per
  stratum. For unequal inclusion probabilities, use the appropriate jackknife
  weights.

---

## 4  Grouped estimates

All three methods work with the `by` argument.

```{r grouped, warning = FALSE}
eff_grp_taylor <- estimate_effort(design, by = day_type, variance = "taylor")
eff_grp_bootstrap <- estimate_effort(design, by = day_type, variance = "bootstrap")

rbind(
  cbind(method = "taylor", as.data.frame(summary(eff_grp_taylor))),
  cbind(method = "bootstrap", as.data.frame(summary(eff_grp_bootstrap)))
)
```

---

## 5  Single-PSU diagnostic

All three variance methods require at least **2 PSUs (sampled days) per
stratum**. When a stratum has only 1 sampled day, tidycreel raises a
structured error naming the stratum:

```{r single-psu, error = TRUE, warning = FALSE}
# Build a minimal design with 1 sampled day per stratum
dates_1psu <- c(as.Date("2024-06-03"), as.Date("2024-06-08"))
cal_1psu <- data.frame(date = dates_1psu, day_type = c("weekday", "weekend"))
d_1psu <- creel_design(cal_1psu, date = date, strata = day_type)
counts_1psu <- data.frame(
  date     = dates_1psu,
  day_type = c("weekday", "weekend"),
  count    = c(10L, 15L)
)
d_1psu <- add_counts(d_1psu, counts_1psu)

# Taylor linearization raises an error when variance cannot be estimated
estimate_effort(d_1psu, variance = "jackknife")
```

The error message names the problematic stratum (`"weekday"`) and suggests
either increasing the sampling rate or switching to `variance = "taylor"`
which can still produce a point estimate via the lonely-PSU adjustment.

---

## 6  Summary

```{r summary-table, echo = FALSE}
knitr::kable(
  data.frame(
    Method = c("taylor", "bootstrap", "jackknife"),
    Deterministic = c("Yes", "No (stochastic)", "Yes"),
    `Min PSUs` = c("1 (with lonely-PSU adjust)", "2", "2"),
    Speed = c("Fast", "Slow (500 replicates)", "Moderate"),
    `When to prefer` = c(
      "Default; simple designs",
      "Skewed counts; complex designs",
      "Small samples; deterministic alternative to bootstrap"
    ),
    check.names = FALSE
  ),
  caption = "Variance method comparison"
)
```
