---
title: "Introduction to aiDIF"
output: rmarkdown::html_vignette
bibliography: aidif.bib
vignette: >
  %\VignetteIndexEntry{Introduction to aiDIF}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>", fig.width = 7,
                      fig.height = 5)
library(aiDIF)
```

## 1. Overview

Classical differential item functioning (DIF) asks whether an item behaves
differently across groups within a single scoring condition. When responses
are scored by an AI system rather than by human raters, a second question
appears: does moving from human to AI scoring change the group contrast?

aiDIF separates three quantities, using the robust scaling framework of
@halpin2024:

- $\mathrm{DIF}^H_i$, the group contrast for item $i$ under human scoring;
- $\mathrm{DIF}^A_i$, the same contrast under AI scoring;
- $\mathrm{DASB}_i = \mathrm{DIF}^A_i - \mathrm{DIF}^H_i$, the Differential
  AI Scoring Bias, which is the change in that contrast.

An item can carry substantial DIF under both conditions and have a DASB of
zero: the AI reproduced whatever the human raters were doing. Conversely, an
item with no DIF under either condition in isolation can still show a
nonzero DASB. The three quantities answer different questions.

## 2. Who should use aiDIF

aiDIF is for psychometricians, assessment researchers, and quantitative
analysts who already have item calibrations from both scoring conditions and
want an integrated workflow for robust DIF estimation, cross-scoring
comparison, diagnostics, and visualisation.

It assumes comfort with IRT calibration output: slope and intercept
estimates and their asymptotic covariance matrices. aiDIF operates
**after** calibration. It never sees respondent-level item responses, which
keeps it independent of any particular estimation package but means the
quality of its input is your responsibility.

## 3. The example data

`make_aidif_eg()` returns item parameter estimates for six items in two
groups under both scoring conditions, with a known structure:

- **Item 1** carries DIF under human scoring (+0.5 in the focal group).
- **Item 3** carries DASB: AI scoring adds +0.4 to the focal group only.
- Latent impact is 0.5 SD.
- A uniform +0.1 AI calibration offset affects all items in both groups.

Item 1 is the interesting contrast. It has real DIF, but the AI reproduces
it exactly, so its DASB should be near zero.

```{r data}
eg <- make_aidif_eg()
str(eg, max.level = 2)
```

## 4. Calibration and identification

This section matters more than its length suggests.

DASB is a difference of differences across four calibrations: two groups
times two scoring conditions. If each of those calibrations carries its own
arbitrary metric origin $c_{gs}$, then the constants contribute

$$(c_{FA} - c_{FH} - c_{RA} + c_{RH})$$

to **every** item's DASB. Raw DASB is therefore identified only up to an
item-common additive offset, and that offset is perfectly confounded with a
uniform DASB affecting all items equally.

There are two ways to proceed.

**Common metric.** If linking was established upstream, so that the four
calibrations are demonstrably on the same metric, the offset is zero and raw
DASB is interpretable as it stands. This is the default.

```{r metric-common}
fit_common <- fit_aidif(eg$human, eg$ai, metric = "common")
```

**Linked.** Otherwise, estimate the offset from the data. aiDIF applies the
same bi-square M-estimator it already uses for within-condition robust
scaling to the DASB vector, and reports DASB relative to the estimated
offset. This assumes most items have no scoring-condition-by-group effect —
the same majority assumption that underlies robust DIF scaling generally.

```{r metric-linked}
fit_linked <- fit_aidif(eg$human, eg$ai, metric = "linked")
attr(fit_linked$scoring_bias, "offset")
```

Do not skip this choice. Subtracting intercepts from separately normalised
latent scales without a defensible linking strategy produces numbers that
look fine and mean nothing.

## 5. Independent and paired scoring designs

The variance of DASB depends on how the two conditions were produced.

**Independent** (the default) sums the four marginal variances. It is
correct when the human and AI calibrations come from disjoint respondent
samples.

**Paired** — the same responses scored twice — requires the cross-condition
covariance within each group. Groups remain independent of each other, since
they consist of disjoint respondents:

$$\mathrm{Var}(\mathrm{DASB}_i) = \sum_{g,s}\mathrm{Var}(d_{igs})
  - 2\mathrm{Cov}(d_{iFH}, d_{iFA}) - 2\mathrm{Cov}(d_{iRH}, d_{iRA}).$$

That covariance is normally positive, so the independence variance is an
**upper bound** for paired data. Using `design = "independent"` on paired
data gives conservative p-values, not anti-conservative ones. If you do not
have the cross-condition covariance, the default is the safe choice; say so
in your write-up rather than pretending the design was independent.

```{r paired, eval = FALSE}
# With cross-condition covariance available:
fit_paired <- fit_aidif(eg$human, eg$ai,
                        design    = "paired",
                        cross_cov = my_cross_cov)
```

## 6. Fitting

```{r fit}
mod <- fit_aidif(eg$human, eg$ai, metric = "linked", adjust = "BH")
print(mod)
```

## 7. Conventional DIF in each condition

```{r dif}
round(mod$dif_human, 4)
round(mod$dif_ai, 4)
```

## 8. DASB, intervals, and multiplicity

Six items means six tests. `adjust` applies a multiplicity correction across
the items in this analysis; the family is the item set, not the union of the
human-DIF, AI-DIF, and DASB tests.

```{r dasb}
mod$scoring_bias
```

Item 3 carries the planted scoring-condition-by-group effect, so its DASB
should depart from zero. Item 1 shows that conventional DIF can persist
across scoring conditions without the AI having introduced any additional
differential functioning.

## 9. Describing the pattern of change

```{r effect}
ai_effect_summary(mod)
```

These labels are descriptive. They compare significance flags between the
two conditions, and a difference in significance is not itself a significant
difference [@gelman2006]. The DASB table in the previous section is the test
of whether the contrast changed; this table describes the pattern once you
have that answer.

## 10. Anchor diagnostics

Items downweighted to near zero were effectively excluded from the robust
scale estimate, which flags them as likely DIF-contaminated.

```{r anchors}
anchor_weights(mod)
```

## 11. Visualisation

```{r plot-forest, fig.alt = "Forest plot of item-level DIF estimates with confidence intervals under human and AI scoring."}
plot(mod, type = "dif_forest")
```

```{r plot-dasb, fig.alt = "Bar chart of DASB estimates per item with confidence intervals."}
plot(mod, type = "dasb")
```

```{r plot-weights, fig.alt = "Dot plot of bi-square anchor weights per item in each scoring condition."}
plot(mod, type = "weights")
```

```{r plot-rho, fig.alt = "Bi-square objective function against candidate location values, with the robust estimate marked."}
plot(mod, type = "rho")
```

## 12. Full report

```{r summary}
summary(mod)
```

`summary()` returns an object, so components can be extracted rather than
scraped from printed output:

```{r summary-object}
s <- summary(mod, adjust = "holm")
s$dasb$p_adj
```

## 13. Using aiDIF after calibration with mirt

`as_aidif()` converts a fitted multiple-group model directly into the
structure aiDIF expects, replacing the manual parameter and covariance
extraction earlier versions required. mirt is in `Suggests`, so this section
runs only when it is installed.

```{r mirt, eval = requireNamespace("mirt", quietly = TRUE)}
set.seed(2026)
n <- 600
J <- 8

a <- runif(J, 0.8, 1.6)
d <- rnorm(J, 0, 0.6)

sim_group <- function(n, a, d, mu = 0, d_shift = rep(0, length(d))) {
  th <- rnorm(n, mu, 1)
  p  <- t(vapply(th, function(t) plogis(a * t + d + d_shift), numeric(length(d))))
  matrix(rbinom(length(p), 1, p), nrow = n)
}

# Human scoring: item 1 has DIF in the focal group.
dif_h <- c(0.5, rep(0, J - 1))
# AI scoring adds a uniform +0.1 drift, plus DASB at item 3.
drift <- rep(0.1, J)
dasb  <- c(0, 0, 0.4, rep(0, J - 5), 0, 0)

human <- rbind(sim_group(n, a, d),
               sim_group(n, a, d + dif_h, mu = 0.5))
ai    <- rbind(sim_group(n, a, d + drift),
               sim_group(n, a, d + dif_h + drift + dasb, mu = 0.5))
grp   <- rep(c("reference", "focal"), each = n)

colnames(human) <- colnames(ai) <- paste0("item", seq_len(J))

# No invariance constraints: each group is calibrated on its own metric with
# the latent mean fixed at 0 and variance at 1. That is exactly the input
# robust scaling expects -- it estimates the linking constant itself, so
# constraining anchors here would pre-empt the method.
human_mod <- mirt::multipleGroup(as.data.frame(human), 1, group = grp,
                                 itemtype = "2PL", SE = TRUE, verbose = FALSE)
ai_mod    <- mirt::multipleGroup(as.data.frame(ai), 1, group = grp,
                                 itemtype = "2PL", SE = TRUE, verbose = FALSE)

human_in <- as_aidif(human_mod, groups = c("reference", "focal"))
ai_in    <- as_aidif(ai_mod,    groups = c("reference", "focal"))

fit_mirt <- fit_aidif(human_in, ai_in, metric = "linked", adjust = "BH")
print(fit_mirt)
fit_mirt$scoring_bias
```

## 14. Simulating item-parameter inputs

`simulate_aidif_data()` does **not** simulate respondent-level item
responses and refit IRT models. It directly generates item-parameter
estimates and their asymptotic covariance matrices, consistent with a 2PL
fitted to `n_obs` observations per group. It is intended for package
demonstration and method benchmarking, not as a substitute for a full
response-level simulation study. For the latter, use the mirt workflow in
the previous section.

```{r sim}
dat <- simulate_aidif_data(n_items = 8, n_obs = 600,
                           dif_items = c(1, 2), dif_mag = 0.5,
                           dasb_items = 5, dasb_mag = 0.4, seed = 123)
sim_mod <- fit_aidif(dat$human, dat$ai, metric = "linked", adjust = "BH")
print(sim_mod)
```

## 15. Limitations

- aiDIF works from calibrated item parameters. It cannot detect problems
  introduced during calibration itself.
- The linked metric assumes a majority of items are free of
  scoring-condition-by-group effects. If most items are affected, the
  estimated offset absorbs the signal and DASB estimates shrink toward zero.
- Variance under `metric = "linked"` is a first-order approximation that
  accounts for each item's leverage on the estimated offset but treats the
  bi-square weights as fixed.
- `design = "paired"` uses only the diagonal of the supplied cross-condition
  covariance.
- Only unidimensional models with a single slope per item are supported.
- The AI-effect classification is descriptive, not inferential.

## References

::: {#refs}
:::
