---
title: "Candidate Item Sets, Diagnostics, and Independent Evaluation"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Candidate Item Sets, Diagnostics, and Independent Evaluation}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(ItemRest)
```

ItemRest 1.0.0 identifies candidate item sets under loading and numerical
screening criteria. It does not establish an optimal measurement solution.
The search considers combinations of flagged items, reassesses remaining items,
and stops a branch when no removable problems remain or the model cannot be
estimated. Factor count is fixed at the discovery baseline. Identical item sets
are evaluated once. A bounded search reports that it is incomplete.
The original correlation matrix and observation counts are computed once per
dataset. Each retained set uses a submatrix, with its own positive-definiteness
check and any explicitly requested smoothing. Validation and bootstrap samples
compute their own matrices. Full psych fits can be kept with `store_fits=TRUE`;
the default keeps the loadings, factor correlations, reliability, and diagnostics.

Automatic parallel analysis retains the `parallel_method="fa"` default;
`parallel_method="pc"` uses full-matrix component eigenvalues. For ordinal data,
independent item permutations preserve category margins and missing positions,
and each reference dataset uses the same correlation estimator. The suggested
count may differ from a normal Pearson reference. These are reference samples,
not additional recomputations of the real-data matrix in the removal search.

## Simulated discovery and holdout observations

This example has two correlated factors, a weak item, a cross-loading item, and
one explicitly reversed item. It checks implementation; a single simulation
does not establish methodological performance or substantively valid deletion.

```{r simulate}
set.seed(5102026)
n <- 600L
F1 <- rnorm(n)
F2 <- .3 * F1 + sqrt(1 - .3^2) * rnorm(n)
d <- data.frame(
  A1 = F1 + rnorm(n, sd=.5), A2 = F1 + rnorm(n, sd=.5),
  A3 = F1 + rnorm(n, sd=.5), A4 = F1 + rnorm(n, sd=.5),
  B1 = F2 + rnorm(n, sd=.5), B2 = F2 + rnorm(n, sd=.5),
  B3 = F2 + rnorm(n, sd=.5), B4 = -F2 + rnorm(n, sd=.5),
  Weak = rnorm(n), Cross = .6 * F1 + .6 * F2 + rnorm(n, sd=.5)
)
split <- itemrest_split(d, n_factors=2, seed=5102026,
  keys=c(B4=-1), retain_items=c("A1", "B1"),
  item_reasons=c(A1="Domain A anchor", B1="Domain B anchor"),
  max_solutions=500, verbose=FALSE)
result <- split$discovery
print(result)
```

Without a supplied seed, the local execution date is encoded as DDMMYYYY.
For 5 October 2026 this means numeric seed 5102026 and label `05102026`.
Use the recorded numeric seed to reproduce the run on a different date. The
caller RNG state is restored. Settings include thresholds, keys, and the number
of parallel-analysis replications (100 by default), along with R/package versions.

## Baseline and diagnostics

```{r diagnostics}
result$removal_summary[, c("Solution_ID", "Baseline", "Removed_Items",
  "Analysis_Status", "Candidate", "Problem_Items", "Diagnostics")]
result$search
result$content_decisions
```

The no-removal baseline is always retained and appears first in the full table.
Failure or underidentification of one subset does not abort other branches.
Numerical diagnostics distinguish inadmissible variances, nonpositive definite
correlations, explicit smoothing, reported/nonreported convergence, high factor
correlations, and too few qualifying items per factor. Default screening requires
three nonflagged primary items per factor. A pattern loading above one requires
review and is not by itself classified as a Heywood case under oblique rotation.
Only qualifying, admissible solutions without review flags enter the candidate
table. These rules do not establish validity or independent stability.

Listwise deletion (default) occurs once on all baseline items, retaining the same
observations across subsets. Pairwise handling records pair-count ranges and
uses the minimum pair count as conservative EFA sample size. Missing-data bias
requires separate consideration. `pd_action="smooth"` is an explicit opt-in;
smoothed solutions are recorded and require manual review.

No winner is selected and default ordering is discovery order. Optional ordering
by removal count or explained variance is presentation. Explained variance is
mean model communality computed using L Phi L', so oblique correlations are
included. Values from different retained sets have different denominators and
are not evidence that one instrument is superior.

## Factor reliability and scoring direction

```{r reliability}
result$solution_details[["S00001"]]$assessment$factor_reliability
result$removal_summary[, c("Solution_ID", "Cronbachs_Alpha",
  "Standardized_Alpha", "Analysis_Correlation_Alpha", "Omega_Total")]
```

Factor reports use qualifying nonflagged primary assignments; low/cross-loading
items are unassigned for subscale reliability. Global alpha describes all retained
items and is insufficient for a multidimensional scale. Raw alpha uses observed
covariances; standardized alpha uses Pearson correlations. Selected-correlation
alpha follows the analysis correlation matrix, which may combine ordinal and
continuous correlations under qgraph automatic detection. These measures are
explicitly distinguished, as described in the
[psych alpha documentation](https://www.personality-project.org/r/psych/help/alpha.html).

Model omega total is common variance of a standardized unit-weighted sum divided
by its model-implied total variance. It includes all common factors, also within
factor subscales, and is not omega hierarchical or proof of unidimensionality.
Scoring direction is specified using `keys`; ItemRest does not automatically
reverse items. Signed and absolute loading ranges are reported separately.

## Holdout EFA and optional CFA

```{r holdout}
split$validation$validation_summary[, c("Solution_ID", "Candidate",
  "Analysis_Status", "Problem_Items", "N_Obs")]
split$validation$solution_details[["S00001"]]$factor_congruence
```

Holdout EFA refits the same retained sets without selecting more items. The full
discovery/holdout factor-congruence matrix is supplied for matching and inspection.
Disjoint row indices are returned by `itemrest_split`. Alternatively, use
`itemrest_validate(result, independent_data)`. The caller must establish that
observations are independent.

```{r cfa}
if (requireNamespace("lavaan", quietly=TRUE)) {
  cfa <- itemrest_validate(result, d[split$validation_rows, ], method="cfa", seed=5102026)
  cfa$validation_summary
}
```

CFA freezes discovery primary assignments with zero cross-loadings and correlated
factors. It uses MLR for continuous data and WLSMV for specified/detected ordinal
items. With pairwise discovery handling, continuous CFA uses FIML; this is labelled
in the output rather than presented as the same missing-data estimator. Fit
indices do not become automatic validity decisions. See the
[lavaan categorical-data documentation](https://lavaan.ugent.be/tutorial/cat.html).

## Bootstrap the complete search

```{r bootstrap}
boot <- itemrest_bootstrap(result, d[split$discovery_rows, ], n_boot=20, seed=5102026)
boot$item_stability
boot$solution_stability
table(boot$replicates$Status)
```

Every bootstrap replicate reruns the threshold-driven search with the discovery
factor count fixed. Frequencies distinguish retention in any candidate, in all
candidates, and average retention among candidates. Each complete replication
has equal weight; candidate-free replications count as zero for any/all retention.
Failed/incomplete searches are reported and excluded from denominators. Exact-set
frequencies describe stability without selecting a rank-based winner. These
frequencies do not replace independent validation or give probabilities of validity.

The complete executable script ships at `examples/reproducible_workflow.R` in
the installed package. Content-based decisions and independent replication remain
necessary before using any candidate as a measurement instrument.
