A complete Gtheory4LLM workflow for the reliability of mean continuous subjective-quality judgments.
All data below are synthetic. This tutorial uses no manuscript results, source-study data, API calls, or real LLM outputs. Its results demonstrate the software workflow; they do not establish an adequate evaluator or prompt count for a real task.
# This vignette is also installed as a runnable script:
# source(system.file("doc", "LLM-workflow.R", package = "Gtheory4LLM"))
library(Gtheory4LLM)The quantity is an item’s mean continuous quality score over evaluators, prompts, and repeated runs. We generate a complete panel of 24 items, four evaluators, three prompts, and two runs: 576 measurements. The model assumes exchangeable random effects for those populations and a Gaussian score model. The scores are simulated continuous numbers, not numeric codes assigned to ordered categories.
| Facet | Meaning in this example |
|---|---|
| Item | The object of measurement. Stable differences between items form the universe-score variance. |
| Evaluator and prompt | Both have main effects and direct item interactions. Different items can respond differently to these facets. |
| Run | Scoped to an evaluator/prompt pair. A run group labels all items. The run source is a common shift across items; remaining item-specific run variation belongs to the residual in this simulation. |
set.seed(20260912)
llm_data <- expand.grid(item = factor(seq_len(24)),
evaluator = factor(paste0("model_", seq_len(4))),
prompt = factor(paste0("prompt_", seq_len(3))),
run = factor(seq_len(2)))
llm_design <- gt_design("item", c("evaluator", "prompt", "run"),
crossed = c("evaluator", "prompt"),
nested = list(run = c("evaluator", "prompt")),
item_interactions = "additive", instrument_interactions = "additive")
llm_design$terms_requested
#> [1] "evaluator" "item" "prompt" "item:evaluator"
#> [5] "item:prompt" "evaluator:prompt:run"This declares six random sources: item, evaluator, prompt, item-by-evaluator, item-by-prompt, and evaluator-by-prompt-by-run. The run group is parent-scoped even though the coded panel is complete. Item-by-run and additional evaluator-by-prompt effects are not silently added. Omitting them is an assumption of this known simulation; a real study needs its own justification. The exact Gaussian engine currently requires a complete coded Cartesian panel with one row per full cell.
source_sd <- c(item = 1.2, evaluator = 0.4, prompt = 0.3,
"item:evaluator" = 0.5, "item:prompt" = 0.4,
"evaluator:prompt:run" = 0.25)
llm_data$quality <- 0
for (source in names(source_sd)) {
members <- strsplit(source, ":", fixed = TRUE)[[1L]]
group <- do.call(interaction, c(llm_data[members], list(drop = TRUE)))
effects <- rnorm(nlevels(group), sd = source_sd[[source]])
llm_data$quality <- llm_data$quality + effects[as.integer(group)]
}
llm_data$quality <- llm_data$quality + rnorm(nrow(llm_data), sd = 0.6)gt_preflight() reports observed counts, source
dimensions, free parameters, configured size limits, and the supported
reliability scale, without building a model matrix or running an
optimizer.
preflight <- gt_preflight(llm_data, "quality", llm_design)
print(preflight)
#> G-theory preflight: 576 rows | exact_balanced_gaussian
#> Observed levels: item=24, evaluator=4, prompt=3, run=2
#> Retained random sources: 6 | Random dimensions: 223 | Model parameters: 8
#> Structural/resource checks: PASS
#> Panel cells (observed coded levels):
#> status cell_count_exact
#> complete 576
#> under_replicated 0
#> over_replicated 0
#> missing 0
#> Audit examples returned: missing_cells=0, replication_issues=0, row_examples=0 | Per-table limit: 10
#> Outcome profiles: 1 | See outcome_profile for category coverage, numeric summaries, and descriptive group variation.
#> Supported reliability scale: observed
#> Reliability of observed Gaussian-score averages.
#> Preflight checks structure and configured size limits. It does not establish identification, approximation adequacy, fit acceptance, or scientific validity.
preflight$sources
#> source observed_groups predictor_dimensions random_dimension covariance_parameters
#> 1 evaluator 4 1 4 1
#> 2 item 24 1 24 1
#> 3 prompt 3 1 3 1
#> 4 item:evaluator 96 1 96 1
#> 5 item:prompt 72 1 72 1
#> 6 evaluator:prompt:run 24 1 24 1
stopifnot(preflight$fitting_feasible)Passing this report means that the listed structural and resource checks pass. It does not prove identification, reliable estimation, numerical acceptance, or approximation adequacy. Read failed checks before changing a resource limit or simplifying a scientific model.
The panel audit counts missing cells and cells above or below the declared replication. It retains only a bounded set of examples, using the observed coded levels; it cannot discover a level that never appeared in the data.
preflight$panel_audit$summary
#> status cell_count cell_count_exact
#> 1 complete 576 576
#> 2 under_replicated 0 0
#> 3 over_replicated 0 0
#> 4 missing 0 0
plot(preflight, type = "cells")incomplete <- gt_preflight(llm_data[-1, ], "quality", llm_design)
incomplete$panel_audit$missing_cells
#> item evaluator prompt run
#> 1 1 model_1 prompt_1 1These diagnostics do not fill cells or remove records. Investigate unexpected codes or missing measurements before deciding how the study should be analysed.
fit <- gt_fit(llm_data, "quality", llm_design,
control = gt_control(gaussian = list(optimizer = "CSOLNP",
extra_tries = 2L, retry_seed = 415L, threads = 1L)))
print(summary(fit))
#> G-theory fit summary: 1 outcome(s), 576 measurement rows
#> Outcomes: quality
#> Families: gaussian | Estimator: REML
#> Random sources: 6 | Instrumentation facets: 3
#> Optimizer: CSOLNP | Attempts: 1 | Selected attempt: 1
#> Optimizer completed: TRUE | Numerically accepted: TRUE
#> Likelihood approximation: exact_balanced_gaussian_likelihood
#> Boundary or nearly singular sources: none recorded
#> -2 log likelihood: 1411
#>
#> Source variances (Gaussian observed-score covariance):
#> source trait variance std_error at_boundary
#> evaluator quality 0.16915 0.15491 FALSE
#> item quality 1.95564 0.60959 FALSE
#> prompt quality 0.06368 0.07587 FALSE
#> item:evaluator quality 0.24121 0.05231 FALSE
#> item:prompt quality 0.10283 0.03178 FALSE
#> evaluator:prompt:run quality 0.04622 0.02085 FALSE
#> Residual quality 0.38950 0.02707 FALSE
#>
#> Fixed location or contrast estimates:
#> quality
#> -0.4553
#>
#> Notes:
#> - Nesting parents define instrumentation groups shared across objects; nested sibling interactions are not automatically included.
#> - No requested source was dropped for this observation family.
#>
#> Use gt_components(fit) for full covariance matrices and gt_diagnostics(fit) for details.
diagnostics <- gt_diagnostics(fit)
print(diagnostics[c("numerically_accepted", "selected_attempt",
"acceptance_failures", "approximation_adequacy")])
#> $numerically_accepted
#> [1] TRUE
#>
#> $selected_attempt
#> [1] "1"
#>
#> $acceptance_failures
#> NULL
#>
#> $approximation_adequacy
#> [1] "exact_balanced_gaussian_likelihood"
stopifnot(isTRUE(diagnostics$numerically_accepted))
gt_components(fit)
#> $evaluator
#> quality
#> quality 0.1691472
#>
#> $item
#> quality
#> quality 1.95564
#>
#> $prompt
#> quality
#> quality 0.06368045
#>
#> $`item:evaluator`
#> quality
#> quality 0.2412094
#>
#> $`item:prompt`
#> quality
#> quality 0.1028346
#>
#> $`evaluator:prompt:run`
#> quality
#> quality 0.04621689
#>
#> $Residual
#> quality
#> quality 0.3895031
#>
#> attr(,"scale")
#> [1] "Gaussian observed-score covariance"
#> attr(,"interpretation")
#> [1] "Cross-category contrasts within one categorical outcome are not separate measured outcomes."The Gaussian default estimator is REML. The trial budget allows the initial fit plus two additional attempts; it is not resampling. The chosen optimizer remains fixed across attempts. Inspect numerical acceptance and source-boundary diagnostics before interpreting the components. A completed optimizer alone is insufficient. If an installation does not support the selected optimizer or a fit is rejected, address that explicitly.
The summary table also carries an asymptotic standard error for each source variance, obtained from the numerically differentiated restricted-likelihood Hessian. A component resting on a variance boundary is flagged, because Wald theory does not apply there.
summary(fit)$variances
#> source trait variance std_error at_boundary
#> 1 evaluator quality 0.16914721 0.15491014 FALSE
#> 2 item quality 1.95564029 0.60958554 FALSE
#> 3 prompt quality 0.06368045 0.07587063 FALSE
#> 4 item:evaluator quality 0.24120936 0.05231458 FALSE
#> 5 item:prompt quality 0.10283459 0.03177643 FALSE
#> 6 evaluator:prompt:run quality 0.04621689 0.02084544 FALSE
#> 7 Residual quality 0.38950314 0.02707234 FALSEreliability <- gt_reliability(fit)
reliability
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2
#> outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se Phi_lower Phi_upper
#> quality 0.9464 0.01778 0.8988 0.9723 0.9173 0.03181 0.8299 0.9619
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
reliability$interpretation
#> [1] "Gaussian observed-score coefficient"Erho2 is the relative coefficient G: common evaluator,
prompt, and run shifts do not change relative item comparisons.
Phi is the absolute coefficient and also includes those
shifts as error. Item-dependent interactions contribute to both error
definitions. Here both coefficients refer to observed Gaussian-score
averages under the declared design. They measure consistency under that
model, not accuracy against human reference labels. A coefficient of
0.80 means that 80% of the variance in item means is universe-score
variance under this model; it is not 80% labeling accuracy.
The interval is a delta-method Wald interval on the logit scale, propagating the fitted parameter covariance matrix through the source covariances. It describes estimation uncertainty in those covariances under the declared model. It is not a guarantee about a future sample of evaluators or prompts.
The dashed reference is an illustrative target. Missing intervals are shown as unavailable, never as a claim of perfect precision. The table records the scale, uncertainty status and any conditional inference, alongside the estimates.
grid <- expand.grid(evaluator = c(2L, 4L, 6L),
prompt = c(2L, 3L, 4L), run = c(1L, 2L))
dstudy <- gt_dstudy(fit, grid)
planning <- as.data.frame(dstudy)
# Prefixed allocation columns avoid collisions with outcome/estimate names.
# These readable aliases are just for this tutorial's known facet names.
planning$evaluator <- planning$allocation_evaluator
planning$prompt <- planning$allocation_prompt
planning$run <- planning$allocation_run
planning$measurements_per_item <- planning$measurements_per_object
planning <- planning[order(planning$measurements_per_item, -planning$Phi), ]
rownames(planning) <- NULL
print(planning)
#> design_id kind outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se
#> 1 1 outcome quality 0.8789244 0.03560952 0.7902498 0.9332760 0.8311242 0.05452953
#> 2 4 outcome quality 0.8989630 0.03088554 0.8204331 0.9454336 0.8543855 0.05049529
#> 3 2 outcome quality 0.9241948 0.02382082 0.8622785 0.9595798 0.8905661 0.03849659
#> 4 7 outcome quality 0.9093288 0.02842758 0.8361284 0.9517192 0.8665114 0.04853116
#> 5 10 outcome quality 0.8985872 0.03138008 0.8185730 0.9456557 0.8508181 0.05218882
#> 6 3 outcome quality 0.9403393 0.01948089 0.8886455 0.9688759 0.9123156 0.03265003
#> 7 5 outcome quality 0.9390021 0.01957600 0.8873654 0.9678246 0.9095813 0.03309628
#> 8 13 outcome quality 0.9125791 0.02788504 0.8403014 0.9539379 0.8681573 0.04884324
#> 9 8 outcome quality 0.9465851 0.01741097 0.9002366 0.9720689 0.9193968 0.03047053
#> 10 11 outcome quality 0.9349508 0.02127879 0.8786388 0.9661408 0.9017489 0.03673160
#> 11 16 outcome quality 0.9197397 0.02612292 0.8513494 0.9582099 0.8770947 0.04728321
#> 12 6 outcome quality 0.9531530 0.01545102 0.9117091 0.9756624 0.9295996 0.02665837
#> 13 9 outcome quality 0.9596917 0.01340455 0.9235008 0.9791477 0.9384895 0.02371661
#> 14 12 outcome quality 0.9477350 0.01771580 0.8999582 0.9733703 0.9201084 0.03137259
#> 15 14 outcome quality 0.9463767 0.01777804 0.8988071 0.9722742 0.9173273 0.03180532
#> 16 17 outcome quality 0.9521950 0.01603064 0.9089930 0.9754426 0.9253201 0.02947603
#> 17 15 outcome quality 0.9582059 0.01419569 0.9196465 0.9786904 0.9349788 0.02569701
#> 18 18 outcome quality 0.9635285 0.01243621 0.9295932 0.9814341 0.9425957 0.02296281
#> Phi_lower Phi_upper allocation_evaluator allocation_prompt allocation_run scale
#> 1 0.6968108 0.9133367 2 2 1 observed
#> 2 0.7259000 0.9285695 2 3 1 observed
#> 3 0.7895704 0.9463806 4 2 1 observed
#> 4 0.7404140 0.9366003 2 4 1 observed
#> 5 0.7181184 0.9273660 2 2 2 observed
#> 6 0.8237973 0.9586001 6 2 1 observed
#> 7 0.8205098 0.9567797 4 3 1 observed
#> 8 0.7404664 0.9382622 2 3 2 observed
#> 9 0.8359359 0.9623144 4 4 1 observed
#> 10 0.8028545 0.9538842 4 2 2 observed
#> 11 0.7512927 0.9440057 2 4 2 observed
#> 12 0.8559650 0.9670397 6 3 1 observed
#> 13 0.8721195 0.9715377 6 4 1 observed
#> 14 0.8330411 0.9637470 6 2 2 observed
#> 15 0.8298542 0.9618948 4 3 2 observed
#> 16 0.8430236 0.9662015 4 4 2 observed
#> 17 0.8626345 0.9705244 6 3 2 observed
#> 18 0.8772614 0.9741760 6 4 2 observed
#> interval_level uncertainty_available uncertainty_reason uncertainty_conditional
#> 1 0.95 TRUE <NA> FALSE
#> 2 0.95 TRUE <NA> FALSE
#> 3 0.95 TRUE <NA> FALSE
#> 4 0.95 TRUE <NA> FALSE
#> 5 0.95 TRUE <NA> FALSE
#> 6 0.95 TRUE <NA> FALSE
#> 7 0.95 TRUE <NA> FALSE
#> 8 0.95 TRUE <NA> FALSE
#> 9 0.95 TRUE <NA> FALSE
#> 10 0.95 TRUE <NA> FALSE
#> 11 0.95 TRUE <NA> FALSE
#> 12 0.95 TRUE <NA> FALSE
#> 13 0.95 TRUE <NA> FALSE
#> 14 0.95 TRUE <NA> FALSE
#> 15 0.95 TRUE <NA> FALSE
#> 16 0.95 TRUE <NA> FALSE
#> 17 0.95 TRUE <NA> FALSE
#> 18 0.95 TRUE <NA> FALSE
#> uncertainty_conditioning Erho2_interval_available Phi_interval_available fixed_facets
#> 1 <NA> TRUE TRUE
#> 2 <NA> TRUE TRUE
#> 3 <NA> TRUE TRUE
#> 4 <NA> TRUE TRUE
#> 5 <NA> TRUE TRUE
#> 6 <NA> TRUE TRUE
#> 7 <NA> TRUE TRUE
#> 8 <NA> TRUE TRUE
#> 9 <NA> TRUE TRUE
#> 10 <NA> TRUE TRUE
#> 11 <NA> TRUE TRUE
#> 12 <NA> TRUE TRUE
#> 13 <NA> TRUE TRUE
#> 14 <NA> TRUE TRUE
#> 15 <NA> TRUE TRUE
#> 16 <NA> TRUE TRUE
#> 17 <NA> TRUE TRUE
#> 18 <NA> TRUE TRUE
#> composite_weights measurements_per_object measurement_count_exact extrapolated batch_status
#> 1 <NA> 4 TRUE FALSE not_declared
#> 2 <NA> 6 TRUE FALSE not_declared
#> 3 <NA> 8 TRUE FALSE not_declared
#> 4 <NA> 8 TRUE TRUE not_declared
#> 5 <NA> 8 TRUE FALSE not_declared
#> 6 <NA> 12 TRUE TRUE not_declared
#> 7 <NA> 12 TRUE FALSE not_declared
#> 8 <NA> 12 TRUE FALSE not_declared
#> 9 <NA> 16 TRUE TRUE not_declared
#> 10 <NA> 16 TRUE FALSE not_declared
#> 11 <NA> 16 TRUE TRUE not_declared
#> 12 <NA> 18 TRUE TRUE not_declared
#> 13 <NA> 24 TRUE TRUE not_declared
#> 14 <NA> 24 TRUE TRUE not_declared
#> 15 <NA> 24 TRUE FALSE not_declared
#> 16 <NA> 32 TRUE TRUE not_declared
#> 17 <NA> 36 TRUE TRUE not_declared
#> 18 <NA> 48 TRUE TRUE not_declared
#> evaluator prompt run measurements_per_item
#> 1 2 2 1 4
#> 2 2 3 1 6
#> 3 4 2 1 8
#> 4 2 4 1 8
#> 5 2 2 2 8
#> 6 6 2 1 12
#> 7 4 3 1 12
#> 8 2 3 2 12
#> 9 4 4 1 16
#> 10 4 2 2 16
#> 11 2 4 2 16
#> 12 6 3 1 18
#> 13 6 4 1 24
#> 14 6 2 2 24
#> 15 4 3 2 24
#> 16 4 4 2 32
#> 17 6 3 2 36
#> 18 6 4 2 48Runs are counted within each evaluator/prompt pair, so the budget per item is evaluators x prompts x runs. A row using more levels than observed extrapolates the fitted variance model to the declared exchangeable population.
screen <- gt_dstudy_target(dstudy, target = 0.80, coefficient = "Phi",
outcome = "quality")
screen[screen$meets_target %in% TRUE, ]
#> design_id kind outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se
#> 1 1 outcome quality 0.8789244 0.03560952 0.7902498 0.9332760 0.8311242 0.05452953
#> 2 2 outcome quality 0.9241948 0.02382082 0.8622785 0.9595798 0.8905661 0.03849659
#> 3 3 outcome quality 0.9403393 0.01948089 0.8886455 0.9688759 0.9123156 0.03265003
#> 4 4 outcome quality 0.8989630 0.03088554 0.8204331 0.9454336 0.8543855 0.05049529
#> 5 5 outcome quality 0.9390021 0.01957600 0.8873654 0.9678246 0.9095813 0.03309628
#> 6 6 outcome quality 0.9531530 0.01545102 0.9117091 0.9756624 0.9295996 0.02665837
#> 7 7 outcome quality 0.9093288 0.02842758 0.8361284 0.9517192 0.8665114 0.04853116
#> 8 8 outcome quality 0.9465851 0.01741097 0.9002366 0.9720689 0.9193968 0.03047053
#> 9 9 outcome quality 0.9596917 0.01340455 0.9235008 0.9791477 0.9384895 0.02371661
#> 10 10 outcome quality 0.8985872 0.03138008 0.8185730 0.9456557 0.8508181 0.05218882
#> 11 11 outcome quality 0.9349508 0.02127879 0.8786388 0.9661408 0.9017489 0.03673160
#> 12 12 outcome quality 0.9477350 0.01771580 0.8999582 0.9733703 0.9201084 0.03137259
#> 13 13 outcome quality 0.9125791 0.02788504 0.8403014 0.9539379 0.8681573 0.04884324
#> 14 14 outcome quality 0.9463767 0.01777804 0.8988071 0.9722742 0.9173273 0.03180532
#> 15 15 outcome quality 0.9582059 0.01419569 0.9196465 0.9786904 0.9349788 0.02569701
#> 16 16 outcome quality 0.9197397 0.02612292 0.8513494 0.9582099 0.8770947 0.04728321
#> 17 17 outcome quality 0.9521950 0.01603064 0.9089930 0.9754426 0.9253201 0.02947603
#> 18 18 outcome quality 0.9635285 0.01243621 0.9295932 0.9814341 0.9425957 0.02296281
#> Phi_lower Phi_upper allocation_evaluator allocation_prompt allocation_run scale
#> 1 0.6968108 0.9133367 2 2 1 observed
#> 2 0.7895704 0.9463806 4 2 1 observed
#> 3 0.8237973 0.9586001 6 2 1 observed
#> 4 0.7259000 0.9285695 2 3 1 observed
#> 5 0.8205098 0.9567797 4 3 1 observed
#> 6 0.8559650 0.9670397 6 3 1 observed
#> 7 0.7404140 0.9366003 2 4 1 observed
#> 8 0.8359359 0.9623144 4 4 1 observed
#> 9 0.8721195 0.9715377 6 4 1 observed
#> 10 0.7181184 0.9273660 2 2 2 observed
#> 11 0.8028545 0.9538842 4 2 2 observed
#> 12 0.8330411 0.9637470 6 2 2 observed
#> 13 0.7404664 0.9382622 2 3 2 observed
#> 14 0.8298542 0.9618948 4 3 2 observed
#> 15 0.8626345 0.9705244 6 3 2 observed
#> 16 0.7512927 0.9440057 2 4 2 observed
#> 17 0.8430236 0.9662015 4 4 2 observed
#> 18 0.8772614 0.9741760 6 4 2 observed
#> interval_level uncertainty_available uncertainty_reason uncertainty_conditional
#> 1 0.95 TRUE <NA> FALSE
#> 2 0.95 TRUE <NA> FALSE
#> 3 0.95 TRUE <NA> FALSE
#> 4 0.95 TRUE <NA> FALSE
#> 5 0.95 TRUE <NA> FALSE
#> 6 0.95 TRUE <NA> FALSE
#> 7 0.95 TRUE <NA> FALSE
#> 8 0.95 TRUE <NA> FALSE
#> 9 0.95 TRUE <NA> FALSE
#> 10 0.95 TRUE <NA> FALSE
#> 11 0.95 TRUE <NA> FALSE
#> 12 0.95 TRUE <NA> FALSE
#> 13 0.95 TRUE <NA> FALSE
#> 14 0.95 TRUE <NA> FALSE
#> 15 0.95 TRUE <NA> FALSE
#> 16 0.95 TRUE <NA> FALSE
#> 17 0.95 TRUE <NA> FALSE
#> 18 0.95 TRUE <NA> FALSE
#> uncertainty_conditioning Erho2_interval_available Phi_interval_available fixed_facets
#> 1 <NA> TRUE TRUE
#> 2 <NA> TRUE TRUE
#> 3 <NA> TRUE TRUE
#> 4 <NA> TRUE TRUE
#> 5 <NA> TRUE TRUE
#> 6 <NA> TRUE TRUE
#> 7 <NA> TRUE TRUE
#> 8 <NA> TRUE TRUE
#> 9 <NA> TRUE TRUE
#> 10 <NA> TRUE TRUE
#> 11 <NA> TRUE TRUE
#> 12 <NA> TRUE TRUE
#> 13 <NA> TRUE TRUE
#> 14 <NA> TRUE TRUE
#> 15 <NA> TRUE TRUE
#> 16 <NA> TRUE TRUE
#> 17 <NA> TRUE TRUE
#> 18 <NA> TRUE TRUE
#> composite_weights measurements_per_object measurement_count_exact extrapolated batch_status
#> 1 <NA> 4 TRUE FALSE not_declared
#> 2 <NA> 8 TRUE FALSE not_declared
#> 3 <NA> 12 TRUE TRUE not_declared
#> 4 <NA> 6 TRUE FALSE not_declared
#> 5 <NA> 12 TRUE FALSE not_declared
#> 6 <NA> 18 TRUE TRUE not_declared
#> 7 <NA> 8 TRUE TRUE not_declared
#> 8 <NA> 16 TRUE TRUE not_declared
#> 9 <NA> 24 TRUE TRUE not_declared
#> 10 <NA> 8 TRUE FALSE not_declared
#> 11 <NA> 16 TRUE FALSE not_declared
#> 12 <NA> 24 TRUE TRUE not_declared
#> 13 <NA> 12 TRUE FALSE not_declared
#> 14 <NA> 24 TRUE FALSE not_declared
#> 15 <NA> 36 TRUE TRUE not_declared
#> 16 <NA> 16 TRUE TRUE not_declared
#> 17 <NA> 32 TRUE TRUE not_declared
#> 18 <NA> 48 TRUE TRUE not_declared
#> coefficient target estimate meets_target fewest_measurements
#> 1 Phi 0.8 0.8311242 TRUE TRUE
#> 2 Phi 0.8 0.8905661 TRUE FALSE
#> 3 Phi 0.8 0.9123156 TRUE FALSE
#> 4 Phi 0.8 0.8543855 TRUE FALSE
#> 5 Phi 0.8 0.9095813 TRUE FALSE
#> 6 Phi 0.8 0.9295996 TRUE FALSE
#> 7 Phi 0.8 0.8665114 TRUE FALSE
#> 8 Phi 0.8 0.9193968 TRUE FALSE
#> 9 Phi 0.8 0.9384895 TRUE FALSE
#> 10 Phi 0.8 0.8508181 TRUE FALSE
#> 11 Phi 0.8 0.9017489 TRUE FALSE
#> 12 Phi 0.8 0.9201084 TRUE FALSE
#> 13 Phi 0.8 0.8681573 TRUE FALSE
#> 14 Phi 0.8 0.9173273 TRUE FALSE
#> 15 Phi 0.8 0.9349788 TRUE FALSE
#> 16 Phi 0.8 0.8770947 TRUE FALSE
#> 17 Phi 0.8 0.9253201 TRUE FALSE
#> 18 Phi 0.8 0.9425957 TRUE FALSE
screen[screen$fewest_measurements %in% TRUE, ]
#> design_id kind outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se
#> 1 1 outcome quality 0.8789244 0.03560952 0.7902498 0.933276 0.8311242 0.05452953
#> Phi_lower Phi_upper allocation_evaluator allocation_prompt allocation_run scale interval_level
#> 1 0.6968108 0.9133367 2 2 1 observed 0.95
#> uncertainty_available uncertainty_reason uncertainty_conditional uncertainty_conditioning
#> 1 TRUE <NA> FALSE <NA>
#> Erho2_interval_available Phi_interval_available fixed_facets composite_weights
#> 1 TRUE TRUE <NA>
#> measurements_per_object measurement_count_exact extrapolated batch_status coefficient target
#> 1 4 TRUE FALSE not_declared Phi 0.8
#> estimate meets_target fewest_measurements
#> 1 0.8311242 TRUE TRUEA threshold of 0.80 is illustrative; choose and justify your own criterion. Read the interval columns alongside the point projection: several allocations here have overlapping intervals, so the ordering of nearby rows is not established by these data.
The screening helper keeps every supplied candidate, including those that miss the target or have an unavailable estimate. It marks all ties for the fewest measurements among qualifying candidates. It performs no search outside this grid and does not guarantee that a future panel will meet the target.
plot(dstudy, coefficient = "Phi", outcome = "quality", target = 0.80,
main = "Candidate measurement allocations")Different allocations can use the same total number of measurements; they are separate points rather than one assumed smooth curve. The figure identifies allocations extending beyond the observed facet counts. Such projections keep the fitted source distributions fixed.
Both result tables contain ordinary scalar columns, suitable for CSV output:
A report collects the design, fitted model context, diagnostic flags, coefficient and decision-study tables, and figures. It does not rerun optimization. Supplied coefficient results are checked against projections from the fit; supplied preflight information is checked for compatible aggregate specifications, which does not prove that it came from identical observations.
analysis_report <- gt_report(fit, preflight = preflight,
reliability = reliability, dstudy = dstudy)
print(analysis_report)
#> G-theory analysis report | schema 1.0
#> Backend: OpenMx exact balanced Gaussian factorial contrasts | Numerically accepted: TRUE
#> Tables: 4 design variables | 1 reliability rows | 18 decision-study rows
#> Export with gt_export_report(report, path).The HTML file contains its figures and styling and can be opened offline. An existing destination is refused unless replacement is explicitly requested. The report omits observations, row examples and group-level identifiers, while keeping variable and source names and aggregate results. This is not an anonymization guarantee. Inspect study-specific names before sharing it. Retained fitting provenance and the report-generation environment are separate; missing fitting provenance is reported as unavailable.
A facet is random when the study generalizes to a population of its levels, and fixed when the universe of generalization is exactly the levels used. LLM evaluation may target exactly a chosen temperature or prompt set, in which case those facets are fixed. Selecting levels deliberately does not by itself decide the universe of generalization.
Declare them with fixed. Following the mixed model of
Brennan (2001), the object-by-fixed-facet variance is averaged over that
facet’s levels and added to universe-score variance, and a source built
only from fixed facets shifts every item equally and leaves the
model.
fixed_prompt <- gt_reliability(fit, fixed = "prompt")
fixed_prompt
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2
#> Fixed facets: prompt | mixed (Brennan 2001): object-by-fixed-facet variance enters the universe score
#> outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se Phi_lower Phi_upper
#> quality 0.963 0.01261 0.9286 0.9811 0.9428 0.02464 0.8707 0.9758
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
fixed_prompt$source_roles
#> evaluator item
#> "absolute_error" "universe"
#> prompt item:evaluator
#> "dropped_fixed_instrumentation_constant" "relative_and_absolute_error"
#> item:prompt evaluator:prompt:run
#> "universe_after_fixed_facet_averaging" "absolute_error"
#> Residual
#> "relative_and_absolute_error"Treating the prompt set as fixed raises both coefficients, because item-by-prompt variation is now part of what the study is trying to measure rather than error. A fixed facet’s count cannot be changed, and a decision study may not project over it.
fixed acts at the reliability stage: it changes how the
already-fitted components are aggregated into a coefficient. It does not
change the G study, so it cannot repair a source whose variance was
misspecified during fitting. Observed seed disagreement can change
across temperatures even when latent variance is constant, because
category probabilities can change. If diagnostics raise doubts about
pooling temperatures, fitting within one temperature is a useful
sensitivity analysis for a conditional estimand; declaring a facet fixed
is a separate decision about generalization.
| Outcome | Current scalar reliability support |
|---|---|
| Gaussian | Observed continuous-score averages. |
| Binary or ordinal | Explicit scale = "latent" after an accepted fit. This
describes latent-response averages, not majority votes, label
proportions, or observed ordinal-score averages. |
| Unordered categorical | No implemented default scalar G/Phi. A scientific score or category-probability estimand is needed. |
Declare category order and references explicitly. Do not recode
binary, ordinal, or nominal labels as the continuous quality score used
above. Discrete fitting uses a bounded dense Laplace engine with
Gaussian latent random effects; gt_preflight() reports its
current limits, and changing the link does not remove the random-effects
distribution assumption. That engine reports point estimates only: it
computes no standard errors. Analytic reliability and D studies require
complete balanced coded panels even when discrete fitting admits missing
whole cells.
At one observation per object-by-facets cell, a discrete fit cannot identify the full-cell source that the default design requests. Declare the same design without it:
A complete panel can still contain little observed outcome variation. This small, deliberately imbalanced example has one positive response. It is a data profile, not an estimation example or a recommended design.
information_data <- expand.grid(item = seq_len(6), rater = seq_len(3))
information_data$label <- as.integer(information_data$item == 1 &
information_data$rater == 1)
information_design <- gt_design("item", "rater", random = ~ item + rater)
information <- gt_preflight(information_data, "label", information_design,
gt_family("binary"), max_examples = 0)
information$outcome_profile$label$categories
#> category count proportion observed
#> 1 0 17 0.94444444 TRUE
#> 2 1 1 0.05555556 TRUE
information$outcome_profile$label$by_variable
#> variable grouping_scope total_groups no_variation_groups single_row_groups
#> 1 item marginal_coded_levels 6 5 0
#> 2 rater marginal_coded_levels 3 2 0
#> no_variation_proportion
#> 1 0.8333333
#> 2 0.6666667
plot(information, type = "outcomes", outcome = "label")The heatmap describes the fraction of groups in which each category appears. A group with one observation necessarily has no observed variation. These summaries supply no minimum-count threshold, independence claim or guarantee of parameter recovery. Declared categories absent from the entire panel are shown with zero counts and block fitting; the preflight report does not merge them or silently change their order. No finite panel can reveal a category never observed and never declared.
Installed help: ?gt_design, ?gt_family,
?gt_preflight, ?gt_control,
?gt_fit, ?gt_diagnostics,
?gt_reliability, ?gt_dstudy, and
?gt_report.
Brennan, R. L. (2001). Generalizability Theory. Springer. doi:10.1007/978-1-4757-3456-0