---
title: "Prompt Template Positional Bias Testing"
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 2
vignette: >
  %\VignetteIndexEntry{Prompt Template Positional Bias Testing}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  fig.align = "center"
)
library(pairwiseLLM)
library(dplyr)
library(readr)
library(tidyr)
library(stringr)
library(knitr)
```

# 1. Motivation

`pairwiseLLM` uses large language models (LLMs) to compare pairs of writing samples and decide which sample is better on a given trait (for example, _Overall Quality_).  

If a **prompt template** systematically nudges the model toward the first or second position, then scores derived from these comparisons may be biased. This vignette documents how we:

- Designed and tested several prompt templates for **positional bias**  
- Quantified both **reverse-order consistency** and **preference for SAMPLE\_1**  
- Recorded descriptive results for several historical provider configurations

The vignette also shows how to:

- Retrieve the tested templates from the package  
- Inspect their full text  
- Access summary statistics from the experiments

For basic function usage, see:

* [Getting Started with pairwiseLLM](https://shmercer.github.io/pairwiseLLM/articles/getting-started.html)

For advanced batch processing workflows, see:

* [Advanced: Submitting and Polling Multiple Batches](https://shmercer.github.io/pairwiseLLM/articles/advanced-batch-workflows.html)

---

# 2. Testing Process Summary

This section describes an archived 2025 experiment bundled with pairwiseLLM
1.3.0. The result artifact was added to the repository on 2025-12-10; the exact
dates on which its provider calls ran were not recorded. It is not a catalog of
currently available models. In particular, `gemini-3-pro-preview` is retired,
and the old Together identifiers below are retained only to identify archived
result rows.

At a high level, the testing pipeline works as follows:

1. **Trait and samples**

   - Choose a trait (here: `"overall_quality"`) and obtain its description with `trait_description()`.
   - Use `example_writing_samples` or your own dataset of writing samples.

2. **Generate forward and reverse pairs**

   - Use `make_pairs()` to generate all unordered pairs.
   - Use `alternate_pair_order()` to build a deterministic "forward" set.
   - Use `sample_reverse_pairs()` with `reverse_pct = 1` to build a fully "reversed" set, where SAMPLE\_1 and SAMPLE\_2 are swapped for all pairs.

3. **Prompt templates**

   - Define multiple templates (e.g., `"test1"`–`"test5"`) and register them in the template registry.
   - Each template is a text file shipped with the package and accessed via `get_prompt_template("testX")`.

4. **Batch calls to LLM providers**

   - For each combination of:
     - Template (`test1`–`test5`)
     - Backend (Anthropic, Gemini, OpenAI, TogetherAI)
     - Historical model recorded in the bundled experiment
     - Thinking configuration (`"no_thinking"` vs `"with_thinking"`, where applicable)
     - Direction (`forward` vs `reverse`)
   
   - Submit the forward and reverse pairs. You can do this using the package's Batch API helpers (for large-scale jobs) or the live API wrapper `submit_llm_pairs()` with `parallel = TRUE` (for faster turnaround on smaller test sets).
   
   - Store responses as CSVs, including the model’s `<BETTER_SAMPLE>` decision and derived `better_id`.

5. **Reverse-order consistency**

   - Within each (template, backend, model, thinking) condition, compare:
     - The model’s decisions for a pair in the forward set
     - The decisions for the same pair in the reverse set (where positions are swapped)
   - Use `compute_reverse_consistency()` to compute:

     - `prop_consistent`: proportion of comparisons where reversing the order yields the same **underlying winner**.

6. **Positional bias statistics**

   - Use `check_positional_bias()` on the reverse-consistency results to quantify:

     - the descriptive proportion of valid forward and reverse outcomes where
       SAMPLE\_1 is chosen as better; and
     - `p_sample1_overall`: the current paired exact test comparing inconsistent
       pairs where position 1 wins both presentations with those where position
       2 wins both presentations.

7. **Summarize and interpret**

   - Aggregate the results across templates and models into a summary table.
   - Look for templates with:
     - High `prop_consistent` (close to 1).
     - `prop_pos1` close to 0.5.
     - Directional preference estimates and their uncertainty, without treating
       a non-significant test as evidence that positional preference is absent.

In the sections below we show how to retrieve the templates, how they are intended to be used, and how to examine the summary statistics for the experiment.

---

# 3. Trait descriptions and custom traits

In the tests, we evaluated samples for overall quality.

```{r}
td <- trait_description("overall_quality")
td
```

In *pairwiseLLM*, every pairwise comparison evaluates writing samples on a
**trait** — a specific dimension of writing quality, such as:

- **Overall Quality**
- **Organization**
- **IRRC**, an overall-writing rubric spanning prompt task, development of
  explanation, organization, and language use

The trait determines *what the model should focus on* when choosing which
sample is better. Each trait has:

- a **short name** (e.g., `"overall_quality"`)
- a **human-readable name** (e.g., `"Overall Quality"`)
- a **textual description** used inside prompts

The function that supplies these definitions is:

```r
trait_description(name, custom_name = NULL, custom_description = NULL)
```
---

## 3.1 Built-in traits

The package includes some predefined traits accessible by name:

```r
trait_description("overall_quality")
trait_description("organization")
trait_description("IRRC")
```

Calling a built-in trait returns a list with:

```r
$name         # human-friendly name
$description  # the textual rubric used in prompts
```

Example:

```r
td <- trait_description("organization")
td$name
td$description
```

This description is inserted into your chosen prompt template wherever
`{TRAIT_DESCRIPTION}` appears.

---

## 3.2 Setting a different built-in trait

To switch evaluations to another trait, simply pass its ID:

```r
td <- trait_description("organization")

prompt <- build_prompt(
  template   = get_prompt_template("test1"),
  trait_name = td$name,
  trait_desc = td$description,
  text1      = sample1,
  text2      = sample2
)
```

This will update all trait-specific wording in the prompt.

---

## 3.3 Creating a custom trait

If your study requires a new writing dimension, you can define your own trait
directly in the call:

```r
td <- trait_description(
  custom_name        = "Clarity",
  custom_description = "Clarity refers to how easily a reader can understand the writer's ideas, wording, and structure."
)

td$name
#> [1] "Clarity"

td$description
#> [1] "Clarity refers to how easily ..."
```

No built-in name needs to be supplied when using custom text:

```r
prompt <- build_prompt(
  template   = get_prompt_template("test2"),
  trait_name = td$name,
  trait_desc = td$description,
  text1      = sample1,
  text2      = sample2
)
```
---

## 3.4 Why traits matter for positional bias testing

Traits determine the **criterion of comparison**, and different traits may
produce different sensitivity patterns in LLM behavior. For example:

- Different rubrics can produce different response patterns
- Rubric wording can be treated as an experimental condition
- Custom traits allow experimentation with alternative rubric wordings

Because positional bias interacts with how the model interprets the trait,
*every trait–template combination* can be evaluated using the same workflow
described earlier in this vignette.

---

# 4. Example data used in tests

The positional-bias experiments in this vignette use the
`example_writing_samples` dataset that ships with the package.

Each row represents a student writing sample and includes:

- an identifying ID,
- a `text` field containing the full written response.

Below we print the 20 writing samples included in the file.  
This dataset provides a reproducible testing base; in real applications,
you would use your own writing samples.

```{r}
data("example_writing_samples", package = "pairwiseLLM")

# Inspect the structure
glimpse(example_writing_samples)

# Print the 20 samples (full text)
example_writing_samples |>
  kable(
    caption = "20 example writing samples included with pairwiseLLM."
  )
```

---

# 5. Built-in prompt templates

The tested templates are stored as plain-text files in the package and exposed via the template registry. You can retrieve them with `get_prompt_template()`:

```{r}
template_ids <- paste0("test", 1:5)
template_ids
```

Use `get_prompt_template()` to view the text:

```{r}
cat(substr(get_prompt_template("test1"), 1, 500), "...\n")
```

The same pattern works for all templates:

```{r, eval = FALSE}
# Retrieve another template
tmpl_test3 <- get_prompt_template("test3")

# Use it to build a concrete prompt for a single comparison
pairs <- example_writing_samples |>
  make_pairs() |>
  head(1)

prompt_text <- build_prompt(
  template   = tmpl_test3,
  trait_name = td$name,
  trait_desc = td$description,
  text1      = pairs$text1[1],
  text2      = pairs$text2[1]
)

cat(prompt_text)
```

---

# 6. Forward and reverse pairs

Here is a small example of how we constructed forward and reverse datasets for each experiment:

```{r}
pairs_all <- example_writing_samples |>
  make_pairs()

pairs_forward <- pairs_all |>
  alternate_pair_order()

pairs_reverse <- sample_reverse_pairs(
  pairs_forward,
  reverse_pct = 1.0,
  seed        = 2002
)

pairs_forward[1:3, c("ID1", "ID2")]
pairs_reverse[1:3, c("ID1", "ID2")]
```

In `pairs_reverse`, SAMPLE\_1 and SAMPLE\_2 are swapped for every pair relative
to `pairs_forward`. Analyze each model, template, trait, and reasoning condition
separately: `compute_reverse_consistency()` does not group on those columns and
would otherwise pool duplicate votes within each unordered pair.

---

# 7. Thinking / Reasoning Configurations Used in Testing

The archived artifact uses a `thinking` column to distinguish two historical
request configurations:

```
thinking = "no_thinking"   # archived grouping label
thinking = "with_thinking" # archived grouping label
```

These strings are labels in the result artifact, not arguments accepted by the
current public wrappers. The underlying request fields were backend-specific.
Below we describe the historical configurations recorded with the experiment;
consult current provider documentation and the package's dated compatibility
registry before constructing a new request.

---

## 7.1 Anthropic (Claude 4.5 models)

The Anthropic experiment recorded the following controls.

### `thinking = "no_thinking"`
- `reasoning = "none"`
- `temperature = 0`  
- Thinking tokens disabled  
- Configured for lower sampling variability; not a guarantee of deterministic
  behavior

### `thinking = "with_thinking"`
- `reasoning = "enabled"`
- `temperature = 1`
- `include_thoughts = TRUE`
- `thinking_budget = 1024` (max internal reasoning tokens)
- Produces Claude’s full structured reasoning trace (not returned to the user)

This configuration used a larger reasoning budget and a higher temperature.

---

## 7.2 Gemini 3 Pro Preview

The Gemini experiment used the `thinkingLevel` field available to that request
shape at the time.

### Only `thinking = "with_thinking"` was used

Settings used:

- `thinkingLevel = "low"`  
- `includeThoughts = TRUE`  
- `temperature` left at **provider default**  
- Gemini’s structured reasoning is stored internally for bias testing

No cross-provider equivalence of reasoning effort was established.

---

## 7.3 OpenAI (gpt-4.1, gpt-4o, gpt-5.1)

The OpenAI experiment used two API shapes:

1. **`chat.completions`** — standard inference  
2. **`responses`** — reasoning-enabled (formerly “Chain of Thought” via `o-series`)

### `thinking = "no_thinking"`
Used for **all models**, including gpt-5.1:

- Endpoint: `chat.completions`
- `temperature = 0`
- No reasoning traces  
- Lower-temperature configuration; not a guarantee of repeatability

### `thinking = "with_thinking"` (gpt-5.1 only)
- Endpoint: `responses`  
- `reasoning = "low"`  
- `include_thoughts = TRUE`  
- No explicit `temperature` parameter (OpenAI ignores it for this endpoint)

This mode returns reasoning metadata that is stripped prior to analysis.

## 7.4 TogetherAI (Deepseek-R1, Deepseek-V3, Kimi-K2, Qwen3)

For Together.ai, the archived experiment used the Chat Completions API
(`/v1/chat/completions`) with the following historical identifiers:

- "deepseek-ai/DeepSeek-R1"
- "deepseek-ai/DeepSeek-V3"
- "moonshotai/Kimi-K2-Instruct-0905"
- "Qwen/Qwen3-235B-A22B-Instruct-2507-tput" 


DeepSeek-R1 emits internal reasoning wrapped in <think>…</think> tags. 
DeepSeek-V3, Kimi-K2, and Qwen3 do not have a separate reasoning switch; 
any “thinking” they do is part of their standard text output. 

Temperature settings used in testing:
- "deepseek-ai/DeepSeek-R1": `temperature = 0.6`
- DeepSeek-V3, Kimi-K2, Qwen3: `temperature = 0.0`

---

## 7.5 Summary of historical request configurations

| Backend   | Thinking Mode       | What It Controls | Temperature Used | Notes |
|-----------|----------------------|------------------|------------------|-------|
| Anthropic | no_thinking          | reasoning=none, no thoughts | **0** | lower-temperature configuration |
| Anthropic | with_thinking        | reasoning enabled, thoughts included, budget=1024 | **1** | rich internal reasoning |
| Gemini    | with_thinking only   | thinkingLevel="low", includeThoughts | provider default | only configuration represented in the artifact |
| OpenAI    | no_thinking          | chat.completions, no reasoning | **0** | lower-temperature configuration |
| OpenAI    | with_thinking (5.1)  | responses API with reasoning=low | ignored / N/A | only applied to gpt-5.1 |
| Together  | with_thinking        | Chat Completions with `<think>…</think>` extracted to `thoughts` | **0.6** (default)   | internal reasoning always on; visible answer in `content` |
| Together  | no_thinking          | Chat Completions, no explicit reasoning toggle | **0** | reasoning not supported in these specific models |

---

# 8. Loading summary results

The archived results are stored in
`inst/extdata/template_test_summary_all.csv`. Only aggregate rows remain, so the
raw judgments cannot be reanalyzed with the current paired test.

```{r}
summary_path <- system.file("extdata", "template_test_summary_all.csv", package = "pairwiseLLM")
if (!nzchar(summary_path)) stop("Data file not found in installed package.")

summary_tbl <- readr::read_csv(summary_path, show_col_types = FALSE)
head(summary_tbl)
```

---

## 8.1 Column definitions

The columns in `summary_tbl` are:

- **`template_id`**  
  ID of the prompt template (e.g., `"test1"`).

- **`backend`**  
  LLM backend (`"anthropic"`, `"gemini"`, `"openai"`, `"together"`).

- **`model`**  
  Exact historical model identifier recorded by the experiment. These values
  must not be interpreted as a current provider catalog.

- **`thinking`**  
  Reasoning configuration (usually `"no_thinking"` or `"with_thinking"`). The exact meaning depends on the provider and dev script (for example, reasoning turned on vs off, or thinking-level settings for Gemini).

- **`prop_consistent`**  
  Proportion of comparisons that remained consistent when the pair order was reversed. Higher values indicate greater order-invariance.

- **`prop_pos1`**  
  Historical descriptive proportion of the 380 forward and reverse outcomes
  where SAMPLE\_1 was chosen. In current function output, calculate the same
  descriptive quantity as
  `total_pos1_wins / total_comparisons`.

- **`p_sample1_overall`**  
  Historical p-value from the former binomial calculation that treated all
  forward and reverse outcomes as independent. It is retained to describe the
  archived artifact, not as current confirmatory evidence. Current
  `check_positional_bias()` instead applies an exact paired test to informative
  inconsistent pairs.

---

## 8.2 Interpreting the statistics

The three key statistics for each (template, provider, model, thinking) combination are:

1. **Proportion consistent (`prop_consistent`)**

   - Measures how often the underlying winner remains the same when a pair is presented forward vs reversed.
   - Values close to 1 indicate strong order-invariance.
   - It is descriptive reversal agreement, not a reliability coefficient or a
     test of judge validity.

2. **Proportion choosing SAMPLE\_1 (`prop_pos1`)**

   - Measures how often the model selects the first position as better.
   - A value near 0.5 can still be compatible with practically important
     positional preference, especially with limited data.
   - Values substantially above 0.5 suggest a systematic preference for SAMPLE\_1; values substantially below 0.5 suggest a preference for SAMPLE\_2.

3. **Binomial test p-value (`p_sample1_overall`)**

   - In the archived table only, tests a 0.5 position-1 probability while
     treating the two presentations and all pairs as independent.
   - Shared items and paired presentations violate that simple independence
     model, so these archived p-values are descriptive historical outputs.
   - A large p-value must not be interpreted as evidence that bias is absent.

As an example, a row with:

- `prop_consistent = 0.93`  
- `prop_pos1 = 0.48`  
- `p_sample1_overall = 0.57`

was historically summarized as:

- Very high reverse-order consistency.  
- The archived unpaired test did not detect a departure from 0.5; this does not
  establish absence of positional preference.

By contrast, a row with:

- `prop_consistent = 0.83`  
- `prop_pos1 = 0.42`  
- `p_sample1_overall = 0.001`

was historically summarized as:

- Somewhat lower consistency.  
- Evidence under the archived unpaired calculation of more SAMPLE\_2 choices;
  paired raw outcomes would be needed for the current test.

---

# 9. Summary results by prompt

In this section we present, for each template:

1. The full template text (as used in the experiments).  
2. A simple summary table with one row per (backend, model, thinking) configuration and columns:

   - `Backend`  
   - `Model`  
   - `Thinking`  
   - `Prop_Consistent`  
   - `Prop_SAMPLE_1`  
   - `Binomial_Test_p`

---

## 9.1 Template `test1`

### 9.1.1 Template text

```{r}
cat(get_prompt_template("test1"))
```

### 9.1.2 Summary table

```{r}
summary_tbl |>
  filter(template_id == "test1") |>
  arrange(backend, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Backend = backend,
    Model = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

## 9.2 Template `test2`

### 9.2.1 Template text

```{r}
cat(get_prompt_template("test2"))
```

### 9.2.2 Summary table

```{r}
summary_tbl |>
  filter(template_id == "test2") |>
  arrange(backend, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Backend = backend,
    Model = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

## 9.3 Template `test3`

### 9.3.1 Template text

```{r}
cat(get_prompt_template("test3"))
```

### 9.3.2 Summary table

```{r}
summary_tbl |>
  filter(template_id == "test3") |>
  arrange(backend, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Backend = backend,
    Model = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

## 9.4 Template `test4`

### 9.4.1 Template text

```{r}
cat(get_prompt_template("test4"))
```

### 9.4.2 Summary table

```{r}
summary_tbl |>
  filter(template_id == "test4") |>
  arrange(backend, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Backend = backend,
    Model = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

## 9.5 Template `test5`

### 9.5.1 Template text

```{r}
cat(get_prompt_template("test5"))
```

### 9.5.2 Summary table

```{r}
summary_tbl |>
  filter(template_id == "test5") |>
  arrange(backend, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Backend = backend,
    Model = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

# 10. Per-backend summary

It is often useful to examine positional-bias metrics **within each backend**
to see whether:

- certain models exhibit more positional bias than others,
- reasoning mode makes a difference,
- a backend shows overall higher or lower reverse-order consistency.

The tables below show, for each provider, the key statistics:

- **Prop_Consistent** — proportion of consistent decisions under pair reversal  
- **Prop_SAMPLE_1** — proportion of comparisons selecting SAMPLE_1  
- **Binomial_Test_p** — significance level for deviation from 0.5

Each row corresponds to a (template, model, thinking) configuration used in testing.

---

## 10.1 Anthropic models

```{r}
summary_tbl |>
  filter(backend == "anthropic") |>
  arrange(template_id, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Template = template_id,
    Model    = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    caption = "Anthropic: Positional-bias summary by template, model, and thinking configuration.",
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

## 10.2 Gemini models

```{r}
summary_tbl |>
  filter(backend == "gemini") |>
  arrange(template_id, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Template = template_id,
    Model    = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    caption = "Gemini: Positional-bias summary by template, model, and thinking configuration.",
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

## 10.3 OpenAI models

```{r}
summary_tbl |>
  filter(backend == "openai") |>
  arrange(template_id, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Template = template_id,
    Model    = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    caption = "OpenAI: Positional-bias summary by template, model, and thinking configuration.",
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

## 10.4 TogetherAI-hosted models

```{r}
summary_tbl |>
  filter(backend == "together") |>
  arrange(template_id, model, thinking) |>
  mutate(
    Prop_Consistent = round(prop_consistent, 3),
    Prop_SAMPLE_1   = round(prop_pos1, 3),
    Binomial_Test_p = formatC(p_sample1_overall, format = "f", digits = 3)
  ) |>
  select(
    Template = template_id,
    Model    = model,
    Thinking = thinking,
    Prop_Consistent,
    Prop_SAMPLE_1,
    Binomial_Test_p
  ) |>
  kable(
    caption = "TogetherAI: Positional-bias summary by template, model, and thinking configuration.",
    align = c("l", "l", "l", "r", "r", "r")
  )
```

---

# 11. Conclusion

This vignette demonstrates a workflow for describing reversal agreement and
testing directional positional preference in prompt templates.

Use the archived tables to inspect the 2025 experiment, not to infer current
provider availability or certify a template for production. For a new study,
retain raw forward/reverse judgments, analyze each configuration separately,
report agreement and position preference as distinct quantities, and review
effect sizes and study design alongside the paired p-value. Because unordered
pairs commonly share items, statistical review is appropriate when formal
inference is required.

---

# 12. Citation

> Mercer, S. H. (2026). *Prompt template positional bias testing* [R package vignette].
> Comprehensive R Archive Network. https://doi.org/10.32614/CRAN.package.pairwiseLLM
