---
title: "table1() & tb(): Cohort Baseline Characteristics and 2x2 Tables"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{table1() & tb(): Cohort Baseline Characteristics and 2x2 Tables}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  fig.width = 7,
  fig.height = 4.5
)
library(SimtablR)
data(epitabl)
```

Epidemiological investigations routinely begin with two foundational reporting tasks: characterizing study participants across exposure or outcome strata (the so-called "Table 1"), and evaluating focused bivariate associations using 2x2 contingency tables.

In this guide, we follow the admission and triage phase of the **ESTROBE-ACS study**: a prospective cohort of 1,500 adult patients presenting with acute chest pain or suspected acute coronary syndrome (ACS) across eight emergency departments. We demonstrate how `table1()` and `tb()` provide fast, design-aware table generation for epidemiologists and researchers, specially for those transitioning to R from software such as Stata, SAS, or SPSS.

Along this guide, we will frame each aspect of the analysis with a clinical context and research question, and show how you can use SimtablR to generate publication-ready tables and explore your dataset with minimal code.

---

## 1. Multi-Variable Baseline Cohort Description with `table1()`

> **Clinical Context & Research Question:** In adult patients presenting to the emergency department with suspected ACS, how do baseline demographics, cardiovascular comorbidities, and arrival delays differ between patients with confirmed ACS versus non-ACS?

The primary baseline table characterizes the cohort by strata of the clinical reference standard (`adjudicated_acs`), preserving sample denominators and distribution shapes:

```{r table1-baseline}
baseline <- table1(
  epitabl, # Input our dataset (dataframe object)
  c(age, sex, bmi, smoking, hypertension, diabetes, renal_impairment, presentation_hours),  # Choose variables of interest
  by = adjudicated_acs, # Stratify by the adjudicated ACS outcome
  test = TRUE, # Add hypothesis testing (p-values) for group comparisons
  labels = c(  # Rename variables for publication-ready output
    age = "Age (years)",
    bmi = "Body mass index (kg/m²)",
    smoking = "Smoking status",
    hypertension = "Hypertension",
    diabetes = "Diabetes mellitus",
    renal_impairment = "Renal impairment (eGFR < 60)",
    presentation_hours = "Time from symptom onset to ED (hours)"
  )
)
baseline
```

### Key Statistical Behaviors of `table1()`

1. **Automatic Continuous Summary Selection (`summary = "auto"`):**
   Continuous metrics are audited for skewness using complete cases. Symmetric variables (such as `age`) are automatically summarized as **Mean (SD)**. In contrast, right-skewed variables (such as `presentation_hours`, reflecting emergency presentation delays) are automatically summarized as **Median (IQR)**. 
   Analysts can override this behavior across all variables using `summary = "mean"` or `summary = "median"`.

2. **Explicit Denominators & Missingness:**
   In clinical datasets, missing values are rarely missing at random. Variables with missing observations (such as `bmi`) display explicit `N Missing (%)` counts directly in the table, preventing distorted clinical denominators.

3. **Hypothesis Testing (`test = TRUE`):**
   Group comparisons select appropriate statistical tests:
   - Categorical variables: $N-1$ Pearson chi-squared test by default, or Fisher's exact test when expected cell frequencies are $< 1$.
   - Symmetric continuous variables: Two-sample Welch $t$-test or one-way ANOVA.
   - Skewed continuous variables: Wilcoxon rank-sum test or Kruskal-Wallis test.

---

### Adding Standardized Mean Differences (SMD) via Verb Closure

In observational studies and propensity-score evaluations, p-values depend heavily on sample size, whereas Standardized Mean Differences (SMDs) quantify covariate balance independently of $N$. 

Through SimtablR's **verb closure**, grammar verbs can be piped directly into computed `simtab_result` objects without re-specifying the analysis:

```{r table1-smd}
# Append SMDs to the computed baseline table without repeating parameters
baseline_smd <- baseline |> test(smd = TRUE) # We may also specify smd = "all" to compute SMDs for all covariates, including categorical variables.
baseline_smd
```

Covariates with $\text{SMD} < 0.10$ are traditionally considered well-balanced between clinical groups.

---

## 2. Focused Bivariate Analyses with `tb()`

> **Clinical Context & Research Question:** Does pre-existing renal impairment (eGFR $< 60\text{ mL/min}/1.73\text{ m}^2$) associate with an increased risk of confirmed ACS, and what is the appropriate epidemiological effect measure under prospective cohort sampling?

While `table1()` provides a high-level overview across many covariates, `tb()` isolates individual exposures and outcomes to compute contingency tables, percentage distributions, and design-appropriate effect measures.

### 2.1 Study Design Resolution: Relative Risk in Prospective Cohorts

In prospective cohorts, the primary parameter of interest is the **Risk Ratio (RR)**. Specifying `design = "cohort"` instructs SimtablR's design resolver to estimate the Relative Risk with Greenland-Robins or Katz log confidence intervals:

```{r tb-rr}
tab_rr <- tb(
  epitabl,
  renal_impairment,
  adjudicated_acs,
  flags = c("row", "rr"),
  design = "cohort", # By specifying our study design, SimtablR automatically selects the appropriate effect measure (RR) and confidence interval method.
  ref = "No"
)
tab_rr
```

With this quick and easy code, we can find that patients presenting with baseline renal impairment experienced a significantly higher absolute incidence of confirmed ACS compared to those with preserved renal function.

### 2.2 Flags for quick coding

If you are transitioning from Stata (e.g., `tabulate, row col`) or Epi Info, you can control table contents using concise string flags just as you would in those software packages. The following flags are available:

- **Percentage Denominators:**
  - `flags = "row"`: Row percentages (essential when the row represents the exposure).
  - `flags = "col"`: Column percentages (standard for case-control studies).
  - `flags = "cell"`: Cell percentage out of the total sample size ($N = 1,500$).
- **Effect Measures:**
  - `flags = "rr"`: Risk Ratio (Relative Risk).
  - `flags = "or"`: Odds Ratio with Woolf / logit intervals.
  - `flags = "pr"`: Prevalence Ratio (for cross-sectional designs).
- **Inference:**
  - `flags = "p"`: Displays the inferential p-value.
  - `flags = "miss"`: Displays missing value rows and columns.

```{r tb-flags}
# Cell percentages with explicit chi-squared p-value
tb(epitabl, 
   smoking, adjudicated_acs, # Evaluate the association between smoking status and confirmed ACS
   flags = c("cell", "p") # Add cell percentages and a chi-squared p-value for the association
   ) 
```

---

## 3. Stratified Analysis & Mantel-Haenszel Pooling

> Does biological sex confound or modify the association between renal impairment and confirmed ACS, and what is the common pooled Risk Ratio after adjusting for sex?

Stratification allows evaluating potential confounding and effect modification across subgroups. Passing `strat = sex` calculates stratum-specific estimates alongside the Greenland-Robins Mantel-Haenszel pooled estimate and a test of homogeneity:

```{r tb-stratified}
tab_strat <- tb( 
  epitabl,
  renal_impairment,
  adjudicated_acs,
  strat = sex, # Stratify by biological sex to evaluate effect measure modification
  flags = c("row", "rr"),
  design = "cohort" 
)
tab_strat
```

The stratum-specific Risk Ratios for males and females remain consistent, and the test of homogeneity confirms no significant effect measure modification by sex.

---

## 4. Stacking Multiple Bivariate Tables with `rbind()`

> How do crude unadjusted effect measures across multiple clinical risk factors compare side-by-side prior to multivariable regression modeling?

In manuscript preparation, investigators standardly summarize a battery of unadjusted bivariate associations in a single summary table. SimtablR implements an `rbind()` method specifically for `simtab_tb` objects:

```{r tb-stack}
t_smoke <- tb(epitabl, smoking, adjudicated_acs, flags = c("row", "or"), ref = "Never")
t_htn   <- tb(epitabl, hypertension, adjudicated_acs, flags = c("row", "or"), ref = "No")
t_dm    <- tb(epitabl, diabetes, adjudicated_acs, flags = c("row", "or"), ref = "No")
t_ckd   <- tb(epitabl, renal_impairment, adjudicated_acs, flags = c("row", "or"), ref = "No")

# Stack into a unified crude association table
stacked_crude <- rbind(t_smoke, t_htn, t_dm, t_ckd)
stacked_crude
```

All underlying numeric estimates, standard errors, and confidence bounds are preserved in the returned result and can be extracted using `as.data.frame(stacked_crude)` or formatted for publication with `as_gt()` or `export_docx()`.
