---
title: "Introduction to ConvergeR"
author: "Paul de Boissier"
date: "`r Sys.Date()`"
output: BiocStyle::html_document
vignette: >
  %\VignetteIndexEntry{Introduction to ConvergeR}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  warning = FALSE,
  message = TRUE
)
```

# Introduction

**ConvergeR** is an analytical toolkit designed to compute, visualize, and statistically evaluate the convergence of single-cell states, based on the methodology established by [Whitfield et al. (2026)](https://doi.org/10.1158/0008-5472.can-25-4403).

In this vignette, we will demonstrate the package workflow using a full dataset of 3,000 Peripheral Blood Mononuclear Cells (PBMC). We will evaluate cells for their balance between two opposing states: Immune Signaling UP versus Immune Signaling DN from MSigDB.

# 1. Environment Setup and Data Preparation

Let's load the required packages and the lightweight subset of the PBMC dataset bundled with the package.

```{r}
library(ConvergeR)
library(Seurat)
library(ggplot2)

# Load the built-in example dataset
seu_path <- system.file("extdata", "pbmc3k_subset.rds", package = "ConvergeR")
seu <- readRDS(seu_path)

# Normalising
seu <- NormalizeData(seu, verbose = FALSE)

# Inject mock clinical metadata to simulate a multi-patient cohort
set.seed(42)
seu$age <- sample(20:70, ncol(seu), replace = TRUE)
seu$tumor_status <- sample(c("Primary", "Metastatic"), ncol(seu), replace = TRUE)
seu$treatment <- sample(c("Treated", "Untreated"), ncol(seu), replace = TRUE)
seu$SS <- paste0("Patient_", sample(1:5, ncol(seu), replace = TRUE))
```

# 2. Importing MSigDB Signatures and Scoring

Instead of manually importing `.gmt` files, we use the `msigdbr` package to directly retrieve the official Hallmark Immune signatures.

```{r}
library(msigdbr)

# S'assurer que l'on travaille sur l'assay RNA principal
DefaultAssay(seu) <- "RNA"

# Récupérer deux signatures immunitaires/inflammatoires bien exprimées dans les PBMC
hallmark_sets <- msigdbr(species = "Homo sapiens", category = "H")

raw_group1 <- hallmark_sets[hallmark_sets$gs_name == "HALLMARK_INTERFERON_ALPHA_RESPONSE", ]$gene_symbol
raw_group2 <- hallmark_sets[hallmark_sets$gs_name == "HALLMARK_INFLAMMATORY_RESPONSE", ]$gene_symbol

# Intersection stricte avec les gènes réellement présents dans pbmc3k_subset
group1_genes <- intersect(raw_group1, rownames(seu))
group2_genes <- intersect(raw_group2, rownames(seu))

# Calcul des scores de modules Seurat
seu <- AddModuleScore(seu, features = list(group1_genes), name = "Score_IFN_")
seu <- AddModuleScore(seu, features = list(group2_genes), name = "Score_Inflam_")

# Renommer proprement les colonnes pour la suite
seu$Score_IFN <- seu$Score_IFN_1
seu$Score_Inflam <- seu$Score_Inflam_1
```

# 3. Computing the Convergence Score

```{r}
seu <- CalculateConvergedScore(
  seurat_obj = seu,
  principal_score = "Score_IFN",
  other_scores = "Score_Inflam",
  principal_name = "Interferon",
  other_name = "Inflammatory",
  output_colname = "Converged_Immune",
  principal_color = "#9b59b6", # Violet pour Interféron
  other_color = "#e67e22"      # Orange pour Inflammatoire
)
```

# 4. Visualization

ConvergeR automatically reads the color metadata generated in the previous step to produce standardized plots.

## Proportion Barplot

Proportion of cells converging to either Immune state across our 10 mock patients.

```{r, fig.width=8, fig.height=5, fig.align='center'}
PlotConvergenceProportion(
  seurat_obj = seu, 
  x_var = "SS", 
  fill_var = "Converged_Immune_Direction", 
  title = "Immune Convergence by Patient"
)
```

## Crossed Proportion Barplot

Faceted by tumor status to identify subgroup trends.

```{r, fig.width=8, fig.height=5, fig.align='center'}
PlotConvergenceCrossedProportion(
  seurat_obj = seu, 
  x_var = "SS",
  fill_var = "Converged_Immune_Direction",
  facet_var = "tumor_status",
  title = "Immune Convergence split by Tumor Status"
)
```

## Density Distribution

Continuous relationship of the convergence direction with patient age.

```{r, fig.width=8, fig.height=5, fig.align='center'}
PlotConvergenceDensity(
  seurat_obj = seu, 
  x_var = "age",
  fill_var = "Converged_Immune_Direction",
  x_label = "Patient Age"
)
```

# 5. Statistical Testing

We evaluate the statistical significance at the patient level (pseudobulk) to prevent pseudoreplication bias caused by single-cell testing.

## Bivariate Test

```{r}
TestConvergenceScore(
  seurat_obj = seu, 
  score_var = "Converged_Immune", 
  test_var = "age",
  level = "patient", 
  patient_id_var = "SS"
)
```

**How to read this result:** The output message displays the statistical test used (e.g., Spearman correlation for continuous variables), the analysis level (patient pseudobulk to avoid single-cell pseudoreplication), the number of observations ($N$ patients), and the p-value. A p-value below 0.05 indicates a statistically significant association between the convergence score and the tested variable.

## Multivariate Linear Model

```{r}
TestMultivariateConvergence(
  seurat_obj = seu, 
  score_var = "Converged_Immune", 
  test_vars = c("age", "tumor_status", "treatment"),
  level = "patient", 
  patient_id_var = "SS"
)
```

**How to read this result:** This model fits a multivariable linear regression on pseudobulk data. It details the global model performance ($R^2$ and global p-value) and breaks down the independent effect (Estimate) and significance (p-value accompanied by conventional significance stars) for each covariate, allowing you to assess the effect of a variable while controlling for potential confounders.

# Session Information

```{r}
sessionInfo()
```
