Main entry point for MPAQT transcript quantification. Estimates transcript abundances from short-read (required) and long-read (optional) RNA-seq data using an EM algorithm with optional positional bias correction and prior integration.

mpaqt_quant(
  index,
  sr_counts,
  lr_counts = NULL,
  positional_bias = NULL,
  n_bins = 50L,
  prior_model = NULL,
  do_umi_correction = FALSE,
  umi_correction_timing = "post",
  normalize = "tpm",
  max_iter = 100L,
  tolerance = 1e-04,
  prior_start = 25L,
  convergence_start = 25L,
  compute_uncertainty = FALSE,
  verbose = TRUE
)

Arguments

index

An mpaqt_index object or path to index RDS file

sr_counts

An mpaqt_counts_sr object or path to short-read counts

lr_counts

An mpaqt_counts_lr object or path to long-read counts (optional)

positional_bias

Type of positional bias correction: NULL (none), "3p" (3' bias), or "5p" (5' bias). Default is NULL.

n_bins

Number of distance bins for bias correction (default: 50)

prior_model

Prior specification: NULL (none), "shrinkage", "long_read", or a custom list. See Details.

do_umi_correction

Normalize each transcript P matrix so its x values sum to 1 before post-quantification. This is mainly intended for UMI-based single-cell data.

umi_correction_timing

When do_umi_correction = TRUE and positional weights are used, normalize transcript probabilities either before applying positional weights ("pre") or after weighting ("post", default).

normalize

Normalization method: "tpm" (default), "depth", or "none"

max_iter

Maximum number of EM iterations (default: 100)

tolerance

Convergence tolerance for log-likelihood change (default: 1e-4)

prior_start

Iteration to start using prior information (default: 25)

convergence_start

Iteration to start checking convergence (default: 25)

compute_uncertainty

Compute uncertainty estimates (default: FALSE)

verbose

Print progress messages (default: TRUE)

Value

An mpaqt_quant_result object containing:

  • abundances: Estimated transcript abundances (beta)

  • tpm: Transcripts per million values

  • log_likelihood: Final log-likelihood

  • converged: Whether algorithm converged

  • iterations: Number of iterations run

  • fitted_values: Expected EC counts

  • prior_variance: Prior variance (sigma^2)

  • positional_weights: Positional bias weights (if used)

  • lr_coverage_probs: Long-read coverage probabilities (if used)

  • uncertainty: Standard errors (if computed)

Details

The MPAQT algorithm solves the following optimization problem: $$maximize L(\beta) = L_{SR}(\beta) + L_{LR}(\beta) - R(\beta)$$

Where:

  • \(L_{SR}(\beta)\): Short-read Poisson likelihood

  • \(L_{LR}(\beta)\): Long-read Poisson likelihood (if available)

  • \(R(\beta)\): L2 regularization penalty on log scale

Workflow Phases

Quantification proceeds in two phases:

  1. Pre-quantification (optional): Estimates positional bias weights

  2. Post-quantification (required): Runs main EM algorithm

You can run these phases separately using mpaqt_prequant() and mpaqt_postquant() for more control.

See also

Examples

if (FALSE) { # \dontrun{
# Basic quantification with short reads only
result <- mpaqt_quant(
  index = "my_index/mpaqt.index.rds",
  sr_counts = "results/mpaqt.short_read.rds"
)

# With long reads
result <- mpaqt_quant(
  index = idx,
  sr_counts = sr,
  lr_counts = lr
)

# With positional bias correction
result <- mpaqt_quant(
  index = idx,
  sr_counts = sr,
  lr_counts = lr,
  positional_bias = "3p",
  n_bins = 50
)

# With shrinkage prior
result <- mpaqt_quant(
  index = idx,
  sr_counts = sr,
  prior_model = "shrinkage"
)
} # }