This article describes the current MPAQT source layout, public R API, S3 objects, and development workflow. Function signatures and object fields below correspond to package version 2.4.0.

Repository Layout

MPAQT/
├── R/                         # Package source
├── tests/testthat/            # Unit and integration tests
├── vignettes/                 # Source articles
├── man/                       # Generated Rd documentation
├── inst/
│   ├── apptainer/             # Apptainer build assets
│   ├── conda/                 # Conda environments and setup scripts
│   └── conda-recipe/          # Conda package recipes
├── DESCRIPTION                # Package metadata and dependencies
├── NAMESPACE                  # Generated exports and imports
└── _pkgdown.yml               # Website navigation and reference groups

The local docs/ directory is ignored build output. The pkgdown workflow builds that directory and publishes it to the gh-pages branch; edit the README and vignette sources instead of generated HTML.

R Source Organization

Files in R/ are grouped by responsibility:

Files Responsibility
api-index.R, api-quant.R Main exported index and quantification entry points
process-short-read.R, process-long-read.R Input-path detection and count preparation
quant-pre.R, quant-post.R Positional-bias prequantification and final quantification
algo-bias.R, algo-em.R, algo-newton.R, algo-prior.R Numerical algorithms and prior fitting
class-index.R, class-counts.R, class-result.R S3 constructors, validation-facing methods, and accessors
infra-kallisto.R, infra-minimap2.R, infra-bambu.R, infra-system.R External-tool wrappers
extract-transcriptome.R, validate.R Reference extraction and shared validation
mpaqt-package.R, zzz.R Package-level documentation and initialization

External commands should remain behind the infra-* wrappers. Primary workflow functions use the mpaqt_* prefix. Public accessors and exported require_*() dependency checks use descriptive names, while internal validators and tool checks use validate_*() and check_*().

Public Workflow

The main object flow is:

mpaqt_index()
    -> mpaqt_index

mpaqt_prepare_short_reads() / mpaqt_prepare_short_reads_sc()
    -> mpaqt_counts_sr or a named list of those objects

mpaqt_prepare_long_reads() / mpaqt_prepare_long_reads_sc()
    -> mpaqt_counts_lr or a named list of those objects

mpaqt_quant() / mpaqt_quant_sc()
    -> mpaqt_quant_result or a named list of result objects

For positional-bias workflows, mpaqt_quant() calls mpaqt_prequant() and then passes its mpaqt_prequant object to mpaqt_postquant(). Users can call the two phases directly when they need to save or reuse estimated weights.

S3 Object Layouts

All primary objects inherit from list; summary information that is not a list component is stored as an attribute.

mpaqt_index

list(
  transcripts = character(),
  genes = character(),
  ec_ids = character(),
  p_matrices = list(),
  ec_counts = numeric(),
  covariates = matrix(),
  distances = data.table::data.table(),
  t2g_matrix = Matrix::sparseMatrix(),
  g2t_matrix = Matrix::sparseMatrix(),
  t2g_normalized = Matrix::sparseMatrix(),
  kallisto_index = NULL,
  gtf_annotation = NULL,
  transcriptome_seqs = NULL,
  kallisto_binary = NULL
)

p_matrices is a list with one data table per transcript. Each table stores the one-based EC row indices in i, design/exposure values in x, and the corresponding dist_5p and dist_3p values. Index attributes include n_transcripts, n_genes, n_ec, created, mpaqt_version, and bundled. Optional components can be NULL.

mpaqt_counts_sr

list(
  counts = numeric(),       # named EC-count vector
  technology = "bulk",     # bulk, 10xv2, 10xv3, or 10xv4
  sample_id = NULL
)

Its classes are mpaqt_counts_sr, mpaqt_counts, and list. Attributes include n_reads, n_ec, and created.

mpaqt_counts_lr

list(
  counts = numeric(),       # named transcript-count vector
  sample_id = NULL
)

Its classes are mpaqt_counts_lr, mpaqt_counts, and list. Attributes include n_reads, n_transcripts, and created.

mpaqt_prequant

list(
  weights = numeric(),
  quantiles = numeric(),
  bias_type = "3p",
  n_bins = 50L,
  converged = FALSE,
  iterations = 100L,
  initial_beta = numeric()
)

The object contains positional-bias results only. Long-read observations and prior models are supplied to mpaqt_postquant(), not mpaqt_prequant().

mpaqt_quant_result

The base result contains:

list(
  abundances = numeric(),
  tpm = numeric(),
  log_likelihood = numeric(),
  converged = logical(),
  iterations = integer(),
  fitted_values = numeric(),
  prior_variance = numeric(),
  positional_weights = NULL,
  positional_bias_type = NULL,
  lr_coverage_probs = NULL,
  lr_model = NULL,
  prior_predictions = NULL,
  diagnostics = list()
)

mpaqt_postquant() adds normalization, normalized_counts, uncertainty, parameters, sr_input, and lr_input. Attributes describe the number of transcripts and whether long reads, bias correction, a prior, and uncertainty are present.

Use public accessors such as tpm(), abundances(), uncertainty(), has_uncertainty(), get_parameters(), and get_input_metadata() instead of depending on internal construction details.

Dependencies

Required R packages are declared under Imports in DESCRIPTION. Optional Bioconductor packages are checked only by workflows that need them. For example, index creation requires Biostrings and rtracklayer, genome-based transcript extraction additionally needs BSgenome-related packages, and raw FLNC processing needs bambu and its supporting packages.

The short-read FASTQ pathway requires kallisto and bustools. Raw long-read alignment additionally requires minimap2 and samtools. Callers that already have count RDS files do not need the corresponding external processing tools.

Error Handling

User-facing functions validate object classes, paths, technology names, and input-path combinations before starting expensive work. Errors and warnings are produced through the R cli package. New pathways should reuse validation and external-tool wrappers rather than invoking system commands directly.

Testing

The current tests are:

tests/testthat/
├── helper.R
├── test-algo-bias.R
├── test-algo-newton.R
├── test-class-counts.R
├── test-class-index.R
├── test-class-result.R
├── test-integration-index.R
├── test-integration-quant.R
├── test-integration-sample-prep.R
└── test-validate.R

Run the package tests from the repository root:

devtools::test()

Run the full package check before release:

devtools::check()

tests/testthat/helper.R defines reusable tool, fixture, temporary-directory, and mock-object helpers. Current integration tests skip directly when BUS parsing data or BSgenome is unavailable.

Documentation Workflow

Roxygen comments immediately above exported functions are the source for man/*.Rd and NAMESPACE. Do not edit generated files by hand.

# Regenerate man/ and NAMESPACE
devtools::document()

# Build the local pkgdown output in docs/
pkgdown::build_site(preview = FALSE)

The source articles are the .Rmd files under vignettes/. When an API signature changes, update the corresponding workflow example and add or update tests in the same change.

Contribution Checklist

  1. Keep changes focused within the relevant api-*, process-*, algo-*, infra-*, or class-* responsibility.
  2. Use four-space indentation and snake_case names.
  3. Add regression tests for changed behavior.
  4. Run devtools::test() and devtools::check().
  5. Regenerate roxygen and pkgdown output when public documentation changes.

Report issues at github.com/csglab/MPAQT/issues.