vignettes/package-structure.Rmd
package-structure.RmdThis article describes the current MPAQT source layout, public R API, S3 objects, and development workflow. Function signatures and object fields below correspond to package version 2.4.0.
MPAQT/
├── R/ # Package source
├── tests/testthat/ # Unit and integration tests
├── vignettes/ # Source articles
├── man/ # Generated Rd documentation
├── inst/
│ ├── apptainer/ # Apptainer build assets
│ ├── conda/ # Conda environments and setup scripts
│ └── conda-recipe/ # Conda package recipes
├── DESCRIPTION # Package metadata and dependencies
├── NAMESPACE # Generated exports and imports
└── _pkgdown.yml # Website navigation and reference groups
The local docs/ directory is ignored build output. The
pkgdown workflow builds that directory and publishes it to the
gh-pages branch; edit the README and vignette sources
instead of generated HTML.
Files in R/ are grouped by responsibility:
| Files | Responsibility |
|---|---|
api-index.R, api-quant.R
|
Main exported index and quantification entry points |
process-short-read.R,
process-long-read.R
|
Input-path detection and count preparation |
quant-pre.R, quant-post.R
|
Positional-bias prequantification and final quantification |
algo-bias.R, algo-em.R,
algo-newton.R, algo-prior.R
|
Numerical algorithms and prior fitting |
class-index.R, class-counts.R,
class-result.R
|
S3 constructors, validation-facing methods, and accessors |
infra-kallisto.R, infra-minimap2.R,
infra-bambu.R, infra-system.R
|
External-tool wrappers |
extract-transcriptome.R, validate.R
|
Reference extraction and shared validation |
mpaqt-package.R, zzz.R
|
Package-level documentation and initialization |
External commands should remain behind the infra-*
wrappers. Primary workflow functions use the mpaqt_*
prefix. Public accessors and exported require_*()
dependency checks use descriptive names, while internal validators and
tool checks use validate_*() and
check_*().
The main object flow is:
mpaqt_index()
-> mpaqt_index
mpaqt_prepare_short_reads() / mpaqt_prepare_short_reads_sc()
-> mpaqt_counts_sr or a named list of those objects
mpaqt_prepare_long_reads() / mpaqt_prepare_long_reads_sc()
-> mpaqt_counts_lr or a named list of those objects
mpaqt_quant() / mpaqt_quant_sc()
-> mpaqt_quant_result or a named list of result objects
For positional-bias workflows, mpaqt_quant() calls
mpaqt_prequant() and then passes its
mpaqt_prequant object to mpaqt_postquant().
Users can call the two phases directly when they need to save or reuse
estimated weights.
All primary objects inherit from list; summary
information that is not a list component is stored as an attribute.
mpaqt_index
list(
transcripts = character(),
genes = character(),
ec_ids = character(),
p_matrices = list(),
ec_counts = numeric(),
covariates = matrix(),
distances = data.table::data.table(),
t2g_matrix = Matrix::sparseMatrix(),
g2t_matrix = Matrix::sparseMatrix(),
t2g_normalized = Matrix::sparseMatrix(),
kallisto_index = NULL,
gtf_annotation = NULL,
transcriptome_seqs = NULL,
kallisto_binary = NULL
)p_matrices is a list with one data table per transcript.
Each table stores the one-based EC row indices in i,
design/exposure values in x, and the corresponding
dist_5p and dist_3p values. Index attributes
include n_transcripts, n_genes,
n_ec, created, mpaqt_version, and
bundled. Optional components can be NULL.
mpaqt_counts_sr
list(
counts = numeric(), # named EC-count vector
technology = "bulk", # bulk, 10xv2, 10xv3, or 10xv4
sample_id = NULL
)Its classes are mpaqt_counts_sr,
mpaqt_counts, and list. Attributes include
n_reads, n_ec, and created.
mpaqt_counts_lr
Its classes are mpaqt_counts_lr,
mpaqt_counts, and list. Attributes include
n_reads, n_transcripts, and
created.
mpaqt_prequant
list(
weights = numeric(),
quantiles = numeric(),
bias_type = "3p",
n_bins = 50L,
converged = FALSE,
iterations = 100L,
initial_beta = numeric()
)The object contains positional-bias results only. Long-read
observations and prior models are supplied to
mpaqt_postquant(), not mpaqt_prequant().
mpaqt_quant_result
The base result contains:
list(
abundances = numeric(),
tpm = numeric(),
log_likelihood = numeric(),
converged = logical(),
iterations = integer(),
fitted_values = numeric(),
prior_variance = numeric(),
positional_weights = NULL,
positional_bias_type = NULL,
lr_coverage_probs = NULL,
lr_model = NULL,
prior_predictions = NULL,
diagnostics = list()
)mpaqt_postquant() adds normalization,
normalized_counts, uncertainty,
parameters, sr_input, and
lr_input. Attributes describe the number of transcripts and
whether long reads, bias correction, a prior, and uncertainty are
present.
Use public accessors such as tpm(),
abundances(), uncertainty(),
has_uncertainty(), get_parameters(), and
get_input_metadata() instead of depending on internal
construction details.
Required R packages are declared under Imports in
DESCRIPTION. Optional Bioconductor packages are checked
only by workflows that need them. For example, index creation requires
Biostrings and rtracklayer, genome-based transcript extraction
additionally needs BSgenome-related packages, and raw FLNC processing
needs bambu and its supporting packages.
The short-read FASTQ pathway requires kallisto and bustools. Raw long-read alignment additionally requires minimap2 and samtools. Callers that already have count RDS files do not need the corresponding external processing tools.
User-facing functions validate object classes, paths, technology
names, and input-path combinations before starting expensive work.
Errors and warnings are produced through the R cli package.
New pathways should reuse validation and external-tool wrappers rather
than invoking system commands directly.
The current tests are:
tests/testthat/
├── helper.R
├── test-algo-bias.R
├── test-algo-newton.R
├── test-class-counts.R
├── test-class-index.R
├── test-class-result.R
├── test-integration-index.R
├── test-integration-quant.R
├── test-integration-sample-prep.R
└── test-validate.R
Run the package tests from the repository root:
devtools::test()Run the full package check before release:
devtools::check()tests/testthat/helper.R defines reusable tool, fixture,
temporary-directory, and mock-object helpers. Current integration tests
skip directly when BUS parsing data or BSgenome is unavailable.
Roxygen comments immediately above exported functions are the source
for man/*.Rd and NAMESPACE. Do not edit
generated files by hand.
# Regenerate man/ and NAMESPACE
devtools::document()
# Build the local pkgdown output in docs/
pkgdown::build_site(preview = FALSE)The source articles are the .Rmd files under
vignettes/. When an API signature changes, update the
corresponding workflow example and add or update tests in the same
change.
api-*,
process-*, algo-*, infra-*, or
class-* responsibility.snake_case names.devtools::test() and
devtools::check().Report issues at github.com/csglab/MPAQT/issues.