Multi-Platform Aggregation and Quantification of Transcripts

MPAQT is an R package for RNA-seq transcript quantification that integrates short-read and long-read sequencing data for improved accuracy. It supports both bulk and single-cell RNA-seq analysis.

Documentation: csglab.github.io/MPAQT


Features

  • Multi-platform integration: Combine Illumina short-reads with PacBio/ONT long-reads
  • Bulk and single-cell: Unified API for both analysis types
  • Positional bias correction: Account for 3’ or 5’ sequencing biases
  • Prior integration: Incorporate custom transcript-specific priors into abundance estimation
  • Flexible inputs: Start from FASTQ or pre-computed counts

Installation

Choose your preferred installation method:

Method Best For Instructions
Local Development, customization Jump to section
Conda Simple environment management Jump to section
Apptainer HPC clusters Jump to section

Local Installation

Step 1: Install R and Dependencies

Install R (>= 4.0.0) from CRAN with compilation tools (make, zlib, curl).

Install the Bioconductor packages required by mpaqt_index():

if (!requireNamespace("BiocManager", quietly = TRUE))
  install.packages("BiocManager")
BiocManager::install(c("Biostrings", "rtracklayer"))

Step 2: Install MPAQT R Package

install.packages("pak")
pak::pak("csglab/MPAQT")

The source repository is public. GitHub credentials are optional for installation and can help avoid API rate limits.

Step 3: Install System Tools

Install kallisto (>= 0.50.1) and bustools (>= 0.43.1).

Verify:

kallisto version
bustools version

Conda

conda create -n mpaqt \
  -c csglab -c conda-forge -c bioconda -c defaults \
  r-mpaqt
conda activate mpaqt

R -e 'library(mpaqt)'

The Conda package includes Biostrings, rtracklayer, kallisto, and bustools, which are required to create an index with mpaqt_index(). The defaults channel supplies the r-gpboost dependency.


Apptainer

For HPC clusters:

# Pull image
apptainer pull mpaqt_2.4.0.sif \
  library://csglab/mpaqt/mpaqt:2.4.0

# R API usage
apptainer exec mpaqt_2.4.0.sif R -e 'library(mpaqt)'

# Run interactively
apptainer shell mpaqt_2.4.0.sif

Quick Start

R API

library(mpaqt)

# 1. Create index (run once)
index <- mpaqt_index(
  annotation = "gencode.v44.gtf",
  transcriptome = "gencode.v44.transcripts.fa",
  output_file = "mpaqt.index.rds"
)

# 2. Process short reads
sr_counts <- mpaqt_prepare_short_reads(
  index = index,
  fastq_1 = "sample_R1.fastq.gz",
  fastq_2 = "sample_R2.fastq.gz",
  output_dir = "results/"
)

# 3. Quantify
result <- mpaqt_quant(
  index = index,
  sr_counts = sr_counts,
  positional_bias = "3p"
)

# 4. Get results
tpm_values <- tpm(result)

Documentation

Full documentation: https://csglab.github.io/MPAQT/

Guide Description
Installation Detailed installation guide
Bulk Workflow Complete bulk RNA-seq analysis
Single-Cell Cluster-level quantification
API Reference All functions

Citation

If you use MPAQT in your research, please cite:

Apostolides, M., Choi, B., Navickas, A., Saberi, A., Soto, L. M., Goodarzi, H., & Najafabadi, H. S. (2024). Accurate isoform quantification by joint short- and long-read RNA sequencing. BioRxiv. https://doi.org/10.1101/2024.07.11.603067


Contributing

We welcome contributions! See our Package Structure guide.

Report issues at github.com/csglab/MPAQT/issues.


License

MIT License