Process long-read data from various input types for MPAQT quantification. Supports three input pathways:
mpaqt_prepare_long_reads(
index,
flnc_fastq = NULL,
genome = NULL,
bambu_counts = NULL,
rds_file = NULL,
output_file = NULL,
keep_bam = FALSE,
keep_bambu_cache = FALSE,
output_dir = tempdir(),
threads = 1L,
minimap_preset = "splice:hq",
discovery = FALSE,
sample_id = NULL,
verbose = TRUE
)An mpaqt_index object (must have transcriptome/GTF for FLNC pathway)
Path to FLNC (full-length non-chimeric) FASTQ file
Path to genome FASTA or BSgenome object/name (for FLNC pathway)
Path to pre-computed Bambu count table (CSV/TSV)
Path to pre-computed mpaqt_counts_lr RDS file
Path to save the final RDS file (default: output_dir/mpaqt.long_read.rds)
Keep minimap2 BAM alignment files (default: FALSE). When FALSE, BAM and BAI files are deleted after Bambu quantification.
Keep Bambu intermediate cache files (default: FALSE). When FALSE, Bambu RDS output is deleted after processing.
Output directory for results (default: tempdir())
Number of threads for minimap2/Bambu (default: 1)
minimap2 preset (default: "splice:hq" for Iso-Seq)
Enable Bambu novel transcript discovery (default: FALSE)
Sample identifier (default: derived from input)
Print progress messages (default: TRUE)
An mpaqt_counts_lr object (invisibly)
Raw FLNC FASTQ: Align with minimap2, quantify with Bambu
Pre-computed Bambu counts: Load existing Bambu count table
Pre-computed RDS: Load previously saved MPAQT counts directly
The function automatically detects which pathway to use based on the arguments provided:
If rds_file is provided, loads and returns the pre-computed counts
If bambu_counts is provided, reads and converts the count table
If flnc_fastq is provided, runs minimap2 alignment + Bambu quantification
For the FLNC pathway, the index must contain stored GTF annotation and
transcriptome sequences (or you must provide a genome reference).
The FLNC pathway requires:
minimap2 and samtools installed and in PATH
Bambu Bioconductor package installed
A genome reference (FASTA or BSgenome)
GTF annotation (stored in index or external file)
By default, intermediate files are cleaned up after processing:
keep_bam = FALSE: Removes minimap2 alignment files (.bam, .bai)
keep_bambu_cache = FALSE: Removes Bambu output (.rds)
Set these to TRUE to retain files for debugging or downstream analysis.
if (FALSE) { # \dontrun{
# Load index
idx <- mpaqt_read_index("my_index/mpaqt.index.rds")
# Pathway A: From raw FLNC FASTQ (full pipeline)
lr_counts <- mpaqt_prepare_long_reads(
index = idx,
flnc_fastq = "sample.flnc.fastq.gz",
genome = "BSgenome.Hsapiens.UCSC.hg38",
output_dir = "results",
threads = 8
)
# With explicit output file and keeping intermediate files
lr_counts <- mpaqt_prepare_long_reads(
index = idx,
flnc_fastq = "sample.flnc.fastq.gz",
genome = "genome.fa",
output_file = "results/my_sample.long_read.rds",
keep_bam = TRUE,
keep_bambu_cache = TRUE,
output_dir = "results"
)
# Pathway B: From pre-computed Bambu counts
lr_counts <- mpaqt_prepare_long_reads(
index = idx,
bambu_counts = "bambu_output/counts.csv",
output_dir = "results"
)
# Pathway C: From pre-computed RDS
lr_counts <- mpaqt_prepare_long_reads(
index = idx,
rds_file = "previous_run/mpaqt.long_read.rds"
)
} # }