Process long-read data from various input types for MPAQT quantification. Supports three input pathways:

mpaqt_prepare_long_reads(
  index,
  flnc_fastq = NULL,
  genome = NULL,
  bambu_counts = NULL,
  rds_file = NULL,
  output_file = NULL,
  keep_bam = FALSE,
  keep_bambu_cache = FALSE,
  output_dir = tempdir(),
  threads = 1L,
  minimap_preset = "splice:hq",
  discovery = FALSE,
  sample_id = NULL,
  verbose = TRUE
)

Arguments

index

An mpaqt_index object (must have transcriptome/GTF for FLNC pathway)

flnc_fastq

Path to FLNC (full-length non-chimeric) FASTQ file

genome

Path to genome FASTA or BSgenome object/name (for FLNC pathway)

bambu_counts

Path to pre-computed Bambu count table (CSV/TSV)

rds_file

Path to pre-computed mpaqt_counts_lr RDS file

output_file

Path to save the final RDS file (default: output_dir/mpaqt.long_read.rds)

keep_bam

Keep minimap2 BAM alignment files (default: FALSE). When FALSE, BAM and BAI files are deleted after Bambu quantification.

keep_bambu_cache

Keep Bambu intermediate cache files (default: FALSE). When FALSE, Bambu RDS output is deleted after processing.

output_dir

Output directory for results (default: tempdir())

threads

Number of threads for minimap2/Bambu (default: 1)

minimap_preset

minimap2 preset (default: "splice:hq" for Iso-Seq)

discovery

Enable Bambu novel transcript discovery (default: FALSE)

sample_id

Sample identifier (default: derived from input)

verbose

Print progress messages (default: TRUE)

Value

An mpaqt_counts_lr object (invisibly)

Details

  1. Raw FLNC FASTQ: Align with minimap2, quantify with Bambu

  2. Pre-computed Bambu counts: Load existing Bambu count table

  3. Pre-computed RDS: Load previously saved MPAQT counts directly

The function automatically detects which pathway to use based on the arguments provided:

  • If rds_file is provided, loads and returns the pre-computed counts

  • If bambu_counts is provided, reads and converts the count table

  • If flnc_fastq is provided, runs minimap2 alignment + Bambu quantification

For the FLNC pathway, the index must contain stored GTF annotation and transcriptome sequences (or you must provide a genome reference).

Requirements for FLNC Pathway

The FLNC pathway requires:

  • minimap2 and samtools installed and in PATH

  • Bambu Bioconductor package installed

  • A genome reference (FASTA or BSgenome)

  • GTF annotation (stored in index or external file)

Intermediate File Handling

By default, intermediate files are cleaned up after processing:

  • keep_bam = FALSE: Removes minimap2 alignment files (.bam, .bai)

  • keep_bambu_cache = FALSE: Removes Bambu output (.rds)

Set these to TRUE to retain files for debugging or downstream analysis.

Examples

if (FALSE) { # \dontrun{
# Load index
idx <- mpaqt_read_index("my_index/mpaqt.index.rds")

# Pathway A: From raw FLNC FASTQ (full pipeline)
lr_counts <- mpaqt_prepare_long_reads(
    index = idx,
    flnc_fastq = "sample.flnc.fastq.gz",
    genome = "BSgenome.Hsapiens.UCSC.hg38",
    output_dir = "results",
    threads = 8
)

# With explicit output file and keeping intermediate files
lr_counts <- mpaqt_prepare_long_reads(
    index = idx,
    flnc_fastq = "sample.flnc.fastq.gz",
    genome = "genome.fa",
    output_file = "results/my_sample.long_read.rds",
    keep_bam = TRUE,
    keep_bambu_cache = TRUE,
    output_dir = "results"
)

# Pathway B: From pre-computed Bambu counts
lr_counts <- mpaqt_prepare_long_reads(
    index = idx,
    bambu_counts = "bambu_output/counts.csv",
    output_dir = "results"
)

# Pathway C: From pre-computed RDS
lr_counts <- mpaqt_prepare_long_reads(
    index = idx,
    rds_file = "previous_run/mpaqt.long_read.rds"
)
} # }