R/process-long-read.R
mpaqt_prepare_long_reads_sc_flnc.RdProcess single-cell long-read FLNC (full-length non-chimeric) FASTQ files with user-provided cluster assignments.
mpaqt_prepare_long_reads_sc_flnc(
index,
flnc_fastq,
genome,
gtf = NULL,
clusters_file,
output_prefix = NULL,
output_dir,
keep_bam = FALSE,
keep_bambu_cache = FALSE,
platform = "PacBio",
discovery = FALSE,
ndr = 1,
threads = 1L,
verbose = TRUE
)An mpaqt_index object
Path to pre-demultiplexed FLNC FASTQ file with barcodes
Path to genome FASTA or BSgenome object/name
Path to GTF annotation file (or uses index GTF if available)
Path to cluster assignment file (CSV with barcode, cluster columns). This is REQUIRED - MPAQT does not perform automatic cell clustering.
Prefix for output RDS files (default: "mpaqt").
Output files will be named {output_prefix}.{cluster_id}.long_read.rds
Output directory for results
Keep minimap2 BAM alignment files (default: FALSE)
Keep Bambu intermediate cache files (default: FALSE)
Sequencing platform: "PacBio" or "ONT" (default: "PacBio"). Determines minimap2 alignment preset.
Enable Bambu novel transcript discovery (default: FALSE)
Novel discovery rate threshold for Bambu (default: 1)
Number of threads for minimap2/Bambu (default: 1)
Print progress messages (default: TRUE)
A list of mpaqt_counts_lr objects, one per cluster (invisibly)
This function processes single-cell long-read data from pre-demultiplexed FLNC FASTQ files. The input FASTQ should have cell barcodes already extracted and associated with each read.
The pipeline:
Reads cluster assignments mapping barcodes to clusters
Aligns long reads with minimap2 (preset based on platform)
Quantifies transcripts with Bambu
Splits counts by cluster using the provided assignments
Creates mpaqt_counts_lr object per cluster
Processing raw FLNC data requires:
minimap2 and samtools installed and in PATH
Bambu Bioconductor package installed
Check requirements with require_sc_long_read_flnc().
MPAQT does NOT perform automatic cell clustering. Users must provide pre-computed cluster assignments, typically from short-read scRNA-seq analysis or other clustering methods.
The clusters_file should be a CSV with at least two columns:
Column 1: Cell barcode
Column 2: Cluster assignment (numeric or character)
By default, intermediate files are cleaned up after processing:
keep_bam = FALSE: Removes minimap2 alignment files (.bam, .bai)
keep_bambu_cache = FALSE: Removes Bambu output (.rds)
if (FALSE) { # \dontrun{
idx <- mpaqt_read_index("my_index/mpaqt.index.rds")
# Process SC long-read FLNC with provided clusters
lr_counts_list <- mpaqt_prepare_long_reads_sc_flnc(
index = idx,
flnc_fastq = "sample.flnc.fastq.gz",
genome = "genome.fa",
gtf = "annotation.gtf",
clusters_file = "clusters.csv",
output_dir = "results",
platform = "PacBio",
threads = 8
)
# Keep intermediate files for debugging
lr_counts_list <- mpaqt_prepare_long_reads_sc_flnc(
index = idx,
flnc_fastq = "sample.flnc.fastq.gz",
genome = "genome.fa",
gtf = "annotation.gtf",
clusters_file = "clusters.csv",
output_prefix = "my_sample",
keep_bam = TRUE,
keep_bambu_cache = TRUE,
output_dir = "results"
)
} # }