Process single-cell long-read FLNC (full-length non-chimeric) FASTQ files with user-provided cluster assignments.

mpaqt_prepare_long_reads_sc_flnc(
  index,
  flnc_fastq,
  genome,
  gtf = NULL,
  clusters_file,
  output_prefix = NULL,
  output_dir,
  keep_bam = FALSE,
  keep_bambu_cache = FALSE,
  platform = "PacBio",
  discovery = FALSE,
  ndr = 1,
  threads = 1L,
  verbose = TRUE
)

Arguments

index

An mpaqt_index object

flnc_fastq

Path to pre-demultiplexed FLNC FASTQ file with barcodes

genome

Path to genome FASTA or BSgenome object/name

gtf

Path to GTF annotation file (or uses index GTF if available)

clusters_file

Path to cluster assignment file (CSV with barcode, cluster columns). This is REQUIRED - MPAQT does not perform automatic cell clustering.

output_prefix

Prefix for output RDS files (default: "mpaqt"). Output files will be named {output_prefix}.{cluster_id}.long_read.rds

output_dir

Output directory for results

keep_bam

Keep minimap2 BAM alignment files (default: FALSE)

keep_bambu_cache

Keep Bambu intermediate cache files (default: FALSE)

platform

Sequencing platform: "PacBio" or "ONT" (default: "PacBio"). Determines minimap2 alignment preset.

discovery

Enable Bambu novel transcript discovery (default: FALSE)

ndr

Novel discovery rate threshold for Bambu (default: 1)

threads

Number of threads for minimap2/Bambu (default: 1)

verbose

Print progress messages (default: TRUE)

Value

A list of mpaqt_counts_lr objects, one per cluster (invisibly)

Details

This function processes single-cell long-read data from pre-demultiplexed FLNC FASTQ files. The input FASTQ should have cell barcodes already extracted and associated with each read.

The pipeline:

  1. Reads cluster assignments mapping barcodes to clusters

  2. Aligns long reads with minimap2 (preset based on platform)

  3. Quantifies transcripts with Bambu

  4. Splits counts by cluster using the provided assignments

  5. Creates mpaqt_counts_lr object per cluster

Requirements

Processing raw FLNC data requires:

  • minimap2 and samtools installed and in PATH

  • Bambu Bioconductor package installed

Check requirements with require_sc_long_read_flnc().

Cluster Assignments

MPAQT does NOT perform automatic cell clustering. Users must provide pre-computed cluster assignments, typically from short-read scRNA-seq analysis or other clustering methods.

The clusters_file should be a CSV with at least two columns:

  • Column 1: Cell barcode

  • Column 2: Cluster assignment (numeric or character)

Intermediate File Handling

By default, intermediate files are cleaned up after processing:

  • keep_bam = FALSE: Removes minimap2 alignment files (.bam, .bai)

  • keep_bambu_cache = FALSE: Removes Bambu output (.rds)

Examples

if (FALSE) { # \dontrun{
idx <- mpaqt_read_index("my_index/mpaqt.index.rds")

# Process SC long-read FLNC with provided clusters
lr_counts_list <- mpaqt_prepare_long_reads_sc_flnc(
    index = idx,
    flnc_fastq = "sample.flnc.fastq.gz",
    genome = "genome.fa",
    gtf = "annotation.gtf",
    clusters_file = "clusters.csv",
    output_dir = "results",
    platform = "PacBio",
    threads = 8
)

# Keep intermediate files for debugging
lr_counts_list <- mpaqt_prepare_long_reads_sc_flnc(
    index = idx,
    flnc_fastq = "sample.flnc.fastq.gz",
    genome = "genome.fa",
    gtf = "annotation.gtf",
    clusters_file = "clusters.csv",
    output_prefix = "my_sample",
    keep_bam = TRUE,
    keep_bambu_cache = TRUE,
    output_dir = "results"
)
} # }