Process single-cell short-read FASTQ files with cluster aggregation. Supports optional UMI deduplication with multiple strategies.

mpaqt_prepare_short_reads_sc(
  index,
  fastq_1 = NULL,
  fastq_2 = NULL,
  clusters_file,
  bus_file = NULL,
  ec_file = NULL,
  whitelist = NULL,
  rds_dir = NULL,
  output_prefix = NULL,
  keep_bus = FALSE,
  do_umi_dedup = FALSE,
  umi_dedup_mode = "fractional_proportional",
  output_dir,
  technology = "10xv3",
  threads = 1L,
  verbose = TRUE
)

Arguments

index

An mpaqt_index object

fastq_1

Path to first FASTQ file (contains barcode + UMI)

fastq_2

Path to second FASTQ file (contains cDNA)

clusters_file

Path to cluster assignment file (barcode, cluster columns)

bus_file

Path to pre-computed BUS file (optional)

ec_file

Path to EC mapping file (optional)

whitelist

Path to barcode whitelist file for correction (optional). If NULL (default), a whitelist is auto-generated from clusters_file.

rds_dir

Directory containing pre-computed cluster RDS files (optional)

output_prefix

Prefix for output RDS files (default: "mpaqt").

keep_bus

Keep intermediate kallisto/bustools files (default: FALSE).

do_umi_dedup

Logical (default: FALSE). If TRUE, perform UMI deduplication using the strategy specified by umi_dedup_mode.

umi_dedup_mode

Character. UMI deduplication strategy (only used when do_umi_dedup = TRUE). One of:

"fractional_proportional"

(default) Read-proportional fractional counting. Mass-conserving, uses within-molecule read evidence.

"fractional_equal"

Equal fractional counting. Mass-conserving, each EC gets weight 1/k.

"unique"

Simple dedup. Each unique (barcode, UMI, EC) contributes 1 count.

"bustools_count"

Dedup via bustools count. May discard counts from newly-created ECs.

output_dir

Output directory for results

technology

Technology string: "10xv2", "10xv3", or "10xv4"

threads

Number of threads (default: 1)

verbose

Print progress messages (default: TRUE)

Value

A list of mpaqt_counts_sr objects, one per cluster (invisibly)

Details

Pipeline (all modes): kallisto bus -> sort -> correct -> sort -> bustools text

Then branch by do_umi_dedup and umi_dedup_mode:

do_umi_dedup = FALSE (default)

Read-level counting. Each read contributes 1 count to its EC.

umi_dedup_mode = "unique"

Collapse duplicate (barcode, UMI, EC) rows, then count rows per (barcode, EC). A UMI mapping to k distinct ECs contributes k counts (one per EC).

umi_dedup_mode = "fractional_equal"

Collapse duplicate (barcode, UMI, EC) rows, then assign each row weight = 1/k where k is the number of distinct ECs for that (barcode, UMI). Enforces mass conservation: each molecule contributes total mass = 1. Counts per (barcode, EC) are sum(weight).

umi_dedup_mode = "fractional_proportional"

Count reads per (barcode, UMI, EC) without deduplicating first, then assign weight = n_reads / total_reads within each molecule. Enforces mass conservation while using within-molecule read support as evidence. Counts per (barcode, EC) are sum(weight).

umi_dedup_mode = "bustools_count"

Uses bustools count for UMI deduplication. Note: the -g flag causes bustools count to create new equivalence classes not present in matrix.ec. These are discarded with a warning. Only counts mapping to original kallisto ECs are retained.

Fractional Counts

When umi_dedup_mode is "fractional_equal" or "fractional_proportional", the resulting EC counts will be fractional (non-integer). Downstream code should not assume integer counts. This is by design: fractional counting enforces mass conservation per molecule and provides a better approximation of latent assignment uncertainty, particularly at isoform-dense loci.