Process single-cell short-read FASTQ files with cluster aggregation. Supports optional UMI deduplication with multiple strategies.
mpaqt_prepare_short_reads_sc(
index,
fastq_1 = NULL,
fastq_2 = NULL,
clusters_file,
bus_file = NULL,
ec_file = NULL,
whitelist = NULL,
rds_dir = NULL,
output_prefix = NULL,
keep_bus = FALSE,
do_umi_dedup = FALSE,
umi_dedup_mode = "fractional_proportional",
output_dir,
technology = "10xv3",
threads = 1L,
verbose = TRUE
)An mpaqt_index object
Path to first FASTQ file (contains barcode + UMI)
Path to second FASTQ file (contains cDNA)
Path to cluster assignment file (barcode, cluster columns)
Path to pre-computed BUS file (optional)
Path to EC mapping file (optional)
Path to barcode whitelist file for correction (optional).
If NULL (default), a whitelist is auto-generated from clusters_file.
Directory containing pre-computed cluster RDS files (optional)
Prefix for output RDS files (default: "mpaqt").
Keep intermediate kallisto/bustools files (default: FALSE).
Logical (default: FALSE). If TRUE, perform UMI
deduplication using the strategy specified by umi_dedup_mode.
Character. UMI deduplication strategy (only used when
do_umi_dedup = TRUE). One of:
"fractional_proportional"(default) Read-proportional fractional counting. Mass-conserving, uses within-molecule read evidence.
"fractional_equal"Equal fractional counting. Mass-conserving, each EC gets weight 1/k.
"unique"Simple dedup. Each unique (barcode, UMI, EC) contributes 1 count.
"bustools_count"Dedup via bustools count. May
discard counts from newly-created ECs.
Output directory for results
Technology string: "10xv2", "10xv3", or "10xv4"
Number of threads (default: 1)
Print progress messages (default: TRUE)
A list of mpaqt_counts_sr objects, one per cluster (invisibly)
Pipeline (all modes):
kallisto bus -> sort -> correct -> sort -> bustools text
Then branch by do_umi_dedup and umi_dedup_mode:
do_umi_dedup = FALSE (default)Read-level counting. Each read contributes 1 count to its EC.
umi_dedup_mode = "unique"Collapse duplicate (barcode, UMI, EC) rows, then count rows per (barcode, EC). A UMI mapping to k distinct ECs contributes k counts (one per EC).
umi_dedup_mode = "fractional_equal"Collapse duplicate (barcode, UMI, EC) rows, then assign each row
weight = 1/k where k is the number of distinct ECs for that
(barcode, UMI). Enforces mass conservation: each molecule contributes
total mass = 1. Counts per (barcode, EC) are sum(weight).
umi_dedup_mode = "fractional_proportional"Count reads per (barcode, UMI, EC) without deduplicating
first, then assign weight = n_reads / total_reads within each
molecule. Enforces mass conservation while using within-molecule
read support as evidence. Counts per (barcode, EC) are
sum(weight).
umi_dedup_mode = "bustools_count"Uses bustools count for UMI deduplication. Note: the
-g flag causes bustools count to create new
equivalence classes not present in matrix.ec. These are
discarded with a warning. Only counts mapping to original kallisto
ECs are retained.
When umi_dedup_mode is "fractional_equal" or
"fractional_proportional", the resulting EC counts will be
fractional (non-integer). Downstream code should not assume integer
counts. This is by design: fractional counting enforces mass
conservation per molecule and provides a better approximation of
latent assignment uncertainty, particularly at isoform-dense loci.