Build a comprehensive index for transcript quantification. Supports multiple input pathways:
Pathway A: GTF + FASTA - Build index from transcriptome file
Pathway B: GTF + genome - Extract transcriptome from BSgenome
Pathway C: GTF + FASTA + pre-built Kallisto index
mpaqt_index(
annotation,
transcriptome = NULL,
genome = NULL,
kallisto_index = NULL,
output_file = "mpaqt.index.rds",
read_length = 75L,
stride = 1L,
threads = 1L,
chunk_size = 10000000L,
store_sequences = FALSE,
store_gtf = FALSE,
temp_dir = tempdir(),
keep_temp = FALSE,
verbose = TRUE
)Path to GTF annotation file (required)
Path to transcriptome FASTA file (optional if genome provided)
BSgenome object or package name for transcriptome extraction (e.g., "BSgenome.Hsapiens.UCSC.hg38"). Required if transcriptome not provided.
Path to pre-built Kallisto index (optional). If provided along with transcriptome, skips Kallisto index building.
Output file path for bundled index. If ends in ".rds" or ".mpaqt.idx", creates a single bundled file. Otherwise creates directory with separate files (legacy behavior).
Expected read length (default: 75)
Stride for read simulation (default: 1)
Number of threads for parallel operations (default: 1)
Number of simulated reads per chunk (default: 10000000)
Store transcriptome sequences in index for long-read workflows (default: FALSE)
Store GTF annotation in index for long-read workflows (default: FALSE)
Temporary directory for intermediate files
Keep temporary files after completion (default: FALSE)
Print progress messages (default: TRUE)
An mpaqt_index object (invisibly)
The index includes:
Kallisto index for pseudoalignment
P matrices mapping equivalence classes to transcripts
Positional distance information for bias correction
Covariate matrices for prior estimation
Optional: GTF annotation and transcriptome sequences (for long-read workflows)
Pathway A (GTF + FASTA): Standard workflow. Provide annotation and
transcriptome. A Kallisto index will be built from the FASTA file.
Pathway B (GTF + genome): When you have a GTF but no transcriptome FASTA.
Provide annotation and genome (BSgenome package name or object).
Transcriptome sequences are extracted from the genome using GTF coordinates.
Pathway C (GTF + FASTA + pre-built index): When you already have a
Kallisto index. Provide annotation, transcriptome, and kallisto_index.
Skips Kallisto index building but still needs the FASTA to build P matrices.
Obtains transcriptome sequences (from FASTA or genome extraction)
Creates a Kallisto index (unless pre-built index provided)
Parses GTF to extract gene-transcript mappings
Simulates reads from each transcript position
Runs kallisto bus to determine equivalence class memberships
Builds P matrices with positional distance information
Creates covariate matrix for prior modeling
Bundles everything into a single portable file
PolyA tails are appended to transcripts of appropriate biotypes to better model 3' bias in poly(A)-selected libraries.
By default, creates a single bundled file containing all index components including the Kallisto index binary. This makes the index fully portable.
For long-read workflows (FLNC FASTQ processing), set store_sequences = TRUE
and store_gtf = TRUE to include the reference data needed for minimap2
alignment and Bambu quantification. Both options default to FALSE.
if (FALSE) { # \dontrun{
# Pathway A: Build from transcriptome FASTA
idx <- mpaqt_index(
annotation = "annotation.gtf",
transcriptome = "transcriptome.fa",
output_file = "my_index.mpaqt.idx"
)
# Pathway B: Extract transcriptome from genome
idx <- mpaqt_index(
annotation = "annotation.gtf",
genome = "BSgenome.Hsapiens.UCSC.hg38",
output_file = "my_index.mpaqt.idx"
)
# Pathway C: Use pre-built Kallisto index
idx <- mpaqt_index(
annotation = "annotation.gtf",
transcriptome = "transcriptome.fa",
kallisto_index = "existing.kallisto.idx",
output_file = "my_index.mpaqt.idx"
)
# View index summary
print(idx)
} # }