Build a comprehensive index for transcript quantification. Supports multiple input pathways:

  • Pathway A: GTF + FASTA - Build index from transcriptome file

  • Pathway B: GTF + genome - Extract transcriptome from BSgenome

  • Pathway C: GTF + FASTA + pre-built Kallisto index

mpaqt_index(
  annotation,
  transcriptome = NULL,
  genome = NULL,
  kallisto_index = NULL,
  output_file = "mpaqt.index.rds",
  read_length = 75L,
  stride = 1L,
  threads = 1L,
  chunk_size = 10000000L,
  store_sequences = FALSE,
  store_gtf = FALSE,
  temp_dir = tempdir(),
  keep_temp = FALSE,
  verbose = TRUE
)

Arguments

annotation

Path to GTF annotation file (required)

transcriptome

Path to transcriptome FASTA file (optional if genome provided)

genome

BSgenome object or package name for transcriptome extraction (e.g., "BSgenome.Hsapiens.UCSC.hg38"). Required if transcriptome not provided.

kallisto_index

Path to pre-built Kallisto index (optional). If provided along with transcriptome, skips Kallisto index building.

output_file

Output file path for bundled index. If ends in ".rds" or ".mpaqt.idx", creates a single bundled file. Otherwise creates directory with separate files (legacy behavior).

read_length

Expected read length (default: 75)

stride

Stride for read simulation (default: 1)

threads

Number of threads for parallel operations (default: 1)

chunk_size

Number of simulated reads per chunk (default: 10000000)

store_sequences

Store transcriptome sequences in index for long-read workflows (default: FALSE)

store_gtf

Store GTF annotation in index for long-read workflows (default: FALSE)

temp_dir

Temporary directory for intermediate files

keep_temp

Keep temporary files after completion (default: FALSE)

verbose

Print progress messages (default: TRUE)

Value

An mpaqt_index object (invisibly)

Details

The index includes:

  • Kallisto index for pseudoalignment

  • P matrices mapping equivalence classes to transcripts

  • Positional distance information for bias correction

  • Covariate matrices for prior estimation

  • Optional: GTF annotation and transcriptome sequences (for long-read workflows)

Input Pathways

Pathway A (GTF + FASTA): Standard workflow. Provide annotation and transcriptome. A Kallisto index will be built from the FASTA file.

Pathway B (GTF + genome): When you have a GTF but no transcriptome FASTA. Provide annotation and genome (BSgenome package name or object). Transcriptome sequences are extracted from the genome using GTF coordinates.

Pathway C (GTF + FASTA + pre-built index): When you already have a Kallisto index. Provide annotation, transcriptome, and kallisto_index. Skips Kallisto index building but still needs the FASTA to build P matrices.

Index Creation Process

  1. Obtains transcriptome sequences (from FASTA or genome extraction)

  2. Creates a Kallisto index (unless pre-built index provided)

  3. Parses GTF to extract gene-transcript mappings

  4. Simulates reads from each transcript position

  5. Runs kallisto bus to determine equivalence class memberships

  6. Builds P matrices with positional distance information

  7. Creates covariate matrix for prior modeling

  8. Bundles everything into a single portable file

PolyA tails are appended to transcripts of appropriate biotypes to better model 3' bias in poly(A)-selected libraries.

Output

By default, creates a single bundled file containing all index components including the Kallisto index binary. This makes the index fully portable.

For long-read workflows (FLNC FASTQ processing), set store_sequences = TRUE and store_gtf = TRUE to include the reference data needed for minimap2 alignment and Bambu quantification. Both options default to FALSE.

Examples

if (FALSE) { # \dontrun{
# Pathway A: Build from transcriptome FASTA
idx <- mpaqt_index(
    annotation = "annotation.gtf",
    transcriptome = "transcriptome.fa",
    output_file = "my_index.mpaqt.idx"
)

# Pathway B: Extract transcriptome from genome
idx <- mpaqt_index(
    annotation = "annotation.gtf",
    genome = "BSgenome.Hsapiens.UCSC.hg38",
    output_file = "my_index.mpaqt.idx"
)

# Pathway C: Use pre-built Kallisto index
idx <- mpaqt_index(
    annotation = "annotation.gtf",
    transcriptome = "transcriptome.fa",
    kallisto_index = "existing.kallisto.idx",
    output_file = "my_index.mpaqt.idx"
)

# View index summary
print(idx)
} # }