Command-line interface#
The installed gpn command exposes a deliberately small maintained surface. It
does not expose dataset construction or every historical scorer module.
gpn ss {train,vep,logits,embedding} ...
gpn msa {vep,logits,embedding} ...
gpn star {train,vep,logits,embedding} ...
ss denotes the single-species GPN family. Cyclopts derives each command from
its typed Python signature. Use gpn <family> <command> --help for the complete
arguments, including the pinned Transformers TrainingArguments surface.
Installation#
Inference commands require the inference extra:
pip install "gpn[inference]"
GPN and GPN-Star training require the training extra:
pip install "gpn[train]"
GPN#
Train GPN from a prepared dataset using one of the maintained recipe profiles:
gpn ss train recipes/gpn_training/cpu-smoke.yaml
Score variants whose pos column uses one-based VCF coordinates:
gpn ss vep \
--input-path variants.parquet \
--genome-path genome.fa.gz \
--window-size 512 \
--model-path songlab/gpn-brassicales \
--output-path scores.parquet \
--per-device-eval-batch-size 64
The same family group exposes masked-nucleotide logits and sequence embeddings:
gpn ss logits \
--input-path POSITIONS_PATH \
--genome-path GENOME_PATH \
--window-size WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
gpn ss embedding \
--input-path WINDOWS_PATH \
--genome-path GENOME_PATH \
--center-window-size CENTER_WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
Embedding averages exactly CENTER_WINDOW_SIZE positions, for both odd and even
values.
VEP inputs must be biallelic SNVs: ref and alt are distinct, uppercase, single
characters from A, C, G, and T. The reference allele must match the local
genome or MSA at chrom:pos; inference stops on a mismatch.
The output score is the raw, uncalibrated alternate-minus-reference
log-likelihood ratio averaged over forward and reverse-complement sequence
orientations. In particular, negative values mean the alternate is less likely
than the reference under the model.
GPN-MSA#
GPN-MSA is deprecated and supports inference only. Its public commands cover variant scoring, masked-nucleotide logits, and embeddings:
gpn msa vep \
--input-path INPUT_PATH \
--msa-path LOCAL_MSA_PATH \
--window-size WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
gpn msa logits \
--input-path INPUT_PATH \
--msa-path LOCAL_MSA_PATH \
--window-size WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
gpn msa embedding \
--input-path INPUT_PATH \
--msa-path LOCAL_MSA_PATH \
--window-size WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
LOCAL_MSA_PATH must point directly to an existing GPN-MSA Zarr store. Its target
genome, species count, species order, and preprocessing must match the selected
checkpoint. The CLI never downloads a whole-genome alignment. Parquet, CSV/TSV,
VCF, GTF, and GFF inputs are detected from existing local paths; otherwise
INPUT_PATH is passed to Hugging Face Datasets as a dataset identifier.
VEP and logits use one-based pos coordinates and require an even WINDOW_SIZE.
Embedding inputs instead use zero-based, half-open start and end coordinates.
The embedding result averages exactly CENTER_WINDOW_SIZE positions, for both
odd and even values.
GPN-MSA VEP emits the same raw, uncalibrated forward/reverse-averaged LLR described
above.
GPN-Star#
Train from prepared intervals and local MSAs with the maintained recipe:
gpn star train recipes/gpn_star_training/cpu-smoke.yaml
The three core inference operations share this shape:
gpn star vep \
--input-path INPUT_PATH \
--msa-path LOCAL_MSA_PATH \
--window-size WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
gpn star logits \
--input-path INPUT_PATH \
--msa-path LOCAL_MSA_PATH \
--window-size WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
gpn star embedding \
--input-path INPUT_PATH \
--msa-path LOCAL_MSA_PATH \
--window-size WINDOW_SIZE \
--model-path MODEL_PATH \
--output-path OUTPUT_PATH
All three inference families support the same durable, process-count-independent
batch checkpoints. Set --checkpoint-batch-size; the directory defaults to
OUTPUT_PATH_checkpoints. See any inference command’s help for cleanup options.
--checkpoint-revision is an optional user-supplied identity for the input state
used to validate resumable output batches.
LOCAL_MSA_PATH is not itself an all.zarr store. It is either a numeric
species-count directory containing all.zarr, or a parent containing one or more
such numeric directories. Every selected alignment must match the target genome,
species order, and evolutionary scale expected by the checkpoint. VEP and logits
use one-based pos coordinates and an even WINDOW_SIZE; embeddings use
zero-based, half-open start and end. Embeddings average exactly
CENTER_WINDOW_SIZE positions, for both odd and even values. GPN-Star VEP outputs
a raw, uncalibrated LLR. Any checkpoint-specific calibration is a separate
downstream operation.
AutoClass-only inference#
PhyloGPN and the sorghum gene-expression fine-tune remain supported through the
explicit register_auto_classes(...) plus Transformers AutoClass APIs documented
in the README and their model cards. They intentionally do not have dedicated CLI
commands.
Devices and distributed execution#
Inference defaults to FP32 without compilation on every device. This makes the
default explicit and predictable. A typical GPU invocation adds
--bf16-full-eval --torch-compile; unsupported compilation fails visibly for
every model family. Other current and future Trainer flags are exposed directly
because this release pins Transformers exactly.
GPN forces the small set of prediction invariants it owns: training and
evaluation phases are disabled, predictions are retained, input columns are not
pruned, incomplete batches are retained, and dataloader results stay in input
order. --push-to-hub is rejected. Direct inference also verifies that it
produced exactly one output row per input row. OUTPUT_PATH is always the
scientific Parquet result; Transformers’ --output-dir is only Trainer working
state and defaults to a temporary directory when omitted.
For multi-GPU inference, launch that same CLI through torchrun. Every process
participates in prediction, while rank zero alone commits checkpoint batches and
the final Parquet output:
torchrun --standalone --nproc-per-node=4 --module gpn.cli \
ss vep \
--input-path variants.parquet \
--genome-path genome.fa.gz \
--window-size 512 \
--model-path songlab/gpn-brassicales \
--output-path scores.parquet \
--per-device-eval-batch-size 64 \
--dataloader-num-workers 4 \
--bf16-full-eval
--per-device-eval-batch-size is per process, so multiply it by the number of
processes for the nominal global batch size. Durable checkpoints use the same
process-count-independent row ranges in single- and multi-GPU runs.
--dataloader-num-workers is also per process and defaults to zero. In the
example above, four local ranks each create four DataLoader workers: 16 worker
processes in total, in addition to the four rank processes. Size the CPU
allocation for all local ranks and workers; the flag is not a job-wide worker
budget.
OMP_NUM_THREADS controls CPU threading separately from the DataLoader worker
count and is not required for correctness. When more than one local process is
launched and the variable is unset, torchrun sets it to 1 to avoid CPU
oversubscription. Keep that default unless profiling identifies a CPU compute
bottleneck; if you tune it, account for the threads used by every local process
within the job’s CPU allocation.
Distributed training uses the public module entry point as well:
torchrun --standalone --nproc-per-node=4 --module gpn.cli \
star train recipes/gpn_star_training/gpu.yaml
The supported command-line contract is the gpn family tree above. Python
modules behind it are implementation details unless documented as public APIs.