Installation

Install the pipline and the dependencies using conda/mamba

Attention

This conda/mamba route is supported on Linux only - it is not supported on macOS. The R packages immunopipe depends on (bioconductor-screpertoire, r-biopipen.utils, r-hitype, r-plotthis, r-scplotter, r-seuratdisk, r-seuratwrappers) are published on the pwwang channel as linux-64 builds only, so on macOS the solve fails with a list of packages that "does not exist (perhaps a typo or a missing channel)" - on Intel and Apple silicon alike.

On macOS use the docker image instead. It is the supported route there: the image carries the whole R environment, so nothing else has to be installed. pip install immunopipe also succeeds on macOS, but it installs the Python package alone - the R dependencies, and therefore every R-backed process, are missing and the pipeline will not run.

This is checked continuously by .github/workflows/install-matrix.yml: on ubuntu-latest it performs the documented installation and runs the pipeline on a minimal dataset; on macos-latest it records the conda failure as the evidence for this limitation and confirms that the pip step still succeeds.

Tip

If you plan to use the docker image to run the pipeline locally, you can skip this section.

immunopipe is built upon pipen framework, and a number of packages written in R and python. It's not recommended to install the packages manually. Instead, you can use the provided environment_base.yml to create a conda environment. Download the environment files first:

$ curl -fLO https://raw.githubusercontent.com/pwwang/immunopipe/master/docker/environment_base.yml
$ curl -fLO https://raw.githubusercontent.com/pwwang/immunopipe/master/docker/environment_rpkgs.yml

Then create the environment with the base file:

$ conda env create \
    -n immunopipe \
    -f ./environment_base.yml

And update it with the essential R packages:

$ conda env update \
    -n immunopipe \
    -f ./environment_rpkgs.yml

Attention

Do not pass the environment file as a URL (that is, do not use -f https://...). In conda env create, -f accepts a list of paths; when the argument is a URL, conda fails in its own pip step with AttributeError: 'list' object has no attribute 'split' (see conda/env/pip_util.py). The dependency solve succeeds and every package is installed before that failure, so the environment is left silently incomplete — the pip: entries (tensorflow, scvelo, liana, keras, ...) are missing, and the error only appears at the very end of a long installation. Downloading the files first, as above, avoids this. Verified with conda 26.7.2.

For more detailed instructions of conda env create, please refer to conda docs.

Attention

The pipeline itself is NOT included in the conda environment. You need to install it separately.

$ conda activate immunopipe
$ pip install -U immunopipe
$ # If you want to create diagram and generate running information
$ # or use the dry run scheduler, install with the extras
$ pip install -U immunopipe[diagram,runinfo,dry]
$ # You also need to install the frontend dependencies to generate reports
$ pipen report update

Use the docker image

You can also use the docker image to run the pipeline. The image is built upon miniconda3 and micromamba is used as the package manager. The image is available at Docker Hub.

To pull the image:

$ docker pull justold/immunopipe:<tag>

If you are using singularity, you can pull and convert the image to sif format:

$ singularity pull docker://justold/immunopipe:<tag>
$ apptainer pull docker://justold/immunopipe:<tag>

To run the pipeline use the image, please refer to Running the pipeline.

The directory structure in the container

The docker image is build upon mambaorg/micromamba:2.5.0. The OS is linux/amd64. Other than the default directories, the following directories are also created or should be mapped during the run:

  • /immunopipe: The directory where the source code of the pipeline is. It is general a clone of the repository. The pipeline is also installed from this directory.
  • /workdir: The working directory. It is the directory where the pipeline is run. It is recommended to map the current directory (.) to this directory.

Prepare to run the pipeline via Google Batch Jobs

There are two ways of running the pipeline via Google Batch Jobs: using the gbatch scheduler (provided by xqute) or using pipen-cli-gbatch. See more details in Running the pipeline via Google Batch Jobs.

In addition to prepare the docker image in the artifact registry (or docker hub if your google cloud project allows pulling from docker hub), you also need to install some dependencies locally.

If you choose to use the gbatch scheduler, in addition to installing the pipeline:

$ pip install -U immunopipe

# install cloud dependencies
$ pip install -U panpath[gs,async-gs]

You still have to install the following dependencies to generate reports:

Then you need to install frontend dependencies for report generation:

$ pipen report update

If you choose to use pipen-cli-gbatch (running the pipeline via immunopipe gbatch), you just need to install the pipeline with the cli-gbatch extra:

$ pip install -U immunopipe[cli-gbatch]