Installation¶
Install the pipline and the dependencies using conda/mamba¶
Attention
This conda/mamba route is supported on Linux only - it is not supported on macOS.
The R packages immunopipe depends on (bioconductor-screpertoire, r-biopipen.utils,
r-hitype, r-plotthis, r-scplotter, r-seuratdisk, r-seuratwrappers) are
published on the pwwang channel as linux-64 builds only, so on macOS the solve
fails with a list of packages that "does not exist (perhaps a typo or a missing
channel)" - on Intel and Apple silicon alike.
On macOS use the docker image instead. It is the supported
route there: the image carries the whole R environment, so nothing else has to be
installed. pip install immunopipe also succeeds on macOS, but it installs the Python
package alone - the R dependencies, and therefore every R-backed process, are missing
and the pipeline will not run.
This is checked continuously by .github/workflows/install-matrix.yml: on
ubuntu-latest it performs the documented installation and runs the pipeline on a
minimal dataset; on macos-latest it records the conda failure as the evidence for
this limitation and confirms that the pip step still succeeds.
Tip
If you plan to use the docker image to run the pipeline locally, you can skip this section.
immunopipe is built upon pipen framework, and a number of packages written in R and python. It's not recommended to install the packages manually. Instead, you can use the provided environment_base.yml to create a conda environment. Download the environment files first:
$ curl -fLO https://raw.githubusercontent.com/pwwang/immunopipe/master/docker/environment_base.yml
$ curl -fLO https://raw.githubusercontent.com/pwwang/immunopipe/master/docker/environment_rpkgs.yml
Then create the environment with the base file:
$ conda env create \
-n immunopipe \
-f ./environment_base.yml
And update it with the essential R packages:
$ conda env update \
-n immunopipe \
-f ./environment_rpkgs.yml
Attention
Do not pass the environment file as a URL (that is, do not use -f https://...). In conda env create, -f accepts a list of paths; when the argument is a URL, conda fails in its own pip step with AttributeError: 'list' object has no attribute 'split' (see conda/env/pip_util.py). The dependency solve succeeds and every package is installed before that failure, so the environment is left silently incomplete — the pip: entries (tensorflow, scvelo, liana, keras, ...) are missing, and the error only appears at the very end of a long installation. Downloading the files first, as above, avoids this. Verified with conda 26.7.2.
For more detailed instructions of conda env create, please refer to conda docs.
Attention
The pipeline itself is NOT included in the conda environment. You need to install it separately.
$ conda activate immunopipe
$ pip install -U immunopipe
$ # If you want to create diagram and generate running information
$ # or use the dry run scheduler, install with the extras
$ pip install -U immunopipe[diagram,runinfo,dry]
$ # You also need to install the frontend dependencies to generate reports
$ pipen report update
Use the docker image¶
You can also use the docker image to run the pipeline. The image is built upon miniconda3 and micromamba is used as the package manager. The image is available at Docker Hub.
To pull the image:
$ docker pull justold/immunopipe:<tag>
If you are using singularity, you can pull and convert the image to sif format:
$ singularity pull docker://justold/immunopipe:<tag>
$ apptainer pull docker://justold/immunopipe:<tag>
To run the pipeline use the image, please refer to Running the pipeline.
The directory structure in the container¶
The docker image is build upon mambaorg/micromamba:2.5.0. The OS is linux/amd64. Other than the default directories, the following directories are also created or should be mapped during the run:
/immunopipe: The directory where the source code of the pipeline is. It is general a clone of the repository. The pipeline is also installed from this directory./workdir: The working directory. It is the directory where the pipeline is run. It is recommended to map the current directory (.) to this directory.
Prepare to run the pipeline via Google Batch Jobs¶
There are two ways of running the pipeline via Google Batch Jobs: using the gbatch scheduler (provided by xqute) or using pipen-cli-gbatch. See more details in Running the pipeline via Google Batch Jobs.
In addition to prepare the docker image in the artifact registry (or docker hub if your google cloud project allows pulling from docker hub), you also need to install some dependencies locally.
If you choose to use the gbatch scheduler, in addition to installing the pipeline:
$ pip install -U immunopipe
# install cloud dependencies
$ pip install -U panpath[gs,async-gs]
You still have to install the following dependencies to generate reports:
nodejs: Follow the instructions at https://nodejs.org/en/download/package-manager/ to installnodejsfor your system (v20+ is required); orbunjs: Follow the instructions at https://bun.sh/docs/install to installbunjsfor your system.
Then you need to install frontend dependencies for report generation:
$ pipen report update
If you choose to use pipen-cli-gbatch (running the pipeline via immunopipe gbatch), you just need to install the pipeline with the cli-gbatch extra:
$ pip install -U immunopipe[cli-gbatch]