module

biopipen.ns.bam

Tools to process sam/bam/cram files

Classes
class

biopipen.ns.bam.CNVpytor(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Detect CNV using CNVpytor

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam fileWill try to index it if it's not indexed.
  • snpfile — The snp file
Output
  • outdir — The output directory
Envs
  • baf_nomask — Do not use P mask in BAF histograms
  • binsizes — The binsizes
  • chrom — The chromosomes to run on
  • chrsize — The geome size file to fix missing contigs in VCF header
  • cnvpytor — Path to cnvpytor
  • filters — The filters to filter the resultSee - https://github.com/abyzovlab/CNVpytor/blob/master/GettingStarted.md#predicting-cnv-regions
  • genome — The genome assembly to put in the VCF file
  • mask_snps — Whether mask 1000 Genome snps
  • ncores — Number of cores to use (-j for cnvpytor)
  • refdir — The directory containing the fasta file for each chromosome
  • samtools — Path to samtools, used to index bam file in case it's not
  • snp — How to read snp data
Requires
  • cnvpytor —
    • check: {{proc.envs.cnvpytor}} --version
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.ControlFREEC(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Detect CNVs using Control-FREEC

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam file
  • snpfile — The snp file
Output
  • outdir — The output directory
Envs
  • arggs — Other arguments for Control-FREEC
  • freec — Path to Control-FREEC executable
  • ncores — Number of cores to use
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.CNAClinic(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Detect CNVs using CNAClinic

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • metafile — The meta file, header included, tab-delimited, includingfollowing columns:
    • - Bam: The path to bam file
    • - Sample: Optional. The sample names,
        if you don't want filename of bam file to be used
    • - Group: Optional. The group names, either "Case" or "Control"
    • - Patient: Optional. The patient names. Since CNAClinic only
        supports paired samples, you need to provide the patient names
        for each sample. Required if "Group" is provided.
    • - Binsizer: Optional. Samples used to estimate the bin size
        "Y", "Yes", "T", "True", will be treated as True
        If not provided, will use envs.binsizer to get the samples
        to use. Either this column or envs.binsizer should be
        provided.
Output
  • outdir — The output directory
Envs
  • binsize — Directly use this binsize for CNAClinic, in bp.
  • binsizer — The samples used to estimate the bin size, it could be:A list of sample names A float number (0 < x <= 1), the fraction of samples to use A integer number (x > 1), the number of samples to use
  • genome — The genome assembly
  • ncores — Number of cores to use
  • plot_args — The arguments for CNAClinic::plotSampleData
  • plot_multi_args — The arguments for CNAClinic::plotMultiSampleData
  • run_args — The arguments for CNAClinic::runSegmentation
  • seed — The seed for random number generator for choosing samplesfor estimating bin size
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.BamSplitChroms(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Split bam file by chromosomes

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam file
Output
  • outdir — The output directory with bam files for each chromosome
Envs
  • chroms — The chromosomes to keep, if not provided, will use all
  • index — Whether to index the output bam files. Requires the input bamfile to be sorted.
  • keep_other_sq — Keep other chromosomes in "@SQ" field in header
  • ncores — Number of cores to use
  • sambamba — Path to sambamba executable
  • samtools — Path to samtools executable
  • tool — The tool to use, either "samtools" or "sambamba"
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.BamMerge(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Merge bam files

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfiles — The bam files
Output
  • outfile — The output bam file
Envs
  • index — Whether to index the output bam fileRequires envs.sort to be True
  • merge_args — The arguments for merging bam filessamtools merge or sambamba merge, depending on tool For samtools, these keys are not allowed: -o, -O, --output-fmt, -@, and --threads, as they are managed by the script For sambamba, these keys are not allowed: -t, and --nthreads, as they are managed by the script
  • ncores — Number of cores to use
  • sambamba — Path to sambamba executable
  • samtools — Path to samtools executable
  • sort — Whether to sort the output bam file
  • sort_args — The arguments for sorting bam filessamtools sort or sambamba sort, depending on tool For samtools, these keys are not allowed: -o, -@, and --threads, as they are managed by the script For sambamba, these keys are not allowed: -t, --nthreads, -o and --out, as they are managed by the script
  • tool — The tool to use, either "samtools" or "sambamba"
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.BamSampling(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Keeping only a fraction of read pairs from a bam file

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam file
Output
  • outfile — The output bam file
Envs
  • fraction (type=float) — The fraction of reads to keep.If 0 < fraction <= 1, it's the fraction of reads to keep. If fraction > 1, it's the number of reads to keep. Note that when fraction > 1, you may not get the exact number of reads specified but a close number.
  • index — Whether to index the output bam file
  • ncores — Number of cores to use
  • samtools — Path to samtools executable
  • seed — The seed for random number generator
  • sort — Whether to sort the output bam file
  • sort_args — The arguments for sorting bam file using samtools sort.These keys are not allowed: -o, -@, and --threads, as they are managed by the script.
  • tool — The tool to use, currently only "samtools" is supported
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.BamSubsetByBed(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Subset bam file by the regions in a bed file

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam file
  • bedfile — The bed file
Output
  • outfile — The output bam file
Envs
  • index — Whether to index the output bam file
  • ncores — Number of cores to use
  • samtools — Path to samtools executable
  • tool — The tool to use, currently only "samtools" is supported
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.BamSort(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Sort bam file

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam file
Output
  • outfile — The output bam file
Envs
  • — Other arguments passed to the sorting toolSee samtools sort or sambamba sort
  • byname (flag) — Whether to sort by read name
  • index (flag) — Whether to index the output bam fileThe index file will be created in the same directory as the output bam file
  • ncores (type=int) — Number of cores to use
  • sambamba — Path to sambamba executable
  • samtools — Path to samtools executable
  • tmpdir — The temporary directory to use
  • tool (choice) — The tool to use.
    • - samtools: Use samtools
    • - sambamba: Use sambamba
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.SamtoolsView(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

View bam file using samtools, mostly used for filtering

This is a wrapper for samtools view command. It will create a new bam file with the same name as the input bam file.

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam file
Output
  • outfile — The output bam file
Envs
  • — Other arguments passed to the view toolSee samtools view or sambamba view.
  • index — Whether to index the output bam fileRequires the input bam file to be sorted.
  • ncores — Number of cores to use
  • samtools — Path to samtools executable
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.SamplotBam(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Plot bam file using samplot

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfiles — The bam files
Output
  • outfile — The output plot fileat chrom:start-end.
Envs
  • — Other arguments passed to the samplot toolSee samplot plot command.
  • chrom — The chromosome to plot
  • end — The end position to plot
  • same_yaxis_scales (flag) — Whether to use the same y-axis scales forall bam files
  • samplot — Path to samplot executable
  • start — The start position to plot
  • titles (list) — The titles for each bam file, in the same order as bamfiles
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.BedtoolsCoverageBam(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Coverage report by bedtools coverage for bam files

The input bedfile and bamfile will be passed to bedtools coverage command bedtools coverage -a <bedfile> -b <bamfile> and this will produce a coverage report for the regions defined in the bedfile.

4 extra columns will be added to the output bed file:

  • * The number of features in bamfile that overlapped (by at least one base pair) the bedfile.
  • * The number of bases in the bedfile that had non-zero coverage from features in the bamfile.
  • * The length of the entry in the bedfile.
  • * The fraction of bases in the bedfile that had non-zero coverage from features in the bamfile.

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • bamfile — The bam file
  • bedfile — The bed file defining regions to calculate coverage for
Output
  • outfile — The output coverage report file
Envs
  • — Other arguments passed to the bedtools coverage toolSee bedtools coverage command.
  • bedtools — Path to bedtools executable
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger
class

biopipen.ns.bam.BedtoolsCoverageBamSummary(*args, **kwds) → Proc

Bases
biopipen.core.proc.Proc pipen.proc.Proc

Coverage summary report by bedtools coverage for bam files

This should run after BedtoolsCoverageBam to summarize the coverage report.

Attributes
  • cache — Should we detect whether the jobs are cached?
  • desc — The description of the process. Will use the summary fromthe docstring by default.
  • dirsig — When checking the signature for caching, whether should we walkthrough the content of the directory? This is sometimes time-consuming if the directory is big.
  • envs — The arguments that are job-independent, useful for common optionsacross jobs.
  • envs_depth — How deep to update the envs when subclassed.
  • error_strategy — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • export — When True, the results will be exported to <pipeline.outdir>Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • forks — How many jobs to run simultaneously?
  • input — The keys for the input channel
  • input_data — The input data (will be computed for dependent processes)
  • lang — The language for the script to run. Should be the path to theinterpreter if lang is not in $PATH.
  • name — The name of the process. Will use the class name by default.
  • nexts — Computed from requires to build the process relationships
  • num_retries — How many times to retry to jobs once error occurs
  • order — The execution order for this process. The bigger the numberis, the later the process will be executed. Default: 0. Note that the dependent processes will always be executed first. This doesn't work for start processes either, whose orders are determined by Pipen.set_starts()
  • output — The output keys for the output channel(the data will be computed)
  • output_data — The output data (to pass to the next processes)
  • output_flatten — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • plugin_opts — Options for process-level plugins
  • requires — The dependency processes
  • scheduler — The scheduler to run the jobs
  • scheduler_opts — The options for the scheduler
  • script — The script template for the process
  • submission_batch — How many jobs to be submited simultaneously.The program entrance for some schedulers may take too much resources when submitting a job or checking the job status. So we may use a smaller number here to limit the simultaneous submissions.
  • template — Define the template engine to use.This could be either a template engine or a dict with key engine indicating the template engine and the rest the arguments passed to the constructor of the pipen.template.Template object. The template engine could be either the name of the engine, currently jinja2 and liquidpy are supported, or a subclass of pipen.template.Template. You can subclass pipen.template.Template to use your own template engine.
Input
  • covfiles — The coverage report files produced by BedtoolsCoverageBamThe metrics (last 4 columns) are named as -
    • - NFeatures: The number of features in bamfile that overlapped (by at least one base pair) the bedfile.
    • - NRegionBases: The number of bases in the bedfile that had non-zero coverage from features in the bamfile.
    • - RegionSize: The length of the entry in the bedfile.
    • - FracRegionBases: The fraction of bases in the bedfile that had non-zero coverage from features in the bamfile.
Output
  • outdir — The output directory containing summary report plots and files
Envs
  • — Other arguments passed to the plot function. See plotthis package in R.
  • cases — Plotting cases. A dict with keys as case names and values as the arguments for the plot function.The arguments will inherit from envs (except cases, groups and save_data), and can be overridden by the case arguments.
  • descr — The description of the plot, showing in the report.
  • devpars (ns) — The device parameters for the clustree plot.
    • - res (type=int): The resolution of the plots.
    • - height: The height of the plots.
    • - width: The width of the plots.
  • groups (type=json) — The groups to add to the data for summary and plotting.It should be a dict with keys as group names and values as lists of sample names. The sample names should match the bam file names (without path and extension). The .coverage in the bam file name will be removed when matching sample names.
  • more_formats (type=list) — The formats to save the plots other than png.
  • plot_type — The type of plot to generate. The supported plot types arefunctions from plotthis package in R.
  • save_code (flag) — Whether to save the code to reproduce the plot.
  • save_data (flag) — Whether to save and report the data used to generate the plot.
Classes
Methods
  • __init_subclass__() — Do the requirements inferring since we need them to build up theprocess relationship </>
  • from_proc(proc, name, desc, envs, envs_depth, cache, export, output_flatten, error_strategy, num_retries, forks, input_data, order, plugin_opts, requires, scheduler, scheduler_opts, submission_batch) (Type) — Create a subclass of Proc using another Proc subclass or Proc itself</>
  • gc() — GC process for the process to save memory after it's done</>
  • log(level, msg, *args, logger) — Log message for the process</>
  • run() — Init all other properties and jobs</>
class

pipen.proc.ProcMeta(name, bases, namespace, **kwargs)

Bases
abc.ABCMeta

Meta class for Proc

Methods
  • __call__(cls, *args, **kwds) (Proc) — Make sure Proc subclasses are singletons</>
  • __repr__(cls) (str) — Representation for the Proc subclasses</>
staticmethod
__repr__(cls) → str

Representation for the Proc subclasses

staticmethod
__call__(cls, *args, **kwds)

Make sure Proc subclasses are singletons

Parameters
  • *args (Any) — and
  • **kwds (Any) — Arguments for the constructor
Returns (Proc)

The Proc instance

classmethod

from_proc(proc, name=None, desc=None, envs=None, envs_depth=None, cache=None, export=None, output_flatten=None, error_strategy=None, num_retries=None, forks=None, input_data=None, order=None, plugin_opts=None, requires=None, scheduler=None, scheduler_opts=None, submission_batch=None)

Create a subclass of Proc using another Proc subclass or Proc itself

Parameters
  • proc (Type) — The Proc subclass
  • name (str | none, optional) — The new name of the process
  • desc (str | none, optional) — The new description of the process
  • envs (Optional, optional) — The arguments of the process, will overwrite parent oneThe items that are specified will be inherited
  • envs_depth (int | none, optional) — How deep to update the envs when subclassed.
  • cache (bool | none, optional) — Whether we should check the cache for the jobs
  • export (bool | none, optional) — When True, the results will be exported to<pipeline.outdir> Defaults to None, meaning only end processes will export. You can set it to True/False to enable or disable exporting for processes
  • output_flatten (bool | none, optional) — Whether to flatten the output when saving to the outputdirectory. Normally, the output will be saved in a subdirectory named after the job index (e.g. <outdir>/0, <outdir>/1, etc.). If output_flatten is True, the output will be saved directly in the output directory without the subdirectories. This is useful when you want the job outputs to be directly revealed in the output directory. Note that this only works for processes with export=True or end processes and make sure the name of the output files won't conflict for jobs with each other when flattening. It takes 3 possible values
    • - None (default): flatten the output for single-job processes only
    • - True: flatten the output for all processes
    • - False: never flatten the output
  • error_strategy (str | none, optional) — How to deal with the errors
    • - retry, ignore, halt
    • - halt to halt the whole pipeline, no submitting new jobs
    • - terminate to just terminate the job itself
  • num_retries (int | none, optional) — How many times to retry to jobs once error occurs
  • forks (int | none, optional) — New forks for the new process
  • input_data (any | none, optional) — The input data for the process. Only when this processis a start process
  • order (int | none, optional) — The order to execute the new process
  • plugin_opts (Optional, optional) — The new plugin options, unspecified items will beinherited.
  • requires (Optional, optional) — The required processes for the new process
  • scheduler (str | none, optional) — The new shedular to run the new process
  • scheduler_opts (Optional, optional) — The new scheduler options, unspecified items willbe inherited.
  • submission_batch (int | none, optional) — How many jobs to be submited simultaneously.
Returns (Type)

The new process class

classmethod

__init_subclass__()

Do the requirements inferring since we need them to build up theprocess relationship

method

run()

Init all other properties and jobs

method

gc()

GC process for the process to save memory after it's done

method

log(level, msg, *args, logger=<LoggerAdapter pipen.core (WARNING)>)

Log message for the process

Parameters
  • level (int | str) — The log level of the record
  • msg (str) — The message to log
  • *args — The arguments to format the message
  • logger (LoggerAdapter, optional) — The logging logger