Skip to contents

Initialize new MitoPilot Project

Usage

new_project(
  path = ".",
  mapping_fn = NULL,
  mapping_id = "ID",
  mapping_geome = "GEOME_BCID",
  fetch_geome = TRUE,
  mapping_gbif = "GBIF_ID",
  fetch_gbif = TRUE,
  mapping_ncbi = "BioSample",
  fetch_ncbi = TRUE,
  link_sources = FALSE,
  data_path = NULL,
  min_depth = 2e+06,
  genetic_code = NULL,
  executor = c("local", "awsbatch", "slurm", "sge", "pbs", "lsf", "NMNH_Hydra",
    "NOAA_SEDNA"),
  container = paste0("macguigand/mitopilot:", utils::packageVersion("MitoPilot")),
  custom_seeds_db = NULL,
  custom_labels_db = NULL,
  config = NULL,
  profile_dir = mitopilot_config_dir(),
  ncbi_api_key = NULL,
  Rproj = TRUE,
  force = FALSE,
  ...
)

Arguments

path

Path to the project directory (default = current working directory)

mapping_fn

Path to a mapping file. Should be a csv that minimally includes an `ID` column with a unique identifier for each sample, a `Taxon` column containing taxonomic information for each sample, and columns `R1` and `R2` specifying the names of the raw paired read inputs. May include additional columns with other sample metadata, and an optional Reference column naming a per-sample MapToRef reference (file path, URL, or NCBI accession). A FASTA reference also needs a Reference_topology column (circular or linear). Both values are stored on the sample and used when its parameter set assembles with MapToRef. Reference is a reserved column name: it is never stored as sample metadata, so rename the column if you use it for something else.

mapping_id

The name of the column in the mapping file that contains the unique sample identifiers (default = "ID").

mapping_geome

Name of the mapping-file column holding GEOME BCIDs (optional). Stored as `GEOME_BCID`. See `vignette("Specimen-Metadata")`. Passed to `new_db()`.

fetch_geome

Fetch GEOME metadata for samples with a BCID during setup (default TRUE). Set FALSE when offline and run [fetch_geome()] later. Passed to `new_db()`.

mapping_gbif

Name of the mapping-file column holding GBIF occurrence IDs (optional). Stored as `GBIF_ID`. See `vignette("Specimen-Metadata")`. Passed to `new_db()`.

fetch_gbif

Fetch GBIF metadata for samples with a GBIF ID during setup (default TRUE). Set FALSE when offline and run [fetch_gbif()] later. Passed to `new_db()`.

mapping_ncbi

Name of the mapping-file column holding NCBI BioSample or SRA accessions (optional). Must not be the sample ID column. Stored as `BioSample`. See `vignette("Specimen-Metadata")`. Passed to `new_db()`.

fetch_ncbi

Fetch NCBI metadata for samples with a BioSample value during setup (default TRUE). Set FALSE when offline and run [fetch_ncbi()] later. Passed to `new_db()`.

Follow links between GEOME, GBIF, and NCBI records to fill in IDs a sample does not have yet (default FALSE). Saved as the project setting. See `vignette("Specimen-Metadata")`. Passed to `new_db()`.

data_path

Path to the directory where the raw data is located. Can be a AWS s3 bucket even if not using AWS for pipeline execution..

min_depth

Minimum number of paired sequences after pre-processing to proceed with assembly (default: 2000000 reads)

genetic_code

Optional NCBI translation table override. Default `NULL` auto-selects the genetic code from each sample's curation ruleset. Supplying a number sets a project-wide override on the default curation options. See https://www.ncbi.nlm.nih.gov/Taxonomy/Utils/wprintgc.cgi

executor

The executor to use for running the nextflow pipeline. May be a built-in template ("local" (default), "awsbatch", "slurm", "sge", "pbs", "lsf", "NMNH_Hydra", "NOAA_SEDNA") or the name of a saved cluster profile created with [generate_config()]. See [list_configs()] for available names.

container

The docker container to use for pipeline execution.

custom_seeds_db

Full path to custom seeds database for GetOrganelle

custom_labels_db

Full path to custom labels database for GetOrganelle

config

(optional) provide a path to an existing custom nextflow config file. If not provided a config file template will be created based on the specified executor.

profile_dir

Directory searched for saved cluster profiles when resolving `executor` (default [mitopilot_config_dir()]).

ncbi_api_key

Optional NCBI API key string. Used to raise NCBI request rate limits for the remote BLAST + GenBank fetch steps. See <https://www.ncbi.nlm.nih.gov/datasets/docs/v2/api/api-keys/>. May be left empty and edited later in `.config` (`params.ncbi_api_key`).

Rproj

(logical) Initialize and open an RStudio project in the project directory (default = TRUE). This option has no effect if not running interactively in RStudio.

force

(logical) Force recreating of existing project database and config files (default = FALSE).

...

Additional arguments passed as default processing parameters to `new_db()`