Bacterial Orthology Finding And Syntenic Analysis (bofasa)
bofasa is specifically designed for investigating orthology between multiple-species of bacteria. It prioritizes high-quality orthology inference at the expense of throughput, designed to run on between 4 and 200 genomes.
For single species analyses, we recommend Panaroo or PPanGGOLiN, which for such cases, offers considerable advantages in terms of speed, accuracy, and scalability.
Comparable and good alternatives to bofasa include PIRATE and SCARAP. PIRATE performs similarly to bofasa when applied to genomes representative of different species from a single genus when using default parameters. SCARAP behaves more similar to non-bacteria specific multi-species orthology inference software such as OrthoFinder and SonicParanoid and is really fast and could be well-suited for more comprehensive and large-scale analysis.
conda create -c conda-forge -c bioconda -p /path/to/bofasa_env/ bofasa
conda activate /path/to/bofasa_env/
export PATH=$CONDA_PREFIX/bin:$PATH
bofasa setup --force
bofasa -hpixi init /path/to/bofasa_env/
pixi add -m /path/to/bofasa_env/ bofasa
pixi shell -m /path/to/bofasa_env/
export PATH=$CONDA_PREFIX/bin:$PATH
bofasa setup --force
bofasa -hCheck out this Wiki page for information on how to install/run bofasa via Docker.
# get test dataset and testing script from bofasa Github repo:
wget https://github.com/raufs/bofasa/raw/refs/heads/main/test_case.tar.gz
wget https://raw.githubusercontent.com/raufs/bofasa/refs/heads/main/run_tests.sh
bash run_tests.shNote
bofasa setup should already have been run. Also, the above will work for conda, but dataset can be downloaded and adapted to test docker installation too!
bofasa provides a unified command-line interface with two main subcommands: bofasa prep and bofasa run. For a more detailed look at the algorithms behind bofasa, check out the ALGORITHM.md document or this Wiki page.
bofasa setupThe geNomad database is downloaded by default. It is large and only needed for bofasa run --run-genomad, so it can be skipped:
bofasa setup --skip-genomadNote
To aid reproducability between two runs, please ensure you are using the same version of the Pfam-A database!
bofasa prep -i genome1.fasta genome2.fasta genome3.fasta genome4.fasta -o prep_output/bofasa run -i prep_output/ -o analysis_output/ -c 8bofasa provides a single entry point with intuitive subcommands:
██████╗ ██████╗ ███████╗ █████╗ ███████╗ █████╗
██╔══██╗██╔═══██╗██╔════╝██╔══██╗██╔════╝██╔══██╗
██████╔╝██║ ██║█████╗ ███████║███████╗███████║
██╔══██╗██║ ██║██╔══╝ ██╔══██║╚════██║██╔══██║
██████╔╝╚██████╔╝██║ ██║ ██║███████║██║ ██║
╚═════╝ ╚═════╝ ╚═╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝
BOFASA: Bacterial Ortholog Finder and Synteny Analysis
A comprehensive tool for identifying ortholog groups and analyzing
syntenic conservation in bacterial genomes representing multiple
species.
Commands:
setup Set up annotation databases
prep Prepare genomic data for analysis
run Run the main BOFASA analysis pipeline
For detailed help on any command, use: bofasa <command> --help
Examples:
bofasa prep -i genome1.fna genome2.fna -o prep_output/
bofasa run -i prep_output/ -o analysis_results/ -c 8
For more information, visit: https://github.com/raufs/bofasa
Note
The concepts in bofasa were developed over many years and the initial release (v1.1.0) developed without the use of AI. For version v1.2.0, we used AI to largely restructure the code and in the process caught some bugs and added a couple optimizations/bells and whistles. Moving forward, we will likely use some AI because its engraved into search engines but will largely try to refrain from relying on it to have a sense of the code so that it is not yet another black box.
Processes input genomes (FASTA or GenBank files) for analysis. Can also take in Prokka and bakta directories!
bofasa prep -i <genome1.fasta> <genome2.gbk> ... -o <output-dir> [OPTIONS]Required Arguments:
-i, --input-genomes: Input genomes (FASTA or GenBank files)-a, --annotation-dirs: Annotation directories (Prokka or Bakta output directories). Required if no input genomes are provided.-o, --output-dir: Output directory for prepared data
Optional Arguments:
-c, --threads: Number of threads (default: 4)-gcm, --gene-calling-method: Gene calling method (pyrodigal/prodigal, default: pyrodigal)-l, --locus-tag-length: Length of locus tags (default: 3)-m, --meta-mode: Use meta-mode for gene calling-ml, --min-length: Minimum length of domain/inter-domain unit to consider (default: 20)-rlt, --rename-locus-tags: Rename locus tags in GenBank files and annotation directories (Prokka/Bakta)-rg, --run-genomad: Run genomad for phage/plasmid annotation-emg, --extract-mge-genomes: Extract mobile genetic element (phage/plasmid) genomes identified by genomad (requires -rg)-sds, --skip-domain-splitting: Skip domain-annotation based splitting of coding sequences-mm, --max-memory: Memory limit in GB (default: 32)x, --ignore-upperbound-limit: Ignore the upper bound limit for the number of genomes to process (default: 200)-y, --auto: Automatically setyfor all interactive prompts
Performs the main orthology analysis on prepared data.
bofasa run -i <bofasa-prep-dir> -o <output-dir> [OPTIONS]Required Arguments:
-i, --bofasa-prep-dir: Input directory produced by bofasa prep-o, --output-dir: Output directory
Optional Arguments:
-sr, --surrounding-bp: Base pairs for syntenic analysis (default: 10000)-ogc, --og-consensus: Determine consensus sequences-cg, --core-genome: Construct core genome alignment-dj, --dog-jaccard: Jaccard index threshold (default: 0.5)-fic, --fixation-index-cutoff: Fixation index cutoff (default: 0.1)-smb, --skip-merge-back: Skip merge back assessment-spr, --skip-phylo-refine: Skip phylogenetic refinement-rs, --rooting-seeds: Rooting seeds (default: 1)-us, --ultra-sens: Use ultra-sensitivity mode-mi, --mcl-inflation: MCL inflation parameter (default: 1.2)-ns, --near-scc-prop: Near SCC proportion (default: 0.95)-c, --threads: Number of threads (default: 4)-mrd, --max-recursion-depth: Max recursion depth (default: 5000)-mm, --max-memory: Memory limit in GB (default: 32)-y, --auto: Automatically setyfor all interactive prompts
bofasa also provides several additional utility scripts:
extract_og_pairs.py- Extract ortholog group pairs for method comparisonextract_scc_og_pairs.py- Extract single-copy-core ortholog group pairscompare_og_pairs.py- Compare ortholog group pairs between methodsdetermine_og_freqs.py- Determine ortholog group frequenciesprint_og_itol_matrix.py- Print ortholog group matrices for iTOL visualizationvisualize_og_context.py- Visualize the context of focal ortholog groups of interest. Is still experimental!
Note: GenBank processing and Prodigal gene calling functionality is now integrated into the main bofasa prep command and no longer requires separate scripts.
High-quality and automated inference of ortholog groups across multiple bacterial species using bofasa. Rauf Salamzade, Eitan Yaffe, Aamuktha Kottapalli, Lindsay Kalan, David Relman, 2026.
Please also consider citing both OrthoFinder 2 and geNomad which are used for determination of coarse domain ortholog groups and the annotation of phages/plasmids, respectively.
BSD 3-Clause License
Copyright (c) 2024, Rauf Salamzade
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice, this
list of conditions and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the copyright holder nor the names of its
contributors may be used to endorse or promote products derived from
this software without specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
