SIDE
1.1 - ABOUT THE GPRO SERVER SIDE
The GPRO Suite is a Client Side & Server Side solution where applications like RNASeq, VariantSeq, DeNovoSeq and STATools (i.e. the Client Side apps) provide tailor made Graphical User Interfaces (GUI) to manage a series of workflows and pipelines installed in a bioinformatic server infrastructure we call the Server Side.
The GPRO Server Side has the following requirements and dependencies:
- Linux Operating System with minimum Bash version 4.
- Mysql Server for databases
- Apache HTTP server 2.2 or later
- PHP 7 or later
- R and Bioconductor
- Perl 5
- Python 2.7
- Pipeline Software dependencies (shown in Table 1)
- Internal GPRO Databases
- A management system of scripts to handle analytical incoming request and manage users accounts as well as other tools for tracking of computational jobs and user support
- The API allowing communication between the server side and client side applications.
| Pipeline Step | Tools | RNASeq | DeNovoSeq | VariantSeq | License |
|---|---|---|---|---|---|
Quality analysis and Preprocessing |
FastQC [v0.11.5] (Andrews 2016) |
✓ | ✓ | ✓ | GPLv3 |
|
FastqMidCleaner [1.0.0] |
✓ | ✓ | ✓ | GPLv3 | |
|
Cutadapt [1.18] (Martin 2011) |
✓ | ✓ | ✓ | MIT | |
|
Prinseq [PRINSEQ-lite 0.20.4] (Schmieder and Edwards 2011) |
✓ | ✓ | ✓ | GPLv3 | |
|
Trimmomatic [0.36] (Bolger, et al. 2014) |
✓ | ✓ | ✓ | GPLv3 | |
|
FastxToolkit [0.0.13] (Hannon Lab 2016) |
✓ | ✓ | ✓ | AGPLv3 | |
|
CANU (Koren, et al. 2017) |
No | ✓ | No | GPLv3 | |
|
FastqCollapser [1.0.0] |
✓ | No | ✓ | GPLv3 | |
|
FastqIntersect [1.0.0] |
✓ | No | ✓ | GPLv3 | |
Mapping on reference genome or transcriptome |
TopHat [v2.1.1] (Kim et al. 2013) |
✓ | No | ✓ | Boost Software 1.0 |
|
Hisat2 [2.2.1] (Kim et al. 2015) |
✓ | No | ✓ | GPLv3 | |
|
Bowtie2 [2.2.9] (Langmead and Salzberg 2012) |
✓ | No | ✓ | GPLv3 | |
|
BWA [0.7.15-r1140] (Li and Durbin 2009) |
✓ | No | ✓ | GPLv3 | |
|
STAR [2.7.0f] (Dobin et al. 2013) |
No | No | ✓ | MIT | |
Quantification |
Corset [1.06] (Davidson and Oshlack 2017) |
✓ | No | No | BSD 2-Clause |
|
Htseq [0.12.4] (Anders 2015) |
✓ | No | No | GPLv3 | |
Post Processing |
Bed Tools [v2.29.2] (Quinlan and Hall 2010) |
No | No | ✓ | GPLv2 |
|
GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011) |
No | No | ✓ | Apache-2.0 | |
|
Picard [2.19.0] (Wysoker et al. 2011) |
No | No | ✓ | MIT | |
|
SAMtools [1.8] (Li et al. 2009) |
No | No | ✓ | MIT | |
Transcriptome assembly |
Cufflinks [v2.2.1](Trapnell et al. 2012) |
✓ | No | No | MIT |
|
Oases (Schulz et al. 2012) |
No | ✓ | No | GPLv3 | |
|
SOAPdenovo-trans (Xie et al. 2014) |
No | ✓ | No | GPLv3 | |
Genome assembly |
Velvet (Zerbino and Birney 2008) |
No | ✓ | No | GPLv3 |
|
SOAPdenovo2 (Luo et al. 2012) |
No | ✓ | No | GPLv3 | |
|
CANU (Koren et al. 2017) |
No | ✓ | No | GPLv2 | |
|
SPAdes (Bankevich et al. 2012) |
No | ✓ | No | GPLv3 | |
Gap filling and scaffolding |
Gap closer (Luo et al. 2012) |
No | ✓ | No | GPLv3 |
|
BESST (Sahlin et al. 2014) |
No | ✓ | No | GPLv3 | |
|
OPERA (Gao et al. 2011) |
No | ✓ | No | MIT | |
Differential expression |
DESeq [2.1.28] (Love et al. 2014) |
✓ | No | No | GPLv3 |
|
EdgeR [3.30.3] (Robinson et al. 2010) |
✓ | No | No | GPLv2 | |
|
Cuffdiff [v2.2.1] (Trapnell et al. 2012) |
✓ | No | No | Boost Software 1.0 | |
|
CummeRbund [2.30.0] (Goff et al. 2013) |
✓ | No | No | Artistic-2.0 | |
Enrichment Analysis |
GOseq [1.40.0] (Young et al. 2010) |
✓ | No | No | LGPLv3 |
Gene Prediction |
Augustus (Stanke et al. 2008) |
No | ✓ | No | Artistic-1.0 |
Gene Annotation |
BLAST (Altschul et al. 1990) |
No | ✓ | No | Public Domain |
|
HMMER (Mistry et al. 2013) |
No | ✓ | No | EASEL | |
|
Variant Effect Predictor [105.0] (McLaren et al. 2016) |
No | No | ✓ | Apache-2.0 | |
Training Sets |
GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011) |
No | No | ✓ | Apache-2.0 |
Variant Calling |
GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011) |
No | No | ✓ | Apache-2.0 |
|
VarScan2 [v2.4.3] (Koboldt et al. 2012) |
No | No | ✓ | VarScan | |
Variant filtering |
GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011) |
No | No | ✓ | Apache-2.0 |
1.2 - SERVER SIDE DOCKER
The GPRO Server Side can be installed in remote servers or in the user PC provided that if it has enough disk space and RAM (at least 500 Gb of hard disk and 16Gb of RAM). Installation of the GPRO Server Side involves complex steps to setup Linux, Apache, MySQL, and PHP (LAMP stack) as well as manual installation of the distinct third-party CLI software shown in Table 1 as well as GPRO databases, and multiple distinct scripts needed to handle requests of client applications to the Server Side. However we have deployed it into a docker container(Merkel 2014) to facilitate its installation. An image of this docker can be downloaded from: https://hub.docker.com/r/biotechvana/gpro. The current version of this docker is limited to one or two users, but we are committed to release a forthcoming version for multiple users.
To install the GPRO Server Side docker, please proceed as follows
First, download the Docker Desktop software from https://www.docker.com/products/docker-desktop and install it.
If you want to install the Server Side docker with default user name and password, please run the following command on the terminal of the Docker Desktop:
|
local_path="/path/to/local_home" > docker run -d -p 80:80 -p 20-22:20-22 -p 65500-65515:65500-65515 -v /path/to/local_home:/home/gpro_user biotechvana/gpro |
In doing so, the Server Side will consider the word g_user as your default user name and password. In other words, your user name will be g_user and your password will be g_user, too.
Otherwise, if you are interested in running the docker with your own user name and password, please run the following command on the Docker Desktop terminal:
|
local_path="/path/to/local_home" GPRO_USER="myUserName" GPRO_USER_PASS="myUserNamePass" > docker run -d -p 80:80 -p 20-22:20-22 -p 65500-65515:65500-65515 -v /path/to/local_home:/home/gpro_user biotechvana/gpro |
For example, if you chose DirkGently as user name and HolisticDetective as password then you will name to run the command as follows:
|
local_path="/path/to/local_home" GPRO_USER="DirkGently" GPRO_USER_PASS="HolisticDetective" > docker run -d -p 80:80 -p 20-22:20-22 -p 65500-65515:65500-65515 -v /path/to/local_home:/home/gpro_user biotechvana/gpro |
Notes
All third-party command line interface software are integrated in the server-side docker image excepting Varscan2. This is because the license for VarScan2 usage distinguishes between academic users and commercial users. Academic users can use VarScan2 without restrictions for academic purposes, while commercial users need to contact the VarScan2 authors to get the corresponding commercial license (for more details see https://github.com/dkoboldt/varscan/releases ). Taking this into consideration, the server-side image does not include VarScan2 and this tool must be integrated by the own user into the running container. The rest of CLI software dependencies used by “RNAseq” and “VariantSeq” are already installed in the docker image and have licenses of use allowing unrestricted use for every kind of author (if they are appropriately cited and accredited). To integrate VarScan2 in the docker please proceed as follows.
To manage the image you need to get image name first if you did not set it in the run command, you can get the image name by running docker ps
$ docker ps
################## Sample output
# CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
# 5121b762b39f d82c97fcc2c7 "/gpro_init" 2 hours ago Up 2 hours 0.0.0.0:20-22->20-22/tcp, :::20-22->20-22/tcp, 0.0.0.0:80->80/tcp, :::80->80/tcp,
0.0.0.0:65500-65515->65500-65515/tcp, :::65500-65515->65500-65515/tcp hungry_boyd
Here the name is hungry_boyd
Adding VarScan to the gpro server : run the following command:
$ docker exec hungry_boyd gpro install varscan
Commercial users should proceed in the same way but must contact first the VarScan2 authors to get the license if they want to comply with the VarScan2 terms of use.
1.3 - LINKING THE CLIENT SIDE WITH THE SERVER SIDE
After installed and running the server side image in the Docker Desktop you must link your GPRO application of interest to the Server Side in order to run the server side analyses, pipelines or workflows. As previously said, the Client Side applications of the Suite that are dependent of the Server Side are RNAseq , VariantSeq , DeNovoSeq (manual in preparation) and STATools (manual in preparation). Please visit their respective manuals for detailed instructions about how to link each application with the Server Side and automatically run the docker desktop each time you open the application linked to the Server Side
1.4 - ACKNOWLEDGEMENTS
This work was supported by the European Union’s Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement No 642095 for the OPATHY consortium, by the pre-doctoral research fellowship from Industrial Doctorates of MINECO (Grant 659 DI-17-09134); by the State Plan for Scientific and Technical Research and Innovation 2017-2020 under the Grant TSI-100903-2019-11 from the Secretary of State for Digital Advancement from Ministry of Economic Affairs and Digital Transformation, Spain; and by Expedient IDI-2021-158274-a from Ministry of Science and Innovation, Spain
1.5 - LITERATURE CITED
- Altschul SF, Gish W, Miller W, Myers EW and Lipman DJ. 1990. Basic local alignment search tool. JMolBiol 215: 403-410. doi:10.1016/S0022-2836(05)80360-2.
- Anders S, Pyl PT, Huber W. 2015. HTSeq — A Python framework to work with high-throughput sequencing data. Bioinformatics. 15;31(2):166-9. doi: 10.1093/bioinformatics/btu638.
- Andrews S. 2016. “FastQC A Quality Control Tool for High Throughput Sequence Data.” http://www.bioinformatics.babraham.ac.uk/projects/fastqc/.
- Bankevich A, Nurk S, Antipov D, et al. 2012. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. Journal of computational biology: 19: 455-477.doi: 0.1089/cmb.2012.0021 .
- Bolger AM., Lohse M, & Usadel B. 2014. Trimmomatic: A flexible trimmer for Illumina Sequence Data. Bioinformatics, btu170.doi: 10.1093/bioinformatics/btu170.
- Cibulskis K, Lawrence MS, Carter SL, Sivachenko A, Jaffe D, Sougnez C, Gabriel S, Meyerson M, Lander ES, Getz G. 2013. Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat Biotechnol, 31:213-219.doi: 10.1038/nbt.2514.
- Davidson NM, and Oshlack A. 2017. Corset: Enabling Differential Gene Expression Analysis for de Novo Assembled Transcriptomes. Accessed June 28. doi:10.1186/s13059-014-0410-6.
- DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR et al. 2011. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature Genetics 43: 491-498. doi:10.1038/ng.806.doi:10.1038/ng.806.
- Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Bttut P, Chaisson M, Gingeras, TR. 2013. STAR: ultrafast universal RNA-seq aligner. Bioinformatics (Oxford, England), 29(1), 15–21doi:10.1093/bioinformatics/bts635.
- Gao S, Sung WK, and Nagarajan N. 2011. Opera: reconstructing optimal genomic scaffolds with high-throughput paired-end sequences. Journal of computational biology. 18: 1681-1691. doi: 10.1089/cmb.2011.0170
- Goff L, Trapnell C, and Kelley D. 2013. cummeRbund: Analysis, exploration, manipulation, and visualization of Cufflinks high-throughput sequencing data. R package version 2.20.0.
- Hannon Lab. 2016. FASTX Toolkit. http://hannonlab.cshl.edu/fastx_toolkit/.
- Kim D, Pertea G, Trapnell C, Pimentel H, Kelley R and Salzberg SL. 2013. “TopHat2: Accurate Alignment of Transcriptomes in the Presence of Insertions, Deletions and Gene Fusions.” Genome Biology 14 (4): R36. doi:10.1186/gb-2013-14-4-r36.
- Kim D, Langmead B, Salzberg SL. 2015. “HISAT: A Fast Spliced Aligner with Low Memory Requirements.” Nature Methods 12 (4). Nature Research: 357–60. doi:10.1038/nmeth.3317.
- Koboldt DC, Zhang Q, Larson DE, Shen D, McLellan MD, Lin L, Miller CA, Mardis ER, Ding L, Wilson RK. 2012. VarScan 2: Somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Research doi: 10.1101/gr.129684.111.
- Koren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM. 2017. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 27, 722–736.doi: 10.1101/gr.215087.116
- Langmead B, and Salzberg SL. 2012. “Fast Gapped-Read Alignment with Bowtie 2.” Article. Nat Methods 9. doi:10.1038/nmeth.1923.
- Li H, Handsaker B, Wysoker A, Fennell T, Ruan J et al. 2009. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25: 2078-2079. doi:10.1093/bioinformatics/btp352
- Li H, and Durbin R. 2009. “Fast and Accurate Long-Read Alignment with Burrows-Wheeler Transform.” Bioinformatics 26 (5): 589–95. doi:10.1093/bioinformatics/btp698
- Love MI, Huber W and Anders S. 2014. “Moderated Estimation of Fold Change and Dispersion for RNA-Seq Data with DESeq2.” Genome Biology 15. doi:10.1186/s13059-014-0550-8.
- Luo R, Liu B, Xie Y, et al. 2012. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1: 18.doi:10.1186/2047-217X-1-18
- Martin M. 2011. Cutadapt Removes Adapter Sequences from High-Throughput Sequencing Reads.EMBnet.journal 17 (1): 10doi:10.14806/ej.17.1.200.
- McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K et al. 2010. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Research 20: 1297-1303. doi:10.1101/gr.107524.110.
- McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, Flicek P and Cunningham F. 2016. The Ensembl Variant Effect Predictor. Genome Biology. 17-122 doi.org/10.1186/s13059-016-0974-4.
- Merkel D. 2014. Docker: lightweight Linux containers for consistent development and deployment. Linux Journal 2014:2 Merkel et al 2014:2.
- Mistry J, Finn RD, Eddy SR, Bateman A, Punta M. 2013. Challenges in homology search: HMMER3 and convergent evolution of coiled-coil regions. Nucleic Acids Res. 2013;41(12):e121. doi:10.1093/nar/gkt263.
- Quinlan AR, and Hall IM. 2010. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 26(6): 841–842. doi:10.1093/bioinformatics/btq033.
- Robinson MD, McCarthy DJ, and Smyth GK. 2010. “edgeR: A Bioconductor Package for Differential Expression Analysis of Digital Gene Expression Data.” Bioinformatics (Oxford, England) 26 (1). Oxford University Press: 139–40. doi:10.1093/bioinformatics/btp616.
- Sahlin K, Vezzi F, Nystedt B, Lundeberg J, and Arvestad L. 2014. BESST - Efficient scaffolding of large fragmented assemblies. BMC Bioinformatics 15:281.doi: 10.1186/1471-2105-15-281.
- Stanke M, Diekhans M, Baertsch R, and Haussler D. 2008. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24: 637-644.doi: 10.1093/bioinformatics/btn013.
- Schmieder R and Edwards R. 2011. Quality Control and Preprocessing of Metagenomic Datasets. Bioinformatics 27 (6): 863–64. doi:10.1093/bioinformatics/btr026.
- Schulz MH, Zerbino DR, Vingron M and Birney E. 2012. Oases: robust de novo RNA-seq assembly across the dynamic range of expression levels. Bioinformatics 28: 1086-1092. doi: 10.1093/bioinformatics/bts094.
- Trapnell C, Roberts A, Goff L, Pertea G, Kim D, Kelley DR, Pimentel H, Salzberg SL, Rinn JL and Pachter L. 2012. Differential Gene and Transcript Expression Analysis of RNA-Seq Experiments with TopHat and Cufflinks. Nature Protocols 7 (3). Nature Research: 562–78. doi:10.1038/nprot.2012.016.
- Wysoker A, Tibbetts K, Fennell T. 2011. PicardTools 1.5.3. Available at http://sourceforge.net/projects/picard/files/picard-tools/.
- Xie Y, Wu G, Tang J, et al. 2014. SOAPdenovo-Trans: de novo transcriptome assembly with short RNA-Seq reads. Bioinformatics 30: 1660-1666. doi: 10.1093/bioinformatics/btu077.
- Young MD, Wakefield MJ, Smyth GK, and Oshlack A. 2010. Gene Ontology Analysis for RNA-Seq: Accounting for Selection Bias. Genome Biology 11. http://genomebiology.com/2010/11/2/R14.
- Zerbino DR, and Birney E. 2008. Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Genome Res 18: 821-829.doi: 10.1101/gr.074492.107.