English
Español
Valencià
GPRO SERVER
Manual

SIDE

1.1 - ABOUT THE GPRO SERVER SIDE

The GPRO Suite is a Client Side & Server Side solution where applications like RNASeq, VariantSeq, DeNovoSeq and STATools (i.e. the Client Side apps) provide tailor made Graphical User Interfaces (GUI) to manage a series of workflows and pipelines installed in a bioinformatic server infrastructure we call the Server Side.

The GPRO Server Side has the following requirements and dependencies:

  • Linux Operating System with minimum Bash version 4.
  • Mysql Server for databases
  • Apache HTTP server 2.2 or later
  • PHP 7 or later
  • R and Bioconductor
  • Perl 5
  • Python 2.7
  • Pipeline Software dependencies (shown in Table 1)
  • Internal GPRO Databases
  • A management system of scripts to handle analytical incoming request and manage users accounts as well as other tools for tracking of computational jobs and user support
  • The API allowing communication between the server side and client side applications.
The table is structured according to the steps that are part of the different pipelines. If a CLI software is included in an application, it will show a ✓ in its corresponding column or a X if it is not included.
Table 1.- Pipeline dependencies hosted at the GPRO Server Side
Pipeline Step Tools RNASeq DeNovoSeq VariantSeq License

Quality analysis and Preprocessing

FastQC [v0.11.5] (Andrews 2016)

✓ ✓ ✓ GPLv3

FastqMidCleaner [1.0.0]

✓ ✓ ✓ GPLv3

Cutadapt [1.18] (Martin 2011)

✓ ✓ ✓ MIT

Prinseq [PRINSEQ-lite 0.20.4] (Schmieder and Edwards 2011)

✓ ✓ ✓ GPLv3

Trimmomatic [0.36] (Bolger, et al. 2014)

✓ ✓ ✓ GPLv3

FastxToolkit [0.0.13] (Hannon Lab 2016)

✓ ✓ ✓ AGPLv3

CANU (Koren, et al. 2017)

No ✓ No GPLv3

FastqCollapser [1.0.0]

✓ No ✓ GPLv3

FastqIntersect [1.0.0]

✓ No ✓ GPLv3

Mapping on reference genome or transcriptome

TopHat [v2.1.1] (Kim et al. 2013)

✓ No ✓ Boost Software 1.0

Hisat2 [2.2.1] (Kim et al. 2015)

✓ No ✓ GPLv3

Bowtie2 [2.2.9] (Langmead and Salzberg 2012)

✓ No ✓ GPLv3

BWA [0.7.15-r1140] (Li and Durbin 2009)

✓ No ✓ GPLv3

STAR [2.7.0f] (Dobin et al. 2013)

No No ✓ MIT

Quantification

Corset [1.06] (Davidson and Oshlack 2017)

✓ No No BSD 2-Clause

Htseq [0.12.4] (Anders 2015)

✓ No No GPLv3

Post Processing

Bed Tools [v2.29.2] (Quinlan and Hall 2010)

No No ✓ GPLv2

GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011)

No No ✓ Apache-2.0

Picard [2.19.0] (Wysoker et al. 2011)

No No ✓ MIT

SAMtools [1.8] (Li et al. 2009)

No No ✓ MIT

Transcriptome assembly

Cufflinks [v2.2.1](Trapnell et al. 2012)

✓ No No MIT

Oases (Schulz et al. 2012)

No ✓ No GPLv3

SOAPdenovo-trans (Xie et al. 2014)

No ✓ No GPLv3

Genome assembly

Velvet (Zerbino and Birney 2008)

No ✓ No GPLv3

SOAPdenovo2 (Luo et al. 2012)

No ✓ No GPLv3

CANU (Koren et al. 2017)

No ✓ No GPLv2

SPAdes (Bankevich et al. 2012)

No ✓ No GPLv3

Gap filling and scaffolding

Gap closer (Luo et al. 2012)

No ✓ No GPLv3

BESST (Sahlin et al. 2014)

No ✓ No GPLv3

OPERA (Gao et al. 2011)

No ✓ No MIT

Differential expression

DESeq [2.1.28] (Love et al. 2014)

✓ No No GPLv3

EdgeR [3.30.3] (Robinson et al. 2010)

✓ No No GPLv2

Cuffdiff [v2.2.1] (Trapnell et al. 2012)

✓ No No Boost Software 1.0

CummeRbund [2.30.0] (Goff et al. 2013)

✓ No No Artistic-2.0

Enrichment Analysis

GOseq [1.40.0] (Young et al. 2010)

✓ No No LGPLv3

Gene Prediction

Augustus (Stanke et al. 2008)

No ✓ No Artistic-1.0

Gene Annotation

BLAST (Altschul et al. 1990)

No ✓ No Public Domain

HMMER (Mistry et al. 2013)

No ✓ No EASEL

Variant Effect Predictor [105.0] (McLaren et al. 2016)

No No ✓ Apache-2.0

Training Sets

GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011)

No No ✓ Apache-2.0

Variant Calling

GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011)

No No ✓ Apache-2.0

VarScan2 [v2.4.3] (Koboldt et al. 2012)

No No ✓ VarScan

Variant filtering

GATK [v4.1.2.0] (MacKena et al. 2010;Cibulskis et al. 2013;DePristo et al. 2011)

No No ✓ Apache-2.0

1.2 - SERVER SIDE DOCKER

The GPRO Server Side can be installed in remote servers or in the user PC provided that if it has enough disk space and RAM (at least 500 Gb of hard disk and 16Gb of RAM). Installation of the GPRO Server Side involves complex steps to setup Linux, Apache, MySQL, and PHP (LAMP stack) as well as manual installation of the distinct third-party CLI software shown in Table 1 as well as GPRO databases, and multiple distinct scripts needed to handle requests of client applications to the Server Side. However we have deployed it into a docker container(Merkel 2014) to facilitate its installation. An image of this docker can be downloaded from: https://hub.docker.com/r/biotechvana/gpro. The current version of this docker is limited to one or two users, but we are committed to release a forthcoming version for multiple users.

To install the GPRO Server Side docker, please proceed as follows

First, download the Docker Desktop software from https://www.docker.com/products/docker-desktop and install it.

If you want to install the Server Side docker with default user name and password, please run the following command on the terminal of the Docker Desktop:

local_path="/path/to/local_home"

> docker run -d -p 80:80 -p 20-22:20-22 -p 65500-65515:65500-65515 -v /path/to/local_home:/home/gpro_user biotechvana/gpro

In doing so, the Server Side will consider the word g_user as your default user name and password. In other words, your user name will be g_user and your password will be g_user, too.

Otherwise, if you are interested in running the docker with your own user name and password, please run the following command on the Docker Desktop terminal:

local_path="/path/to/local_home"

GPRO_USER="myUserName"

GPRO_USER_PASS="myUserNamePass"

> docker run -d -p 80:80 -p 20-22:20-22 -p 65500-65515:65500-65515 -v /path/to/local_home:/home/gpro_user biotechvana/gpro

For example, if you chose DirkGently as user name and HolisticDetective as password then you will name to run the command as follows:

local_path="/path/to/local_home"

GPRO_USER="DirkGently"

GPRO_USER_PASS="HolisticDetective"

> docker run -d -p 80:80 -p 20-22:20-22 -p 65500-65515:65500-65515 -v /path/to/local_home:/home/gpro_user biotechvana/gpro

Notes

All third-party command line interface software are integrated in the server-side docker image excepting Varscan2. This is because the license for VarScan2 usage distinguishes between academic users and commercial users. Academic users can use VarScan2 without restrictions for academic purposes, while commercial users need to contact the VarScan2 authors to get the corresponding commercial license (for more details see https://github.com/dkoboldt/varscan/releases ). Taking this into consideration, the server-side image does not include VarScan2 and this tool must be integrated by the own user into the running container. The rest of CLI software dependencies used by “RNAseq” and “VariantSeq” are already installed in the docker image and have licenses of use allowing unrestricted use for every kind of author (if they are appropriately cited and accredited). To integrate VarScan2 in the docker please proceed as follows.

To manage the image you need to get image name first if you did not set it in the run command, you can get the image name by running docker ps


$ docker ps
################## Sample output
# CONTAINER ID   IMAGE          COMMAND        CREATED       STATUS       PORTS                                                                                                                                                     NAMES
# 5121b762b39f   d82c97fcc2c7   "/gpro_init"   2 hours ago   Up 2 hours   0.0.0.0:20-22->20-22/tcp, :::20-22->20-22/tcp, 0.0.0.0:80->80/tcp, :::80->80/tcp,  
0.0.0.0:65500-65515->65500-65515/tcp, :::65500-65515->65500-65515/tcp   hungry_boyd

Here the name is hungry_boyd

Adding VarScan to the gpro server : run the following command:

$ docker exec hungry_boyd gpro install varscan

Commercial users should proceed in the same way but must contact first the VarScan2 authors to get the license if they want to comply with the VarScan2 terms of use.

1.3 - LINKING THE CLIENT SIDE WITH THE SERVER SIDE

After installed and running the server side image in the Docker Desktop you must link your GPRO application of interest to the Server Side in order to run the server side analyses, pipelines or workflows. As previously said, the Client Side applications of the Suite that are dependent of the Server Side are RNAseq , VariantSeq , DeNovoSeq (manual in preparation) and STATools (manual in preparation). Please visit their respective manuals for detailed instructions about how to link each application with the Server Side and automatically run the docker desktop each time you open the application linked to the Server Side

1.4 - ACKNOWLEDGEMENTS

This work was supported by the European Union’s Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement No 642095 for the OPATHY consortium, by the pre-doctoral research fellowship from Industrial Doctorates of MINECO (Grant 659 DI-17-09134); by the State Plan for Scientific and Technical Research and Innovation 2017-2020 under the Grant TSI-100903-2019-11 from the Secretary of State for Digital Advancement from Ministry of Economic Affairs and Digital Transformation, Spain; and by Expedient IDI-2021-158274-a from Ministry of Science and Innovation, Spain

1.5 - LITERATURE CITED


  • Altschul SF, Gish W, Miller W, Myers EW and Lipman DJ. 1990. Basic local alignment search tool. JMolBiol 215: 403-410. doi:10.1016/S0022-2836(05)80360-2.
  • Anders S, Pyl PT, Huber W. 2015. HTSeq — A Python framework to work with high-throughput sequencing data. Bioinformatics. 15;31(2):166-9. doi: 10.1093/bioinformatics/btu638.
  • Andrews S. 2016. “FastQC A Quality Control Tool for High Throughput Sequence Data.” http://www.bioinformatics.babraham.ac.uk/projects/fastqc/.
  • Bankevich A, Nurk S, Antipov D, et al. 2012. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. Journal of computational biology: 19: 455-477.doi: 0.1089/cmb.2012.0021 .
  • Bolger AM., Lohse M, & Usadel B. 2014. Trimmomatic: A flexible trimmer for Illumina Sequence Data. Bioinformatics, btu170.doi: 10.1093/bioinformatics/btu170.
  • Cibulskis K, Lawrence MS, Carter SL, Sivachenko A, Jaffe D, Sougnez C, Gabriel S, Meyerson M, Lander ES, Getz G. 2013. Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat Biotechnol, 31:213-219.doi: 10.1038/nbt.2514.
  • Davidson NM, and Oshlack A. 2017. Corset: Enabling Differential Gene Expression Analysis for de Novo Assembled Transcriptomes. Accessed June 28. doi:10.1186/s13059-014-0410-6.
  • DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR et al. 2011. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature Genetics 43: 491-498. doi:10.1038/ng.806.doi:10.1038/ng.806.
  • Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Bttut P, Chaisson M, Gingeras, TR. 2013. STAR: ultrafast universal RNA-seq aligner. Bioinformatics (Oxford, England), 29(1), 15–21doi:10.1093/bioinformatics/bts635.
  • Gao S, Sung WK, and Nagarajan N. 2011. Opera: reconstructing optimal genomic scaffolds with high-throughput paired-end sequences. Journal of computational biology. 18: 1681-1691. doi: 10.1089/cmb.2011.0170
  • Goff L, Trapnell C, and Kelley D. 2013. cummeRbund: Analysis, exploration, manipulation, and visualization of Cufflinks high-throughput sequencing data. R package version 2.20.0.
  • Hannon Lab. 2016. FASTX Toolkit. http://hannonlab.cshl.edu/fastx_toolkit/.
  • Kim D, Pertea G, Trapnell C, Pimentel H, Kelley R and Salzberg SL. 2013. “TopHat2: Accurate Alignment of Transcriptomes in the Presence of Insertions, Deletions and Gene Fusions.” Genome Biology 14 (4): R36. doi:10.1186/gb-2013-14-4-r36.
  • Kim D, Langmead B, Salzberg SL. 2015. “HISAT: A Fast Spliced Aligner with Low Memory Requirements.” Nature Methods 12 (4). Nature Research: 357–60. doi:10.1038/nmeth.3317.
  • Koboldt DC, Zhang Q, Larson DE, Shen D, McLellan MD, Lin L, Miller CA, Mardis ER, Ding L, Wilson RK. 2012. VarScan 2: Somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Research doi: 10.1101/gr.129684.111.
  • Koren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM. 2017. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 27, 722–736.doi: 10.1101/gr.215087.116
  • Langmead B, and Salzberg SL. 2012. “Fast Gapped-Read Alignment with Bowtie 2.” Article. Nat Methods 9. doi:10.1038/nmeth.1923.
  • Li H, Handsaker B, Wysoker A, Fennell T, Ruan J et al. 2009. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25: 2078-2079. doi:10.1093/bioinformatics/btp352
  • Li H, and Durbin R. 2009. “Fast and Accurate Long-Read Alignment with Burrows-Wheeler Transform.” Bioinformatics 26 (5): 589–95. doi:10.1093/bioinformatics/btp698
  • Love MI, Huber W and Anders S. 2014. “Moderated Estimation of Fold Change and Dispersion for RNA-Seq Data with DESeq2.” Genome Biology 15. doi:10.1186/s13059-014-0550-8.
  • Luo R, Liu B, Xie Y, et al. 2012. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1: 18.doi:10.1186/2047-217X-1-18
  • Martin M. 2011. Cutadapt Removes Adapter Sequences from High-Throughput Sequencing Reads.EMBnet.journal 17 (1): 10doi:10.14806/ej.17.1.200.
  • McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K et al. 2010. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Research 20: 1297-1303. doi:10.1101/gr.107524.110.
  • McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, Flicek P and Cunningham F. 2016. The Ensembl Variant Effect Predictor. Genome Biology. 17-122 doi.org/10.1186/s13059-016-0974-4.
  • Merkel D. 2014. Docker: lightweight Linux containers for consistent development and deployment. Linux Journal 2014:2 Merkel et al 2014:2.
  • Mistry J, Finn RD, Eddy SR, Bateman A, Punta M. 2013. Challenges in homology search: HMMER3 and convergent evolution of coiled-coil regions. Nucleic Acids Res. 2013;41(12):e121. doi:10.1093/nar/gkt263.
  • Quinlan AR, and Hall IM. 2010. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 26(6): 841–842. doi:10.1093/bioinformatics/btq033.
  • Robinson MD, McCarthy DJ, and Smyth GK. 2010. “edgeR: A Bioconductor Package for Differential Expression Analysis of Digital Gene Expression Data.” Bioinformatics (Oxford, England) 26 (1). Oxford University Press: 139–40. doi:10.1093/bioinformatics/btp616.
  • Sahlin K, Vezzi F, Nystedt B, Lundeberg J, and Arvestad L. 2014. BESST - Efficient scaffolding of large fragmented assemblies. BMC Bioinformatics 15:281.doi: 10.1186/1471-2105-15-281.
  • Stanke M, Diekhans M, Baertsch R, and Haussler D. 2008. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24: 637-644.doi: 10.1093/bioinformatics/btn013.
  • Schmieder R and Edwards R. 2011. Quality Control and Preprocessing of Metagenomic Datasets. Bioinformatics 27 (6): 863–64. doi:10.1093/bioinformatics/btr026.
  • Schulz MH, Zerbino DR, Vingron M and Birney E. 2012. Oases: robust de novo RNA-seq assembly across the dynamic range of expression levels. Bioinformatics 28: 1086-1092. doi: 10.1093/bioinformatics/bts094.
  • Trapnell C, Roberts A, Goff L, Pertea G, Kim D, Kelley DR, Pimentel H, Salzberg SL, Rinn JL and Pachter L. 2012. Differential Gene and Transcript Expression Analysis of RNA-Seq Experiments with TopHat and Cufflinks. Nature Protocols 7 (3). Nature Research: 562–78. doi:10.1038/nprot.2012.016.
  • Wysoker A, Tibbetts K, Fennell T. 2011. PicardTools 1.5.3. Available at http://sourceforge.net/projects/picard/files/picard-tools/.
  • Xie Y, Wu G, Tang J, et al. 2014. SOAPdenovo-Trans: de novo transcriptome assembly with short RNA-Seq reads. Bioinformatics 30: 1660-1666. doi: 10.1093/bioinformatics/btu077.
  • Young MD, Wakefield MJ, Smyth GK, and Oshlack A. 2010. Gene Ontology Analysis for RNA-Seq: Accounting for Selection Bias. Genome Biology 11. http://genomebiology.com/2010/11/2/R14.
  • Zerbino DR, and Birney E. 2008. Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Genome Res 18: 821-829.doi: 10.1101/gr.074492.107.
Sign in to your account