The gDRcore
is the part of the gDR
suite. The package provides set of tools to proces and analyze drug response data.
The data model is built on the MultiAssayExperiments (MAE) structure. Within an MAE, each SummarizedExperiment (SE) contains a different unit type (e.g. single-agent, or combination treatment). Columns of the MAE are defined by the cell lines and any modification of them and are shared with the SEs. Rows are defined by the treatments (e.g drugs, perturbations) and are specific to each SE. Assays of the SE are the different levels of data processing (raw, control, normalized, averaged data, as well as metrics). Each nested element of the assays of the SEs comprises the series themselves as a table (data.table in practice). Although not all elements need to have a series or the same number of elements, the attributes (columns of the table) should be consistent across the SE.
For drug response data, the input files need to be merged such that each measurement (data) is associated with the right metadata (cell line properties and treatment definition). Metadata can be added with the function cleanup_metadata
if the right reference databases are in place.
When the data and metadata are merged into a long table, the wrapper function runDrugResponseProcessingPipeline
can be used to generate an MAE with processed and analyzed data.
.
In practice runDrugResponseProcessingPipeline does the following steps:
create_SE
creates the structure of the MAE and the associated SEs by assigning metadata into the row and column attributes. The assignment is performed in the function split_SE_components (see details below for the assumption made when building SE structures).
create_SE also dispatches the raw data and controls into the right nested tables. Note that data may be duplicated between different SEs to make them self-contained.normalize_SE
normalizes the raw data based on the control. Calculation of the GR value is based on a cell line division time provided by the reference database if no pre-treatment control is provided. If both information are missing, GR values cannot be calculated. Additional normalization can be added as new rows in the nested table.average_SE
averages technical replicates that are stored in the same nested table are averaged.fit_SE
fits the dose-response curves and calculates response metrics for each normalization type.fit_SE.combinations
calculates synergy scores for drug combination data and, if the data is appropriate, fits along the two drugs and matrix-level metrics (e.g. isobolograms) are calculated. This is also performed for each normalization type independently..
The functions to process the data have parameters for specifying the names of the variables and assays. Additional parameters are available to personalize the processing steps such as force the nesting (or not) of an attribute, specify attributes that should be considered as technical replicates or not.
Please familiarize with gDRimport
package containing bunch of tools allowing to prepare input data for gDRcore
.
This example is made up based on the artificial dataset called data1
available within gDRimport
package. gDR
required three types of data
that should be used as the raw input: Template, Manifest, and RawData. More info about these three types of data you could find in our general
documentation.
td <- gDRimport::get_test_data()
Provided dataset needs to be merged into the one data.table
object to be able to run gDR pipeline. This process can be done using two functions –
gDRimport::load_data()
and gDRcore::merge_data()
.
We provide an all-in-one function that splits data into appropriate data types, creates the SummarizedExperiment object for each data type, splits data into treatment and control assays, normalizes, averages, calculates gDR metrics, and finally, creates the MultiAssayExperiment object. This function is called runDrugResponseProcessingPipeline
.
mae <- runDrugResponseProcessingPipeline(input_df)
mae
#> A MultiAssayExperiment object of 1 listed
#> experiment with a user-defined name and respective class.
#> Containing an ExperimentList class object of length 1:
#> [1] single-agent: SummarizedExperiment with 4 rows and 6 columns
#> Functionality:
#> experiments() - obtain the ExperimentList instance
#> colData() - the primary/phenotype DataFrame
#> sampleMap() - the sample coordination DataFrame
#> `$`, `[`, `[[` - extract colData columns, subset, or experiment
#> *Format() - convert into a long or wide DataFrame
#> assays() - convert ExperimentList to a SimpleList of matrices
#> exportClass() - save data to flat files
And we can subset the MultiAssayExperiment to receive the SummarizedExperiment specific to any data type, e.g.
mae[["single-agent"]]
#> class: SummarizedExperiment
#> dim: 4 6
#> metadata(5): identifiers experiment_metadata Keys fit_parameters
#> .internal
#> assays(5): RawTreated Controls Normalized Averaged Metrics
#> rownames(4): G00002_drug_002_moa_A_168_0
#> G00002_drug_002_moa_A_168_0.149999910525364
#> G00004_drug_004_moa_A_168_0
#> G00004_drug_004_moa_A_168_0.149999910525364
#> rowData names(5): Gnumber DrugName drug_moa Duration drug_011
#> colnames(6): CL00011_cellline_BA_breast_cellline_BA_unknown_26
#> CL00012_cellline_CA_breast_cellline_CA_unknown_30 ...
#> CL00015_cellline_FA_breast_cellline_FA_unknown_42
#> CL00018_cellline_IB_breast_cellline_IB_unknown_54
#> colData names(6): clid CellLineName ... subtype ReferenceDivisionTime
Extraction of the data from either MultiAssayExperiment
or SummarizedExperiment
objects into more user-friendly structures as well as other data transformations can be done using gDRutils
. We encourage to read gDRutils
vignette to familiarize with these functionalities.
sessionInfo()
#> R version 4.4.1 (2024-06-14)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 22.04.4 LTS
#>
#> Matrix products: default
#> BLAS: /home/biocbuild/bbs-3.20-bioc/R/lib/libRblas.so
#> LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.10.0
#>
#> locale:
#> [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C
#> [3] LC_TIME=en_GB LC_COLLATE=C
#> [5] LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8
#> [7] LC_PAPER=en_US.UTF-8 LC_NAME=C
#> [9] LC_ADDRESS=C LC_TELEPHONE=C
#> [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C
#>
#> time zone: America/New_York
#> tzcode source: system (glibc)
#>
#> attached base packages:
#> [1] stats graphics grDevices utils datasets methods base
#>
#> other attached packages:
#> [1] gDRcore_1.3.11 gDRtestData_1.3.2 BiocStyle_2.33.1
#>
#> loaded via a namespace (and not attached):
#> [1] fastmap_1.2.0 BumpyMatrix_1.13.0
#> [3] TH.data_1.1-2 digest_0.6.37
#> [5] lifecycle_1.0.4 gDRutils_1.3.9
#> [7] survival_3.7-0 magrittr_2.0.3
#> [9] compiler_4.4.1 rlang_1.1.4
#> [11] sass_0.4.9 drc_3.0-1
#> [13] tools_4.4.1 plotrix_3.8-4
#> [15] utf8_1.2.4 yaml_2.3.10
#> [17] data.table_1.15.4 knitr_1.48
#> [19] lambda.r_1.2.4 S4Arrays_1.5.7
#> [21] DelayedArray_0.31.11 abind_1.4-5
#> [23] multcomp_1.4-26 BiocParallel_1.39.0
#> [25] purrr_1.0.2 BiocGenerics_0.51.0
#> [27] grid_4.4.1 stats4_4.4.1
#> [29] fansi_1.0.6 colorspace_2.1-1
#> [31] scales_1.3.0 gtools_3.9.5
#> [33] MASS_7.3-61 MultiAssayExperiment_1.31.5
#> [35] SummarizedExperiment_1.35.1 cli_3.6.3
#> [37] mvtnorm_1.2-6 rmarkdown_2.28
#> [39] crayon_1.5.3 httr_1.4.7
#> [41] readxl_1.4.3 cachem_1.1.0
#> [43] stringr_1.5.1 zlibbioc_1.51.1
#> [45] splines_4.4.1 gDRimport_1.3.2
#> [47] assertthat_0.2.1 parallel_4.4.1
#> [49] formatR_1.14 BiocManager_1.30.24
#> [51] cellranger_1.1.0 XVector_0.45.0
#> [53] matrixStats_1.3.0 vctrs_0.6.5
#> [55] Matrix_1.7-0 sandwich_3.1-0
#> [57] jsonlite_1.8.8 carData_3.0-5
#> [59] bookdown_0.40 car_3.1-2
#> [61] IRanges_2.39.2 S4Vectors_0.43.2
#> [63] testthat_3.2.1.1 jquerylib_0.1.4
#> [65] rematch_2.0.0 glue_1.7.0
#> [67] codetools_0.2-20 stringi_1.8.4
#> [69] futile.logger_1.4.3 GenomeInfoDb_1.41.1
#> [71] GenomicRanges_1.57.1 UCSC.utils_1.1.0
#> [73] munsell_0.5.1 tibble_3.2.1
#> [75] pillar_1.9.0 htmltools_0.5.8.1
#> [77] brio_1.1.5 GenomeInfoDbData_1.2.12
#> [79] R6_2.5.1 evaluate_0.24.0
#> [81] lattice_0.22-6 Biobase_2.65.0
#> [83] futile.options_1.0.1 backports_1.5.0
#> [85] bslib_0.8.0 SparseArray_1.5.31
#> [87] checkmate_2.3.2 xfun_0.47
#> [89] MatrixGenerics_1.17.0 zoo_1.8-12
#> [91] pkgconfig_2.0.3