Optimizing data integration improves Gene Regulatory Network inference in Arabidopsis thaliana
Abstract Motivations Gene Regulatory Networks (GRN) are traditionnally inferred from gene expression profiles monitoring a specific condition or treatment. In the last decade, integrative strategies have successfully emerged to guide GRN inference from gene expression with complementary prior data. However, datasets used as prior information and validation gold standards are often related and limited to a subset of genes. This lack of complete and independent evaluation calls for new criteria to robustly estimate the optimal intensity of prior data integration in the inference process.Results We address this issue for two common regression-based GRN inference models, an integrative Random Forest (weigthedRF) and a generalized linear model with stability selection estimated under a weighted LASSO penalty (weightedLASSO). These approaches are applied to data from the root response to nitrate induction inArabidopsis thaliana. For each gene, we measure how the integration of transcription factor binding motifs influences model prediction. We propose a new approach, DIOgene, that uses model prediction error and a simulated null hypothesis for optimizing data integration strength in a hypothesis-driven, gene-specific manner. The resulting integration scheme reveals a strong diversity of optimal integration intensities between genes. In addition, it provides a good trade-off between prediction error minimization and validation on experimental interactions, while master regulators of nitrate induction can be accurately retrieved.Availability and implementation The R code and notebooks demonstrating the use of the proposed approaches are available in the repository<jats:ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://github.com/OceaneCsn/integrative_GRN_N_induction">https://github.com/OceaneCsn/integrative_GRN_N_induction</jats:ext-link>..
Medienart: |
Preprint |
---|
Erscheinungsjahr: |
2023 |
---|---|
Erschienen: |
2023 |
Enthalten in: |
bioRxiv.org - (2023) vom: 04. Okt. Zur Gesamtaufnahme - year:2023 |
---|
Sprache: |
Englisch |
---|
Beteiligte Personen: |
Cassan, Océane [VerfasserIn] |
---|
Links: |
Volltext [kostenfrei] |
---|
Themen: |
---|
doi: |
10.1101/2023.09.29.558791 |
---|
funding: |
|
---|---|
Förderinstitution / Projekttitel: |
|
PPN (Katalog-ID): |
XBI040990761 |
---|
LEADER | 01000caa a22002652 4500 | ||
---|---|---|---|
001 | XBI040990761 | ||
003 | DE-627 | ||
005 | 20231205144102.0 | ||
007 | cr uuu---uuuuu | ||
008 | 231002s2023 xx |||||o 00| ||eng c | ||
024 | 7 | |a 10.1101/2023.09.29.558791 |2 doi | |
035 | |a (DE-627)XBI040990761 | ||
035 | |a (biorXiv)10.1101/2023.09.29.558791 | ||
040 | |a DE-627 |b ger |c DE-627 |e rakwb | ||
041 | |a eng | ||
100 | 1 | |a Cassan, Océane |e verfasserin |0 (orcid)0000-0002-4595-2457 |4 aut | |
245 | 1 | 0 | |a Optimizing data integration improves Gene Regulatory Network inference in Arabidopsis thaliana |
264 | 1 | |c 2023 | |
336 | |a Text |b txt |2 rdacontent | ||
337 | |a Computermedien |b c |2 rdamedia | ||
338 | |a Online-Ressource |b cr |2 rdacarrier | ||
520 | |a Abstract Motivations Gene Regulatory Networks (GRN) are traditionnally inferred from gene expression profiles monitoring a specific condition or treatment. In the last decade, integrative strategies have successfully emerged to guide GRN inference from gene expression with complementary prior data. However, datasets used as prior information and validation gold standards are often related and limited to a subset of genes. This lack of complete and independent evaluation calls for new criteria to robustly estimate the optimal intensity of prior data integration in the inference process.Results We address this issue for two common regression-based GRN inference models, an integrative Random Forest (weigthedRF) and a generalized linear model with stability selection estimated under a weighted LASSO penalty (weightedLASSO). These approaches are applied to data from the root response to nitrate induction inArabidopsis thaliana. For each gene, we measure how the integration of transcription factor binding motifs influences model prediction. We propose a new approach, DIOgene, that uses model prediction error and a simulated null hypothesis for optimizing data integration strength in a hypothesis-driven, gene-specific manner. The resulting integration scheme reveals a strong diversity of optimal integration intensities between genes. In addition, it provides a good trade-off between prediction error minimization and validation on experimental interactions, while master regulators of nitrate induction can be accurately retrieved.Availability and implementation The R code and notebooks demonstrating the use of the proposed approaches are available in the repository<jats:ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://github.com/OceaneCsn/integrative_GRN_N_induction">https://github.com/OceaneCsn/integrative_GRN_N_induction</jats:ext-link>. | ||
650 | 4 | |a Biology |7 (dpeaa)DE-84 | |
650 | 4 | |a 570 |7 (dpeaa)DE-84 | |
700 | 1 | |a Lecellier, Charles-Henri |0 (orcid)0000-0002-0229-5434 |4 aut | |
700 | 1 | |a Martin, Antoine |0 (orcid)0000-0002-6956-2904 |4 aut | |
700 | 1 | |a Bréhélin, Laurent |0 (orcid)0000-0002-2582-2831 |4 aut | |
700 | 1 | |a Lèbre, Sophie |0 (orcid)0000-0003-3444-2416 |4 aut | |
773 | 0 | 8 | |i Enthalten in |t bioRxiv.org |g (2023) vom: 04. Okt. |
773 | 1 | 8 | |g year:2023 |g day:04 |g month:10 |
856 | 4 | 0 | |u http://dx.doi.org/10.1101/2023.09.29.558791 |z kostenfrei |3 Volltext |
912 | |a GBV_XBI | ||
951 | |a AR | ||
952 | |j 2023 |b 04 |c 10 |