Deep learning models to design synthetic enhancers

da | Giu 6, 2024 | Artificial Intelligence, Biologia Molecolare, Deep Learning, Machine Learning

Fig. 1: Exploiting deep learning models, trained on scATAC-seq datasets, to design synthetic enhancers starting from random DNA sequences. (Image created with BioRender.com)

Abstract

Enhancers are non-coding elements in the genome that regulate the spatiotemporal activation of transcription on their target genes. However, the logic and the grammar behind this regulatory mechanism are still a challenge to overcome. Advances in the use of deep learning and transfer learning models in molecular biology have been recently applied for the de novo design of tissue-, cell type- and cell state-specific synthetic enhancers. Through these models, the strength, combination and arrangement of transcription factors activator and repressor motifs have been evaluated, allowing the synthesis of functional enhancers in adult and embryonic Drosophila melanogaster tissues, as well as in human cells.

Review

 

Introduction

Enhancers are cis-regulatory elements fundamental for transcription regulation. The DNA motifs in their sequence act as binding sites for specific transcription factors (TFs). The combination of TFs permits enhancer activation or repression, leading to cell-type-specific gene expression1. Enhancers regulatory capacity is contained in their sequence, and, for this reason, the development of models that are able to decode enhancer logic and predict gene expression is of primary interest. Thanks to the improvement in computational methods, the design of functional synthetic enhancers has been made possible by exploiting deep learning models.

Two recent studies2,3, both published in December 2023, reported the usage of Convolutional Neural Networks (CNNs) as a tool to design synthetic enhancers. CNNs are deep learning models, made by many layers convolutionally connected, that recognise local structure in image-like data. Applications of CNNs include modelling and prediction of TF binding sites using chromatin profiling data. Since the DNA sequence has just one spatial dimension, the CNN applies a ‘filter’ or ‘kernel’ that slides in one direction in search for TF binding sites4–6. Thus, the model, starting from a sequence, is able to predict the activity of an enhancer for a specific cell-type.

Deep learning models

Genome wide profiling of chromatin accessibility at single cell resolution by transposase-accessible chromatin with sequencing (scATAC-seq), histone modifications, TF binding and enhancer activity datasets are used to train deep learning models to discover TF motifs and enhancer rules.

In the first study2, Taskiran et al. employed two deep learning models, trained and validated in previous works: DeepFlyBrain7 and DeepMEL28. Specifically, DeepFlyBrain is trained on differentially accessible regions for various cell types of Drosophila melanogaster, while DeepMEL2 on chromatin accessibility data of human melanoma cell lines.

In the second study3, carried out by a different research group, De Almeida et al. made an important upgrade in the investigation of developmental enhancers using transfer learning. This strategy improves prediction performance by taking a pre-trained model that uses large datasets, sharing similarities with the target tasks, and refines its training by specific-adjustments computed on a smaller dataset. The authors first trained a sequence-to-accessibility model based on scATAC-seq datasets (in Drosophila). Next, the pre-trained model was refined using data from in vivo enhancer activity assays, defining a sequence-to-activity model (Fig. 2).

The models used by Taskiran et al. differ from the one developed in the second article for the fact that they base their prediction only on chromatin accessibility, whereas the second one combines accessibility with knowledge coming from in vivo enhancer activity assays.

Fig. 2: The upper part shows the sequence-to-accessibility model trained on DNA accessibility data. The bottom part represents the sequence-to-activity model obtained by transfer learning, which refines its knowledge with data from in vivo enhancer activity assays.

Discussion

Taskiran et al.2 initially developed synthetic enhancers to specifically target Kenyon cells (KC) in the mushroom body of the fruit fly brain. The first strategy employed was the nucleotide-by-nucleotide sequence evolution approach, evolving a random sequence (500 bp) from scratch, towards the chosen cell-type, using DeepFlyBrain as a guide (i.e. obtaining a sequence with high predicted activity in the chosen cell-type). At each iteration, the authors performed saturation mutagenesis: all nucleotides were mutated one by one, and each sequence variation was scored by DeepFlyBrain, to select the mutation with the greatest positive delta score. This method led to an optimal score after 10 to 15 mutations, resulting in the creation of known TF activator motifs. The first generated mutations were the ones that disrupted the TF repressor motifs, which made the enhancers non-functional, significantly lowering DeepFlyBrain prediction scores. Building a GFP reporter system, the authors validated, through in vivo testing, the designed enhancers with high prediction score, resulting in 10 out of 13 active enhancers in vivo.

Next, Taskiran et al. investigated how fast genomic regions with already high prediction scores, but no activity, could evolve towards functional enhancers. Interestingly, they observed that only six mutations were needed. This suggested that cell-type specific enhancers can arise de novo in the genome requiring only few mutations. Using the sequence evolution method once again, the authors were able to evolve an enhancer, active in one cell-type, to be functional simultaneously in a second cell-type, thereby obtaining a dual-code enhancer. Vice versa, enhancers active in several cell types could be evolved towards a single cell-type code.

The second enhancer design strategy used was the motif implantation. Starting from a random sequence, the authors implanted the known TF activator motifs for KC, selecting the locations with the highest prediction score. Since the best position for the activator motifs resulted to be in the central region of the random sequence, it was also possible to create functional minimal enhancers of just 49 pb.

For the design of melanocyte or melanocyte-like melanoma (MEL) enhancers, Taskiran et al. used another deep learning model: DeepMEL2. The approaches were the same as in Drosophila, and led to the synthesis of active enhancers and demonstration of the dominance of the TF repressor motifs, specifically ZEB2. The model predictions were validated by in vitro testing using luciferase assay, resulting in seven out of ten active enhancers. Later, the authors confirmed that the same principles can be applied on genomic enhancers, using the MEL enhancer in the IRF4 intron as an example. Indeed, modifying the number of ZEB2 sites within the genomic enhancer sequence resulted in a significant increase (no ZEB2) or loss (many ZEB2) of the enhancer activity. As in Drosophila, after applying the motif implantation strategy successfully, Taskiran et al. created minimal enhancers, in which the activator motifs are all adjacent.

In the second study, de Almeida et al.3 thought of the possibility to exploit similar strategies to synthesise tissue-specific developmental enhancers. The authors trained a sequence-to-activity model (Fig. 2) to design enhancers of 4 distinct tissues (CNS, gut, muscle, and epidermis) of Drosophila melanogaster embryos between 10-to-12 hours. They generated 1001 bp random sequence (501 bp enhancer and 250 bp at both ends to determine a context) as input for the transfer learning model, to obtain enhancer sequences with high prediction score of accessibility and activity. In the in vivo testing, the designed enhancers showed tissue-specific activity in CNS, muscle, and epidermis, while gut enhancers showed partial additional activities in other tissues. This suggested that gut ‘enhancer grammar’ is more complex, since the same TFs are broadly used in many tissues.

Conclusions

These studies showed that it is possible to design synthetic enhancers for Drosophila melanogaster and human cells by using deep learning-based strategies. In addition, the process of combining genome-wide and smaller-scale datasets by transfer learning is generally applicable in the design of tissue-specific enhancers, although some enhancers are active in non-target tissues. This highlights the potential for developing more refined models that can discriminate between closely related tissue subtypes and individual cell types, especially those that share transcription factors. Therefore, the usage of deep learning can be applied to better understand gene expression patterns. Simultaneously, an increase in dataset dimension could lead to the generation of synthetic enhancers exploitable both in disease and homeostasis.

References

  1. Shlyueva, D., Stampfel, G. & Stark, A. Transcriptional enhancers: from properties to genome-wide predictions. Nat. Rev. Genet. 15, 272–286 (2014).
  2. Taskiran, I. I. et al. Cell-type-directed design of synthetic enhancers. Nature 626, 212–220 (2024).
  3. de Almeida, B. P. et al. Targeted design of synthetic enhancers for selected tissues in the Drosophila embryo. Nature 626, 207–211 (2024).
  4. Eraslan, G., Avsec, Ž., Gagneur, J. & Theis, F. J. Deep learning: new computational modelling techniques for genomics. Nat. Rev. Genet. 20, 389–403 (2019).
  5. Zou, J. et al. A primer on deep learning in genomics. Nat. Genet. 51, 12–18 (2019).
  6. Greener, J. G., Kandathil, S. M., Moffat, L. & Jones, D. T. A guide to machine learning for biologists. Nat. Rev. Mol. Cell Biol. 23, 40–55 (2022).
  7. Janssens, J. et al. Decoding gene regulation in the fly brain. Nature 601, 630–636 (2022).
  8. Atak, Z. K. et al. Interpretation of allele-specific chromatin accessibility using cell state–aware deep learning. Genome Res. 31, 1082–1096 (2021)

Angelica Castino

Master Cellular and Molecular Biology student

Arianna Maggiora

Master Cellular and Molecular Biology student

Michela Rossi

Master Cellular and Molecular Biology student