{"id":1244269,"date":"2022-11-05T01:48:52","date_gmt":"2022-11-05T01:48:52","guid":{"rendered":"https:\/\/www.ghanamma.com\/2022\/11\/05\/stability-ai-backs-effort-to-bring-machine-learning-to-biomed\/"},"modified":"2022-11-05T01:48:52","modified_gmt":"2022-11-05T01:48:52","slug":"stability-ai-backs-effort-to-bring-machine-learning-to-biomed","status":"publish","type":"post","link":"https:\/\/www.ghanamma.com\/2022\/11\/05\/stability-ai-backs-effort-to-bring-machine-learning-to-biomed\/","title":{"rendered":"Stability AI backs effort to bring machine learning to biomed \u2022"},"content":{"rendered":"<p><\/p>\n<div>\n<p id=\"speakable-summary\">Stability AI, the venture-backed startup behind the text-to-image AI system Stable Diffusion, is funding a wide-ranging effort to apply AI to the frontiers of biotech. Called OpenBioML, the endeavor\u2019s first projects will focus on machine learning-based approaches to DNA sequencing, protein folding and computational biochemistry.<\/p>\n<p>The company\u2019s founders describe OpenBioML as an \u201copen research laboratory\u201d \u2014 and aims to explore the intersection of AI and biology in a setting where students, professionals and researchers can participate and collaborate, according to Stability AI CEO Emad Mostaque.<\/p>\n<p>\u201cOpenBioML is one of the independent research communities that Stability supports,\u201d Mostaque told  in an email interview. \u201c<span style=\"font-size: 1rem; letter-spacing: -0.1px;\">Stability looks to develop and democratize AI, and through OpenBioML, we see an opportunity to advance the state of the art in sciences, health and medicine.\u201d<\/span><\/p>\n<p>Given the controversy surrounding Stable Diffusion \u2014 Stability AI\u2019s AI system that generates art from text descriptions, similar to OpenAI\u2019s DALL-E 2 \u2014 one might be understandably wary of Stability AI\u2019s first venture into healthcare. The startup has taken a laissez-faire approach to governance, allowing developers to use the system however they wish, including for celebrity deepfakes and pornography.<\/p>\n<p>Stability AI\u2019s ethically questionable decisions to date aside, machine learning in medicine is a minefield. While the tech has been successfully applied to diagnose conditions like skin and eye diseases, among others, research has shown that algorithms can develop biases leading to worse care for some patients. An April 2021 study, for example, found that statistical models used to predict suicide risk in mental health patients performed well for white and Asian patients but poorly for Black patients.<\/p>\n<p><span style=\"font-size: 1rem; letter-spacing: -0.1px;\">OpenBioML is starting with safer territory, wisely. Its first projects are:<\/span><\/p>\n<ul>\n<li><span style=\"font-size: 1rem; letter-spacing: -0.1px;\"><strong>BioLM<\/strong>, which seeks to apply natural language processing (NLP) techniques to the fields of computational biology and chemistry<\/span><\/li>\n<li><span style=\"font-size: 1rem; letter-spacing: -0.1px;\"><strong>DNA-Diffusion<\/strong>, which aims to develop AI that can generate DNA sequences from text prompts<\/span><\/li>\n<li><span style=\"font-size: 1rem; letter-spacing: -0.1px;\"><strong>LibreFold<\/strong>, which looks to increase access to AI protein structure prediction systems similar to DeepMind\u2019s AlphaFold 2<\/span><\/li>\n<\/ul>\n<p>Each project is led by independent researchers, but Stability AI is providing support in the form of access to its AWS-hosted cluster of over 5,000 Nvidia A100 GPUs to train the AI systems. According to Niccol\u00f2 Zanichelli, a computer science undergraduate at the University of Parma and one of the lead researchers at <span style=\"font-size: 1rem; letter-spacing: -0.1px;\">OpenBioML, this will be<\/span><span style=\"font-size: 1rem; letter-spacing: -0.1px;\"> enough processing power and storage to eventually train up to 10 different AlphaFold 2-like systems in parallel.<\/span><\/p>\n<p>\u201cA lot of computational biology research already leads to open-source releases. However, much of it happens at the level of a single lab and is therefore usually constrained by insufficient computational resources,\u201d Zanichelli told  via email. \u201cWe want to change this by encouraging large-scale collaborations and, thanks to the support of Stability AI, back those collaborations with resources that only the largest industrial laboratories have access to.\u201d<\/p>\n<h2>Generating DNA sequences<\/h2>\n<p>Of <span style=\"font-size: 1rem; letter-spacing: -0.1px;\">OpenBioML\u2019s ongoing projects, <\/span>DNA-Diffusion \u2014 led by pathology professor Luca Pinello\u2019s lab at the Massachusetts General Hospital &amp; Harvard Medical School \u2014 is perhaps the most ambitious. The goal is to use generative AI systems to learn and apply the rules of \u201cregulatory\u201d sequences of DNA, or segments of nucleic acid molecules that influence the expression of specific genes within an organism. Many diseases and disorders are the result of misregulated genes, but science has yet to discover a reliable process for identifying \u2014 much less changing \u2014 these regulatory sequences.<\/p>\n<p>DNA-Diffusion proposes using a type of AI system known as a diffusion model to generate cell-type-specific regulatory DNA sequences. Diffusion models \u2014 which underpin image generators like Stable Diffusion and OpenAI\u2019s DALL-E 2 \u2014 create new data (e.g. DNA sequences) by learning how to destroy and recover many existing samples of data. As they\u2019re fed the samples, the models get better at recovering all the data they had previously destroyed to generate new works.<\/p>\n<p>\u201cDiffusion has seen widespread success in multimodal generative models, and it is now starting to be applied to computational biology, for example for the generation of novel protein structures,\u201d Zanichelli said. \u201cWith DNA-Diffusion, we\u2019re now exploring its application to genomic sequences.\u201d<\/p>\n<p>If all goes according to plan, the DNA-Diffusion project will produce a diffusion model that can generate regulatory DNA sequences from text instructions like \u201cA sequence that will activate a gene to its maximum expression level in cell type X\u201d and \u201cA sequence that activates a gene in liver and heart, but not in brain.\u201d Such a model could also help interpret the components of regulatory sequences, Zanichelli says \u2014 improving the scientific community\u2019s understanding of the role of regulatory sequences in different diseases.<\/p>\n<p>It\u2019s worth noting that this is largely theoretical. While preliminary research on applying diffusion to protein folding seems promising, it\u2019s very early days, Zanichelli admits \u2014 hence the push to involve the wider AI community.<\/p>\n<h2>Predicting protein structures<\/h2>\n<p><span style=\"font-size: 1rem; letter-spacing: -0.1px;\">OpenBioML\u2019s LibreFold, while smaller in scope, is more likely to bear immediate fruit. The project seeks to arrive at a better understanding of machine learning systems that predict protein structures in addition to ways to improve them.<\/span><\/p>\n<p>As my colleague Devin Coldewey covered in his piece about DeepMind\u2019s work on AlphaFold 2, AI systems that accurately predict protein shape are relatively new on the scene but transformative in terms of their potential. Proteins comprise sequences of amino acids that fold into shapes to accomplish different tasks within living organisms. The process of determining what shape an acids sequence will create was once an arduous, error-prone undertaking. AI systems like AlphaFold 2 changed that; thanks to them, over 98% of protein structures in the human body are known to science today, as well as hundreds of thousands of other structures in organisms like E. coli and yeast.<\/p>\n<p>Few groups have the engineering expertise and resources necessary to develop this kind of AI, though. DeepMind spent days training AlphaFold 2 on tensor processing units (TPUs), Google\u2019s costly AI accelerator hardware. And acid sequence training data sets are often proprietary or released under non-commercial licenses.<\/p>\n<div id=\"attachment_2181067\" style=\"width: 1034px\" class=\"wp-caption aligncenter\"><img aria-describedby=\"caption-attachment-2181067\" decoding=\"async\" loading=\"lazy\" class=\"size-full wp-image-2181067\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/11\/GettyImages-1279331936.jpg\" alt=\"\" width=\"1024\" height=\"561\"\/><\/p>\n<p id=\"caption-attachment-2181067\" class=\"wp-caption-text\">Proteins folding into their three-dimensional structure. <strong>Image Credits:<\/strong> Christoph Burgstedt\/Science Photo Library \/ Getty Images<\/p>\n<\/div>\n<p>\u201cThis is a pity, because if you look at what the community has been able to build on top of the AlphaFold 2 checkpoint released by DeepMind, it\u2019s simply incredible,\u201d Zanichelli said, referring to the trained AlphaFold 2 model that DeepMind released last year. \u201c<span style=\"font-size: 1rem; letter-spacing: -0.1px;\">For example, just days after the release, Seoul National University professor Minkyung Baek reported a trick on Twitter that allowed the model to predict quaternary structures \u2014 something which few, if anyone, expected the model to be capable of. There are many more examples of this kind, so who knows what the wider scientific community could build if it had the ability to train entirely new AlphaFold-like protein structure prediction methods?\u201d<\/span><\/p>\n<p>Building on the work of RoseTTAFold and OpenFold, two ongoing community efforts to replicate AlphaFold 2, <span style=\"font-size: 1rem; letter-spacing: -0.1px;\">LibreFold will facilitate \u201clarge-scale\u201d experiments with various protein folding prediction systems. Spearheaded by researchers at University College London, Harvard and Stockholm, LibreFold\u2019s focus will be to gain a better understanding of what the systems can accomplish and why, according to Zanichelli.\u00a0<\/span><\/p>\n<p>\u201cLibreFold is at its heart a project for the community, by the community. The same holds for the release of both model checkpoints and data sets, as it could take just one or two months for us to start releasing the first deliverables or it could take significantly longer,\u201d he said. \u201cThat said, my intuition is that the former is more likely.\u201d<\/p>\n<h2>Applying NLP to biochemistry<\/h2>\n<p>On a longer time horizon is <span style=\"font-size: 1rem; letter-spacing: -0.1px;\">OpenBioML\u2019s <\/span>BioLM project, which has the vaguer mission of \u201capplying language modeling techniques derived from NLP to biochemical sequences.\u201d In collaboration with EleutherAI, a research group that\u2019s released several open source text-generating models, BioLM hopes to train and publish new \u201cbiochemical language models\u201d for a range of tasks, including generating protein sequences.<\/p>\n<p><span style=\"font-size: 1rem; letter-spacing: -0.1px;\">Zanichelli points to Salesforce\u2019s ProGen as an example of the types of work BioLM might embark on. ProGen treats amino acid sequences like words in a sentence. Trained on a dataset of more than 280 million protein sequences and associated metadata, the model predicts the next set of amino acids from the previous ones, like a language model predicting the end of a sentence from its beginning.<\/span><\/p>\n<p>Nvidia earlier this year released a language model, MegaMolBART, that was trained on a dataset of millions of molecules to search for potential drug targets and forecast chemical reactions. Meta also recently trained an NLP called ESM-2 on sequences of proteins, an approach the company claims allowed it to predict sequences for more than 600 million proteins in just two weeks.<\/p>\n<div id=\"attachment_2434957\" style=\"width: 1034px\" class=\"wp-caption aligncenter\"><img aria-describedby=\"caption-attachment-2434957\" decoding=\"async\" loading=\"lazy\" class=\"size-full wp-image-2434957\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/11\/310702728_517882743006587_8593866092610342553_n.png\" alt=\"Meta protein folding\" width=\"1024\" height=\"576\"\/><\/p>\n<p id=\"caption-attachment-2434957\" class=\"wp-caption-text\">Protein structures predicted by Meta\u2019s system. <strong>Image Credits:<\/strong> Meta<\/p>\n<\/div>\n<h2>Looking ahead<\/h2>\n<p>While OpenBioML\u2019s interests are broad (and expanding), Mostaque says that they\u2019re unified by a desire to \u201cmaximize the positive potential of machine learning and AI in biology,\u201d following in the tradition of open research in science and medicine.<\/p>\n<p>\u201cWe are looking to enable researchers to gain more control over their experimental pipeline for active learning or model validation purposes,\u201d Mostaque continued. \u201cWe\u2019re also looking to push the state of the art with increasingly general biotech models, in contrast to the specialized architectures and learning objectives that currently characterize most of computational biology.\u201d<\/p>\n<p>But \u2014 as might be expected from a VC-backed startup that recently raised over $100 million \u2014 Stability AI doesn\u2019t see OpenBioML as a purely philanthropic effort. Mostaque says that the company is open to exploring commercializing tech from OpenBioML \u201cwhen it\u2019s advanced enough and safe enough and when the time is right.\u201d<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Stability AI, the venture-backed startup behind the text-to-image AI system Stable Diffusion, is funding a wide-ranging effort to apply AI to the frontiers of biotech. Called OpenBioML, the endeavor\u2019s first projects will focus on machine learning-based approaches to DNA sequencing, protein folding and computational biochemistry. The company\u2019s founders describe OpenBioML as an \u201copen research laboratory\u201d [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1244271,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[21],"tags":[],"class_list":["post-1244269","post","type-post","status-publish","format-standard","has-post-thumbnail","category-celebrity-gossip"],"_links":{"self":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/posts\/1244269","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/comments?post=1244269"}],"version-history":[{"count":0,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/posts\/1244269\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/media?parent=1244269"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/categories?post=1244269"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/tags?post=1244269"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}