Development of a random background to understand ligand optimization
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript.
Ligand optimization is central to drug discovery, with hundreds of analogues often designed and synthesized between an initial hit and a therapeutic candidate1,2. The efficiency of this process is unclear, partly because there is no random background for optimization to compare against. Such a random background might emerge from systematic random small substitutions across starting ligands, measuring the likelihood of achieving a substantial improvement in affinity or potency, or other property by any single perturbation. Recent literature has suggested that perhaps 10% of analogues with minor modifications improve upon the potency of a parent by tenfold or more3,4, but this number is clouded by reporting bias, intentional improvement and inter-group variability. To begin to establish a background expectation for ligand optimization, here we systematically modified 18 lead molecules across six targets with single-atom changes; 257 compounds were synthesized. Unexpectedly, 11.3% of these random small perturbation analogues improved potency by tenfold or more. Conversely, they typically had worse in vitro pharmacokinetics. Although it was possible to find analogues where the potency increase compensated for inferior exposure and half-life, resulting in more potent compounds in vivo, overall, a frustrated landscape for ligand optimization is revealed. This study begins to establish a background expectation for ligand potency optimization and offers a simple strategy to do so. It also begins to quantify the challenges confronting the field in moving beyond in vitro potency.
Ligand optimization is central to chemical probe and drug discovery, as initial active molecules rarely have the potency or pharmacokinetic (PK) properties to be viable in vivo1,2. Between the discovery of an initial hit and a clinical candidate many hundreds, occasionally thousands of optimized analogues might be synthesized1,5. Several strategies6,7,8,9,10 have been developed to improve this process, ranging from early empirical approaches such as Topliss Trees11 and quantitative structureâactivity relationships12 to contemporary artificial intelligence (AI)-guided PK13 prediction and free energy calculations14,15. The efficiency of these strategies remains uncertain as there is no random background against which to compare them. Using tenfold improvement between analogue and parent as a benchmark for substantial impact, 10% of analogues meet this standard in the ChEMBL database16 (Extended Data Fig. 1a). As these ChEMBL results suffer from success bias, sample multiple types of perturbation and suffer from inter-group irreproducibility17, the public domain offers no sure guidance on what level of improvement one might expect in ligand optimization from random conservative changes.
Random backgrounds have long been used in biology to quantify significance. In genomics, they distinguish between artificial and natural selection18. In epidemiology, random incidence distinguishes genetic diseases from those driven by environmental factors19,20. In protein engineering, random backgrounds help to evaluate improvements in enzyme function21, and alanine scanning introduces a minimal perturbation to find hot spots for binding and function22. A random background for ligand optimization might quantify progress across chemotypes and targets and compare different optimization strategies.
Like alanine scanning, an unbiased background for ligand optimization might involve conservative perturbations, systematically and comprehensively applied across multiple parents with different targets. This resembles a âpositional analogue scanningâ approach that substitutes ligand non-polar hydrogens with groups such as CH3, OH, Cl, F and Br, and aromatic carbon atoms with nitrogen23. In a retrospective analysis of 110,000 matched molecular pairs in the ChEMBL24 database3,4, about 30% of analogues improved over threefold in affinity or potency versus their parent. This observation is tempered by design and reporting biases towards compound improvement, by the variability in experimental conditions across datasets that can reduce reproducibility17, and by the few parent compounds that had multiple conservative changes. Thus, it seemed interesting to create a set of parent ligands that were systematically and comprehensively modified by small perturbations, without design and without a bias towards improvement, using the positional analogue scanning approach23,24 where the differences between parent and the suite of analogues were quantified by a single laboratory.
To create what we refer to as a random background set for ligand optimization, we chose 18 parent ligands spanning six targets, including three GPCRs (α2A adrenergic receptor (α2AAR), ÎŒ-opioid receptor (ÎŒOR) and cannabinoid receptor type 2 (CB2)), one transporter (the serotonin transporter (SERT)) and two soluble enzymes (AmpC ÎČ-lactamase (AmpC) and macrodomain 1 of SARS-CoV-2 (Mac1)). This set was chosen for targets that we had under experimental control and is admittedly far from comprehensive. The 18 parents were chosen for similar pragmatic reasons; they were molecules to which we had ready access. Still, these proteins provide a range of target classes while the 18 parents cover a wide range of physicochemical properties (Extended Data Fig. 2): ranging from 1 nM to 43 ”M activity (Table 1) and representing 11 unrelated scaffolds. For each parent, we systematically explored all analogues accessible by a single-atom or methyl modification where such modifications were synthetically feasible and where they did not modify the net charge of the parent at physiological pH. The single-atom modifications involved substituting carbon hydrogens with one of methyl25, hydroxyl26, chloro27, fluoro28 or rarely bromo, or substituting aromatic carbon atoms with nitrogens (Fig. 1). These modifications are not comprehensive but do represent those preferred by medicinal chemists.
Full size imagea, Parent ligands from six drug targets (the aminergic GPCR α2AAR, the lipid GPCR CB2R, the peptide GPCR ”OR, the serotonin SLC transporter SERT, and the soluble enzymes AmpC and Mac1) were modified by one-at-a-time, one-atom substitutions: HâCH3, HâCl, HâOH, HâF and ring CâN. These modifications were made systematically and comprehensively; only compounds that were too expensive to synthesize or that changed the charge state of the parent were left out. The analogues were tested for changes in affinity or potency, solubility, plasma protein binding, plasma stability, hepatic microsomal stability, permeability and hERG inhibition, versus their parent ligands. Relative affinity and in vitro PK changes were also evaluated computationally using FEP+ and ADMET prediction models. cpd, compound. b, For comparison with the experimental set, matched parentâanalogue pairs differing by one heavy-atom modification were identified in the ChEMBL database and analysed using the same activity-improvement framework. Figure created in BioRender; Xu, X. https://biorender.com/b6zqsv7 (2026).
These substitutions were systematically and comprehensively applied without design, ensuring an approach unbiased by anything other than synthetic accessibility and price. Overall, 257 single-atom analogues were synthesized and tested for changes in activity on the target, allowing us to measure how often we might expect affinity or potency to improve tenfold or more by minimal perturbation. This begins to provide a background expectation for how often a designed analogue might meaningfully improve affinity or potency. As we show, these undesigned changes had an unexpectedly high success rate, suggesting a systematic strategy for optimization, even though that was not our goal. As affinity or potency maturation is only one criterion for advancing from hit to candidate, we also measured in vitro PK properties for the analogues, including permeability, metabolic stability, plasma protein binding and three other terms, versus the parent compounds. These studies explore trade-offs between PK and target affinity or potency in ligand optimization.
We began with 18 parent ligands, most of which are unrelated to one another topologicallyâthat is, by their patterns of atoms and bonds (ECFP4 Tanimoto coefficients †0.35), with affinities ranging from 1 nM to 43 ”M and molecular weights ranging from 200 to 400 amu (Extended Data Fig. 2). Each of the 18 parents was modified one atom at a time with five types of single-atom substitutions (Fig. 1a, central panel), subject only to expense constraints (all molecules cost US$400 or less). About 7% of the possible molecules met these criteria and of those approximately 80% were successfully synthesized, resulting in 257 analogues. In the analyses that follow, we pooled all 257 analogues for analysis, improving statistical power (Supplementary Fig. 1).
The 257 analogues were tested for activity changes (dissociation constant (Kd), half-maximal inhibitory concentration (IC50) or half-maximal effective concentration (EC50)) versus the parents across the six targets. For ÎŒOR, EC50 values were determined in cAMP Glo-sensor assays. α2A and CB2 ligands were assessed by cAMP assays and radioligand competition binding, while SERT activity was assessed using a transporter uptake assay. AmpC inhibition was measured enzymatically and that of Mac1 was measured by ligand displacement. For every parent, we have only reported changes for one type of activity. For instance, if a parent inhibitory constant (Ki) is reported, then only Ki values are reported for its analogues, and if a parent is an agonist, only EC50 values are reported for its analogues. Each functional or inhibition assay included a literature positive-control molecule (Supplementary Fig. 2) with representative concentrationâresponse curves and Z values for each assay (Supplementary Fig. 3). The effects of receptor expression on agonism were controlled for (Supplementary Figs. 4 and 5). Full concentrationâresponse curves and analyses are shown for every parentâanalogue pair, allowing for direct inspection.
Overall, 23% of the analogues affected activity by twofold or less (fold changes rounded to the nearest integer; Supplementary Table 1) and another 30% reduced activity by tenfold or more (Extended Data Fig. 1b). Intriguingly, 29 of the 257 analogues (11.3%) improved activity by tenfold or more (Table 1). We note that four of these 29 had fold changes of between 9.7 and 9.9; we consider these effectively tenfold improvements and used the same convention in analysing the public ChEMBL data (see Table 1 and Supplementary Table 1 for details). Such tenfold or more improvement analogues were found for five of the six targets, and for 10 of the 18 parents (Table 1). The effects and chemical identities of all analogues versus their parents are listed in Supplementary Table 1.
We were interested in which perturbations had the biggest effects, how the rates of affinity or potency improvements compared with the literature, with its admitted biases, and if physical properties of the parents or of the binding sites correlated with greater likelihoods of improvement. Among the analogues, methyl substitutions were the most likely to improve activity tenfold or more, with chlorine substitutions the second most likely; for threefold or more improvements, of which there were 69 among the 257 analogues, both chlorine and methyl substitutions were the most common (Fig. 2a). Given the similar size and similar hydrophobicity of methyls and chlorines, their similar effects seem sensible. Conversely, fluorine substitutions, which are often used to increase metabolic stability, yielded only two tenfold or more and five threefold or more activity improvements after rounding fold changes to the nearest integer. Aromatic carbon-to-nitrogen substitutions also rarely improved affinity or potency, although they often had the greatest effects on PK properties.
Full size imagea, Observed percentages of compounds achieving threefold or more or tenfold or more activity improvement by single modification (CH3, Cl, F, N and OH). The open bars and points show observed percentages; the black error bars show two-sided 95% confidence intervals estimated by percentile bootstrap (10,000 resamples). Sample sizes were n = 72, 52, 39, 48 and 38 analogues for CH3, Cl, F, N and OH, respectively. b, Pearson productâmoment correlations between the parent-level rate of analogues achieving tenfold or more affinity or potency improvement and parent molecular weight (MW), parent pKi, binding-site volume and D-score (n = 18 parent compounds for each correlation). R denotes Pearsonâs correlation coefficient. P values are two-sided: P = 0.7104, 0.6096, 0.6937 and 0.1526 for MW, pKi, binding-site volume and D-score, respectively. c, The fraction of single-atom modifications (CH3, F, Cl, N and OH) that improve activity, comparing the experimental observations in this study with results from the ChEMBL database. The lines show the observed frequencies; the shaded regions indicate two-sided 95% confidence intervals estimated from 10,000 iterations of bootstrap resampling. The ChEMBL ligands are a composite of 54,021 parents and 79,984 small-change analogues; for only 4.7% of these were there more than three analogues per parent.
Because hydrophobicity increases will increase affinity simply by disfavouring the unbound state, and will have physical property liabilities, it is useful to consider the improved affinity or potency of the analogues in light of ligand efficiency and lipophilic efficiency (LipE). If affinity or potency improvements were largely driven by hydrophobicity, we would expect ligand efficiency to track with hydrophobicity (here, the calculated octanolâwater partition coefficient (clog P)) and LipE to remain constant or deteriorate. Instead, analogues that improved tenfold or more had ligand efficiencies that increased much more than hydrophobicity, and had substantially increased LipEs. For instance, an analogue with an added methyl, chloro or other atom that improved in binding or in EC50 by even threefold had a ligand efficiency of 0.65 kcal per atom for the added atom, well above the 0.3 kcal per atom thought to be useful for drug optimization29,30, whereas for analogues that improved activity tenfold or more, the single-atom ligand efficiency was 1.4 or more, close to the efficiency limit31. Similarly, every analogue that improved activity by more than tenfold also improved LipE (Extended Data Fig. 3f). More broadly, among the 220 analogueâparent pairs with complete numerical activity data, ÎclogP showed little quantitative correlation with ÎpActivity (Extended Data Fig. 3a).
We investigated whether there were correlations between likelihood of potency improvement and ligand and binding-site properties. For instance, one might expect that weaker ligands might be easier to optimize, or that smaller binding sites might be better suited to big affinity jumps. We calculated whether the parent molecular weight, binding-site volume, the ligandability of the binding sites (calculated using D-score32) or the affinity or potency of the starting parent compound was correlated with the likelihood of finding analogues with substantial affinity or potency improvement. Using the fraction of analogues achieving tenfold or more improvement as the outcome, none of these properties was meaningfully correlated (Fig. 2b). If one considers the magnitude of the affinity or potency change overall, and not just whether it improved tenfold or more, correlations did emerge, with weaker versus stronger parents more likely to support improved affinity (R = â0.34, P = 3.4 Ă 10â7) and with larger binding sites also tending to do so (R = 0.24, P = 4.0 Ă 10â4; Supplementary Fig. 6). Although these trends were echoed among matched pairs in ChEMBL (Fig. 2c and Extended Data Fig. 1), the correlations overall were weak. Together, these results suggest that: (1) as a strategy, systematic small perturbations can be unexpectedly successful at improving ligand affinity or potency23,24, perhaps even compared with design-heavy approaches33; (2) this effect is common across diverse parents and binding sites; (3) it can find echoes in the literature; and (4) random expectations for substantial affinity or potency improvement by small perturbation may be as high as 11% (95% CI 7.5â15.2%).
To investigate the structural bases for the affinity or potency changes, we determined the structure of the parent inhibitors in complex with Mac1 and six analogues from the â3122 series, 12 from the â3453 series and four from the â3794 series, each to between 0.97 and 1.02 Ă resolution (Extended Data Tables 1â4). The analogues with structures determined represented all six types of single-atom modifications in this study; 15 of the analogues differed in affinity or potency from their parents by threefold or more, two by tenfold or more and among themselves by up to 680-fold.
For some parents, analogues superposed well, facilitating analysis. For the Mac1 â3453 (IC50 of 8.9 ”M) inhibitor series (Extended Data Fig. 4a), the introduced chloro at the ortho-position of the phenyl ring (â0676, IC50 of 21 ”M) projects into solvent and loses activity versus the parent (Extended Data Fig. 4b). Moving this chloro to the meta-position (â1304), where it packs with Phe132, Ile131 and Gly48, improves inhibition 11-fold (P < 0.01) to 1.9 ”M (Extended Data Fig. 4c). Moving it one atom further over, to the para-position, had little effect (â6343, IC50 of 2.1 ”M), although a methyl at the same position (â9870, IC50 of 7.6 ”M) loses fourfold affinity (P < 0.01), presumably reflecting its poorer packing (Extended Data Fig. 4d).
Many of the effects of the small perturbation analogues could be explained post hoc from their structures; however, fewer were easily anticipated. Even the apparent simplicity of the â3453 series is belied by structural accommodations in the enzyme. For instance, the para-chloro of the 2 ”M â6343 would have clashed with the enzyme without Phe132 and Ile131 rotating away. In doing so, Phe132 adopts a partially eclipsed conformation. Presumably, the resulting strain is overcome by improved ligand packing, but this balance of terms is non-obvious (encouragingly, free energy calculations did correctly predict that the para-chloro and meta-chloro would improve affinity, whereas the ortho-chloro would lose it, although the agreement was more qualitative than quantitative; see below). Conversely, the methyl analogue (â9249, 53 ”M) of â0676 (IC50 of 22 ”M) is buried from solvent and might be expected to be more potent than its chloro cousin, but instead it loses another 2.5-fold in Ki (P < 0.01; here, the opposite effect was anticipated by the free energy calculation). Meanwhile, in the complex between â4905 and the α2AAR, an added methyl group fits into a pre-existing sub-pocket and qualitatively its improved potency is reasonable (Fig. 5a). Quantitatively, its 52-fold effect on EC50 is at the outer edge of what might be expected, especially as the methyl buries a serine hydroxyl. This is echoed by the free energy calculation, which here too correctly predicted improved activity but by twofold not 52-fold. These are examples of complexes where the analogues superpose on the parent and each other. In complexes in which the single-atom changes led to substantial inhibitor movement or to the adoption of multiple ligand conformations, prediction was harder still (Extended Data Fig. 4f,g). For example, the Mac1 methyl analogue â3176 repositions the inhibitors in the pocket (Extended Data Fig. 4g), whereas compounds â9249 and â3194 bind in two distinct conformations (Extended Data Fig. 4g), a phenomenon that is probably underreported in the Protein Data Bank34 but is readily seen at the ultra-high resolution of these structures. Together, this set of perturbations may provide an interesting and sometimes challenging test for prediction methods.
In optimizing chemical probes and leads, compound PK are as important as molecular target affinity or potency35,36. We thus explored how the small-perturbation analogues affected PK properties versus those of the parent ligands, and how these changes related to alterations of affinity or potency. We measured the following in vitro PK properties: metabolic stability in microsomes (half-life), plasma stability (half-life), plasma protein binding (fraction unbound), solubility (”M), hERG inhibition (IC50, ”M) and membrane permeability (parallel artificial membrane permeability assay (PAMPA); cm sâ1), for both the parent ligands and their small-change analogues (Fig. 3 and Supplementary Tables 2â7). These admittedly in vitro measurements are considered predictive of in vivo PK37,38,39. Some of the small-change analogues showed meaningful improvements relative to the parent ligands for each of these properties (Supplementary Table 8). For example, 12.4% of the analogues had threefold or more improvement in microsomal stability, 20.8% improved threefold or more in plasma stability, 13.8% improved solubility by threefold or more, 6.2% improved threefold or more in permeability, and 13.2% of the analogues improved threefold or more in plasma fraction unbound. Thus, each property can be meaningfully improved by simple changes. Correspondingly, a substantial fraction of the analogues experienced threefold or more decreases in these same properties (Extended Data Fig. 5). Considering what are perhaps the three most impactful in vitro PK measurements, microsomal stability, fraction unbound and permeability, 41.4% of the analogues experienced more than threefold losses in at least one of these, almost all (93.1%) experienced at least some deterioration by the same criterion, and 34.5% deteriorated in all three in vitro PK properties.
Full size imagea, Fold changes in PK properties for the 29 analogues classified as tenfold or more improved after integer rounding in affinity or potency versus a parent. b, Correlations between affinity or potency fold changes and those of the other PK properties. Correlations were assessed by two-sided Spearman rank correlation; exact pairwise-complete sample sizes (analogueâparent pairs) were: activity orpotency versus microsomal stability, permeability, plasma stability, fraction unbound, solubility and hERG IC50, n = 203, 203, 203, 193, 191 and 143, respectively; microsomal stability versus permeability, plasma stability, fraction unbound, solubility and hERG IC50, n = 240, 240, 213, 220 and 166, respectively; permeability versus plasma stability, fraction unbound, solubility and hERG IC50, n = 240, 213, 220 and 166, respectively; plasma stability versus fraction unbound, solubility and hERG IC50, n = 213, 220 and 166, respectively; fraction unbound versus solubility and hERG IC50, n = 200 and 143, respectively; and solubility versus hERG IC50, n = 154. *P < 0.05, **P < 0.01 and ***P < 0.001. Exact two-sided P values for all correlations are provided in the Source Data. c, Observed percentages of single-atom changes achieving threefold or more improvement in each parameter. The open bars and points show observed percentages; the black error bars show two-sided 95% confidence intervals estimated by percentile bootstrap resampling (10,000 resamples). Sample sizes by modification were n = 52, 39, 48, 72 and 38 analogues for chlorine, fluorine, nitrogen, methyl and hydroxyl, respectively.
Although many of the small-perturbation analogues substantially improved individual PK properties, none of the 29 analogues classified as tenfold or more improved after integer rounding showed improvements across all three key properties: metabolic stability, PAMPA permeability and plasma fraction unbound (Fig. 3a and Supplementary Fig. 7); many were worse in all three properties. Indeed, the changes in in vitro PK were, at best, orthogonal to affinity or potency fold change (Fig. 3b), and most were, if anything, anti-correlated with it, including plasma fraction unbound, microsomal stability, hERG inhibition and, to a minor extent, solubility. This is borne out at the level of individual substitutions: the substituent groups that most often improved affinity or potency substantially, such as methyl and chloro groups (Fig. 2a), were least associated with improved PK properties (Fig. 3c and Supplementary Fig. 8). Meanwhile, those changes least likely to improve affinity or potency substantially, such as aromatic CâN, were most likely to improve microsomal stability or fraction unbound (Supplementary Fig. 8). More broadly, most of the PK properties were largely independent of one another, with no correlations exceeding |Ï | > 0.5 and those that were statistically significant having Spearman Ï values ranging only from |0.16| to |0.38|. What emerges is a frustrated landscape for ligand optimization, with affinity or potency improvements counterbalanced by worsening PK. Although this trade-off has been well noted in medicinal chemistry40,41, this set can quantify its systematic impact.
The current benchmark only comprises several hundred analogues. If we could predict the results of the small perturbations, rather than having to experimentally test them, it could be much extended. We therefore asked how well modern methods could predict the affinity or potency and PK changes that we observed. In blinded experiments, we used free energy perturbation, arguably the highest level of theory available to the field, with the program FEP+ (Schrödinger) to predict the changes in activity between the parents and analogues (Fig. 4aâc). Similarly, we asked how well we could predict in vitro PK changes between the parents and analogues, using widely accessible tools (Fig. 4d).
Full size imagea, Comparison of FEP-predicted and experimentally measured binding free energies (ÎG). The dashed lines indicate ±1 kcal molâ1 and ±2 kcal molâ1 error margins. b, Correlation between FEP-predicted and experimentally measured relative binding free energy changes (ÎÎG). Dashed lines are as in panel a. c, An example of good structural alignment between the predicted (cyan) and experimental (light blue) poses, accompanied by a small error in the predicted ÎÎG. Dashed lines indicate hydrogen-bonding interactions. Abs., absolute; Exp., experimental. d, Heatmap of Spearman correlations between AI-predicted and experimentally measured fold changes across six in vitro PK properties, evaluated using three models. The colour intensity indicates correlation strength; correlations were assessed by two-sided Spearman rank correlation; pairwise complete-case n for Deep-PK, ADMET-AI and ADMETlab 3.0 were 218, 218 and 218, respectively, for solubility; 222 for the comparable ADMETlab 3.0 plasma-stability end point (the Deep-PK and ADMET-AI end points were not directly comparable); 222, 222 and 222, respectively, for permeability; 222, 222 and 222, respectively, for microsomal stability; 221, 221 and 221, respectively, for hERG; and 195, 195 and 195, respectively, for fraction unbound. Exact two-sided P values for all correlations are provided in the Source Data. **P < 0.01 and ***P < 0.001. n = 219 analogues (a,b).
There was a strong overall correlation between the FEP+-predicted and experimentally measured binding free energies (ÎG), with 188 of the 219 predictions falling within ±2 kcal molâ1 of the experimental values, and 129 within ±1 kcal molâ1, whereas the set had a mean unsigned error for ÎG of 1.06 kcal molâ1 (Fig. 4a,b and Supplementary Table 9). Of the 18 parents, the FEP+ predictions on 14 had ÎG mean unsigned error below 1.2 kcal molâ1 (Supplementary Table 10). If one considers the correlations qualitatively, of the 35 analogues predicted to lose affinity by 1 kcal molâ1 or more by FEP+ (that is, in the range of confident prediction), 20 did so experimentally. Meanwhile, of the 34 predicted to increase in affinity or potency by more than 1 kcal molâ1 by FEP+, 14 did so experimentally, whereas 12 were experimentally neutral and eight showed decreased potency. Categorization with a 1 kcal molâ1 threshold revealed a significant association between the direction and strength of FEP+ predictions and experimentally observed affinity or potency changes (Ï2 = 36.4 and P < 0.001; Supplementary Table 11).
Several challenges faced FEP predictions here, which future work may share. The predictive power was probably weakened by our inability to measure activity for weak analogues, reducing activity range and making it impossible to measure affinities predicted to be very low by FEP+. We used an automated workflow to generate initial poses for FEP; on inspection, most of the outliers reflected errors in these poses. FEP accuracy is impacted by the quality of the ligand-complexed structures42, and the predictions here began with docked poses, which are less reliable than experimental poses. For the 25 analogue structures that we determined by crystallography (for Mac1) and by cryo-electron microscopy (cryo-EM; for an α2A analogue, below), where predicted and experimental poses closely superimposed, FEP+ energies were largely consistent with experiment (Fig. 4c).
To predict in vitro PK43,44, we used ADMET-AI, ADMETlab 3.0 and Deep-PK45,46,47. The predictions of all three correlated with the experimental measurements of plasma fraction unbound and with permeability, whereas microsomal stability, plasma stability and solubility, which are difficult to predict48,49, were essentially uncorrelated. Correlations in the 0.4â0.5 range for fraction unbound and permeability may often be strong enough to guide compound selection and these open-access tools may have broad impact. Still, even with a correlation of 0.53 in fraction unbound, the highest that we observed, 28 analogues that had substantially increased fraction unbound experimentally were predicted to have lower levels than their parents, and 16 analogues that had higher protein binding than their parents experimentally were predicted to have less. Together, although the FEP and in vitro PK predictions can profitably model the effects of the small-perturbation analogues, their correlations with experiment are not yet strong enough to allow this benchmark to be extended by calculation alone.
Notwithstanding the frustrated landscape for ligand affinity or potency and PK improvement, medicinal chemists can often navigate this multi-parameter landscape to improve the properties of lead molecules. Here too, there were analogues that improved in affinity or potency with only a modest decrease in PK. Among these was an analogue of an α2A parent, compound â4905, where a single methyl group substitution improved potency 52-fold (Fig. 5a,b). In a cryo-EM structure of the â4905âα2AâG protein complex determined to 2.8 Ă resolution (Extended Data Table 5), â4905 superposed with the docking prediction for the parent â3629 and with the experimental structure of a previously described member of this family, â9087. This suggests that the methyl perturbation did not change the orientation of the analogue and that this family may be understood in the same structural context. A slight upwards deflection at high concentrations of â4905 in the cAMP doseâresponse curves may reflect secondary Gs coupling, supported by pertussis toxin-treated cAMP assays showing a residual, concentration-dependent increase in cAMP for â4905 but not â3629 (Supplementary Fig. 9a). Structurally, the added methyl group fills a hydrophobic sub-pocket (Fig. 5a and Supplementary Fig. 9b), helping to explain its improved potency.
Full size imagea, Docked parent â3629 superposed on the cryo-EM structure of the methyl analogue â4905. b, Concentrationâresponse curves for α2AAR activation by â3629 and â4905; the high-concentration upturn for â4905 may reflect Gs activation. Data are mean ± s.e.m. from n = 3 independent experiments, each with two technical replicates. câe, Tail-flick latencies after vehicle and â4905 (c), PS75 (d) or â3629 (e) at the indicated doses. f, Mechanical withdrawal threshold at baseline, after spared nerve injury (SNI), and after SNI + vehicle, â4905, â3629 or PS75 (5 mg kgâ1 each). g, Hotplate latency after vehicle, â3629 or â4905 (5 mg kgâ1). Mice were male C57BL/6J, 7â8 weeks of age; assignment was randomized and testing was blinded. For câg, data are mean ± s.d.; points are individual mice; and group sizes (left to right) were: n = 10, 10, 5, 5, 10 and 10 (c); n = 10, 5 and 10 (d); n = 10, 5, 5 and 5 (e); n = 10, 10, 12, 5, 5 and 5 (f); n = 10, 10 and 10 (g). Tests were two-way analysis of variance (ANOVA) with Dunnettâs test (c), KruskalâWallis with Dunnâs test (d,e), two-way ANOVA with Tukeyâs test (f) and Friedman with Dunnâs test (g). Exact P values: vehicle versus increasing â4905 doses of 0.9759, 0.9841, 0.7141 and <0.0001 (c); vehicle versus PS75 of 0.0301 and 0.0043 (d); vehicle versus â3629 of >0.9999, 0.0393 and 0.0164 (e); SNI versus SNI + vehicle of 0.7416, SNI + vehicle versus SNI + â4905 of 0.0037, SNI + vehicle versus SNI + PS75 of 0.0333, SNI + â4905 versus SNI + â3629 of 0.0188 (f); vehicle versus â3629, vehicle versus â4905 and â4905 versus â3629 of 0.0417, <0.0001 and 0.2209, respectively (g). *P < 0.05, **P < 0.01, ****P < 0.0001 and not significant (NS).
Although the improved potency of â4905 came at a cost to its in vitro PK, the effects were modest: compared with its parent, â3629, the fraction unbound decreased by 57% and the microsomal half-life was reduced by 41% (Supplementary Tables 2â7). The in vivo PK of â4905 qualitatively reflected the in vitro results, with cerebrospinal fluid concentrationsâa proxy for fraction unbound in the brain50âreduced by 36%, whereas half-life was reduced from 196 to 15.2 min (Supplementary Table 12). Accordingly, we compared the in vivo analgesia of â4905 to that of its parent â3629 and to PS75, the most potent analgesic known for this family of α2A agonists51. Despite the deterioration in PK, in mouse nociception assays including tail flick, neuropathic pain (SNI) and hotplate, â4905 was far more potent than the parent â3629 and at least as potent as PS75 (Fig. 5câg). If ligand optimization is a frustrated landscape, it remains a navigable one41. In addition to beginning to provide a random background for such efforts, this study supports a strategy by which that landscape may be reconnoitered23,24.
Although ligand optimization is at the heart of drug discovery, quantifying the success of optimization strategies has been difficult owing to a lack of a random background. Four key observations emerge from this study. First, systematic and unbiased small perturbations to 18 scaffolds revealed that 11.3% of the analogues improved tenfold or more in affinity or potency versus the parent (Table 1); âmagic methylsâ25,52 were common in this set. These improvements were observed in five of six targets, across ligand physical properties and across five orders of magnitude of parent activity. This begins to establish a background expectation for affinity or potency optimization against which methods and designs may be compared. It may also suggest a simple, unbiased approach for achieving substantial jumps in affinity or potency, improvements that will not always be obvious from structural analyses or simulations (Figs. 4 and 5 and Extended Data Fig. 4) but may complement them. Second, improved affinity or potency typically came at the cost of ligand PK, with plasma fraction unbound, stability to liver microsomes and permeability, among other terms, declining as analogue affinity or potency increased. The individual in vitro PK terms were orthogonal to each other and several were anti-correlated with affinity or potency improvement (Fig. 3). This suggests a frustrated landscape for ligand optimization. Although the qualitative challenges and trade-offs in ligand optimization are known to medicinal chemists8,40,41, this study has begun to quantify them across a range of target classes and ligand properties. Third, computational prediction of affinity or potency and PK had some success anticipating the trends that we observed (Fig. 4). Still, their correlations with experiment remained loose enough to preclude replacing systematic exploration and testing of molecules as this background set is expanded, something seen with other potency prediction methods53. Fourth, although potency trades off against PK, it is possible to find analogues where it rises sufficiently to overcome drops in exposure and half-life. Thus, the α2A analogue â4905 is the most potent analgesic in a large series, and is far more potent in vivo than its parent, despite experiencing these trade-offs.
At first glance, this study may recall classical medicinal chemistry where conservative and ligand-based approaches such as Topliss Trees11 dominated the field. However, this would be to misremember the classic age of medicinal chemistry on two counts. First, the systematic small perturbations explored here were rarely used then, partly because the building blocks and reactions to support them were lacking. The ability to systematically explore small perturbations across scaffolds reflects advances in synthetic chemistry and the new building-block approaches that have expanded readily available compounds54,55. Second, the logic behind Topliss Trees and other quantitative structureâactivity approaches11,12 was anchored in ligand physical chemistry without reference to receptor structure. For instance, in Topliss Trees the success of methyl and chloro derivatives taught opposite lessons because of their different electronic effects on the ligand, whereas the steric and non-polar properties of the two groups will seem similar from the view âinside the receptorâ52,56. Meanwhile, the role of a random background would have seemed foreign. Thus, although a strategy of systematic small perturbations may smack of the âmethyl, ethyl, propyl, butyl, futileâ approach with which classic medicinal chemistry has been tarred57,58, it has a different basis in theory and reflects new synthetic opportunities.
Certain limitations of this study merit airing. At 257 analogues and six targets, the set remains small and inevitably, if unintentionally, biased. Expanding to additional targets and ligand classes might reveal different trends. We note that the six targets do recognize a range of chemotypes, including large polar co-factors (Mac1), large hydrophobic lipids (CB2), anionic ÎČ-lactams (AmpC), cationic neurotransmitters (α2A and SERT) and peptides (”OR). Experimentally, we assessed activity changes using a transporter uptake assay (SERT), radioligand competition binding assays (α2A and CB2), second messenger assays (α2A, CB2 and MOR), and binding and enzyme activity (Mac1 and AmpC) assays. Although ligand competition correlates with Ki, agonist EC50 values are affected by receptor expression. Because we compared the relative activities of analogues and parents, controlled for receptor expression (Supplementary Fig. 4), and measured effects at low expression levels (Supplementary Fig. 5), the effects of receptor expression on relative activity should be modest, although they cannot be completely discounted. For these and related reasons, changes of EC50 between agonist parent and analogues cannot be read as changes in affinity the way that changes in Ki can be, although EC50 changes remain the relevant metric for agonist activity. Overall, the impact of these effects may be inspected in full concentrationâresponse curves for parentâligand pairs with more than threefold activity improvement (Supplementary Fig. 2). In the area of in vitro PK, our use of liver microsomes rather than hepatocytes to measure metabolic stability meant that some types of metabolism were missed, including glucuronidation. This might make groups such as phenolic hydroxyls seem less labile than they would be in vivo. More broadly, for all but two molecules, we only measured in vitro not in vivo PK. Whereas the latter are widely used in ligand optimization, their prediction of in vivo behaviour is only approximate.
These limitations should not obscure the main observations of this study. Overall, 11.3% of systematic and unbiased small perturbations improved analogue affinity or potency tenfold or more, beginning to establish a background expectation for the likelihood of substantial ligand affinity or potency improvement and a systematic approach to doing so23,24. Balancing this was a concomitant deterioration in ligand PK, which will lower the exposure and half-life of a ligand in vivo, counteracting the improvements in affinity or potency. Navigating this multi-parameter space is at the heart of medicinal chemistry; this study supports the development of quantitative models to do so.
From the PostgreSQL version of CHEMBL34 (ref. 4) (https://doi.org/10.6019/CHEMBL.database.34), we filtered the database for compounds with activities reported against a single protein, with either Ki, Kd, IC50 or EC50 as the activity type. To be considered a molecular pair, compounds had to have the same assay ID, the same target ID, the same activity type and the same reference document (publication) ID, in addition to having a parentâanalogue relationship as defined in this work, meaning a CâH group replaced by CâOH, CâF, CâCl, CâBr, CâCH3 or an aromatic carbon replaced by a nitrogen. This left us with 191,732 parentâanalogue pairs.
Eighteen parent compounds were selected from previously published literature51,59,60 or datasets (https://asapdiscovery.org/outputs/molecules/#ASAP-SARS-COV-2-NSP3-MAC1) based on their known binding activity, structural relevance or representation of diverse chemical scaffolds. Starting from a parent compound, we used RDKit (http://www.rdkit.org/) to identify all CâH bonds and iteratively replaced the hydrogen atom with a methyl, hydroxyl, fluoro, chloro or bromo group. We also identified aromatic carbons with two heavy-atom neighbours and replaced them with nitrogen. Every analogue generated was represented as a canonical isomeric SMILES string and added to a set to remove duplicates. For each of the 18 parents, the full set of possible analogues that could be synthesized for less than US$400 for 10 mg was ordered.
When reported, 95% confidence intervals were derived from bootstrap resampling with 10,000 iterations. In cases in which the observed frequency of success is exactly 0, bootstrap will fail to give an upper bound for the interval. In such cases, we used the ClopperâPearson method to estimate an upper bound based on the sample size61. Pearson and Spearman correlations with their associated P values were computed using the scipy.stats module from the SciPy package62.
These were conducted using FEP+ within the Schrödinger software suite (v2025-2) with the OPLS4 force field63 and the modified simple point charge water model. The default setting was used for the number of lambda windows selection where it depends on the type of perturbations; charge-changing, core hopping and all other perturbations have 24, 16, and 12 lambda windows, respectively. For alchemical transformations with charge changes, the total charge of the simulation box was kept constant by transmuting a Na+ or Clâ ion to water or vice versa. In addition, a 0.15 M concentration of NaCl was added to the simulation box of charge-changing perturbations. For α2A, CB2 and SERT, the FEP+ membrane protocol was applied where a POPC membrane was added to the system in simulation. All other settings were kept default except that the simulation time was extended from 5 ns to 10 ns.
The default FEP map generation protocol was used with the parent compound selected as the biased node. In preparing proteins and ligands for FEP+, the Schrödinger protein preparation workflow and LigPrep were used. The initial binding poses of parent compounds were from poses generated by DOCK3.8, and analogues were aligned to the parents with severe steric clashes resolved using the FEP+ Pose Builder workflow.
Whereas no structural information was used in the design of analogues from parent compounds, for FEP+ calculations and for post hoc structural analysis, we generated ligand-bound complexes of the parents and relevant ligands using DOCK.3.8 (refs. 54,64). Ligands were docked into the receptor-binding site using grids prepared in previous studies51,59,60,65,66.
We used one consistent readout per parent series (no mixing within a series). SERT and AmpC are reported as Ki; Mac1 is reported as IC50 from peptide displacement (a scalable functional hydrolysis assay is unavailable); GPCR agonist series are reported as EC50 and antagonist series as Ki, such that ÎŒOR is EC50 only, whereas α2A and CB2 include EC50 and Ki depending on the parent series. Representative concentrationâresponse curves and the corresponding Z values for each assay are provided in Supplementary Fig. 3.
SERT activity was measured using the Neurotransmitter Transporter Uptake Assay Kit from Molecular Devices (#R8174), following the manufacturerâs protocol with slight modifications as previously described65. HEK293T cells (ATCC CRL-3216) stably expressing human SERT were plated in poly-L-lysine-coated 384-well black, clear-bottom plates at a density of 15,000 cells in 40 ”l per well, using DMEM supplemented with 1% dialysed FBS. Cells were incubated overnight at 37 °C with 5% CO2 to allow adherence and recovery. The following day, the medium was removed by flicking, and cells were incubated with 25 ”l per well of test compound solutions prepared in assay buffer (1Ă HBSS, 20 mM HEPES, pH 7.4, supplemented with 1 mg mlâ1 BSA) for 30 min at 37 °C. After drug treatment, 25 ”l per well of dye solution (as provided in the kit) was added directly to the wells, followed by an additional 30-min incubation at 37 °C. Fluoxetine (10 ”M) was used as a positive control for SERT inhibition. Fluorescence was measured using the FlexStation II microplate reader with excitation at 440 nm and emission at 520 nm. Relative fluorescence units were exported and analysed using GraphPad Prism 10.0 to calculate IC50 values, from which Ki values were derived using the ChengâPrusoff equation.
CB2 receptor-binding assays were performed using membrane preparations from HEK293 cells (ATCC CRL-1573) stably expressing human CB2, following previously published methods67,68. Membranes were resuspended in TME buffer containing 0.1% BSA (w/v) and 25 ”g of membrane protein was added per well. The assay was conducted using the radioligand [3H]CP-55,940 at a final concentration of 0.75 nM, prepared in assay buffer. Nonspecific binding was defined in the presence of 5 ”M unlabelled CP-55,940. Test compounds were applied at increasing concentrations to assess competition. Reactions were incubated at 30 °C for 1 h with gentle shaking. After incubation, samples were transferred to Unifilter GF/B 96-well filter plates and filtered using a Packard Filtermate-196 cell harvester (PerkinElmer). Plates were washed four times with ice-cold wash buffer (50 mM Tris-HCl, 5 mM MgCl2 and 0.5% BSA, pH 7.4). Radioactivity bound to the filters was quantified via liquid scintillation counting. Specific binding was calculated by subtracting nonspecific binding from total binding. IC50 and Ki values were calculated using nonlinear regression in GraphPad Prism 9 using the ChengâPrusoff equation.
α2AAR binding was performed using membrane preparations from insect cells (Expression Systems, 94-001S) expressing human α2A receptors, as previously described51. Membranes were incubated with increasing concentrations of test compounds and 5 nM [3H]rauwolscine in buffer containing 20 mM HEPES (pH 7.5) and 100 mM NaCl, at room temperature for 2 h. After incubation, samples were filtered onto GF/B filter plates, washed with ice-cold buffer and radioactivity was quantified by liquid scintillation counting. IC50 and Ki values were derived using nonlinear regression in GraphPad Prism.
The GloSensor cAMP assay was performed following the manufacturerâs instructions (Promega) with slight modifications as previously described51,69. In brief, wild-type human α2A, CB2 and ÎŒOR were cloned into the pcDNA3.1 vector and co-transfected with the 22F cAMP GloSensor plasmid into HEK293T cells (ATCC CRL-3216) cultured in six-well plates. After 24 h, cells were reseeded into 96-well white plates in CO2-independent medium and equilibrated with GloSensor cAMP reagent as per the manufacturerâs protocol. Cells were incubated for 1 h at 37 °C followed by 1 h at room temperature. Where applicable, 10 ÎŒM forskolin was used to elevate basal cAMP levels for assessing receptor-mediated inhibition. Serially diluted test compounds were added, and luminescence signals were recorded using a PerkinElmer microplate reader. Data were analysed using GraphPad Prism 9.0 to calculate EC50 or IC50 values.
The AmpC ÎČ-lactamase inhibition assay was performed as previously described60. The candidate inhibitors were dissolved in DMSO (20 mM stock) and diluted to maintain a constant 1% DMSO (v/v) in 50 mM sodium cacodylate buffer (pH 6.5). Assays were performed in the presence of 0.01% Triton X-100 to reduce aggregation artefacts. AmpC enzymatic activity was monitored spectrophotometrically using CENTA or nitrocefin as substrates. Initial screening was performed at 200 ”M, 100 ”M and 40 ”M compound concentrations. Substrate concentrations were selected based on known Km values to achieve defined [S]:Km ratios: for CENTA ([S] = 50 ”M, Km = 27.6 ”M) and nitrocefin ([S] = 100 ”M or 28 ”M, Km = 180 ”M). Reactions were carried out in 96-well format on a BMG Labtech CLARIOstar plate reader, with substrate and enzyme injected into wells containing the inhibitor, followed by kinetic measurement over 50 s. IC50 values were determined by fitting inhibition curves in GraphPad Prism using a fixed Hill coefficient of 1, and Ki values were calculated using the ChengâPrusoff equation70.
Binding of the compounds to Mac1 was assessed by the displacement of an ADPr-conjugated biotin peptide from His6-tagged protein using a homogeneous time-resolved fluorescence (HTRF)-based assay, as previously described71. The expression sequences used for SARS-CoV-2 Mac1 are listed below. All proteins were expressed and purified as previously described for SARS-CoV-2 Mac1 (ref. 71). Compounds were dispensed into ProxiPlate-384 Plus (PerkinElmer) assay plates using an Echo 650 Liquid Handler (Beckman Coulter). Binding assays were conducted in a final volume of 16 ÎŒl with 12.5 nM NSP3 Mac1 protein, 200 nM peptide ARTK(Bio)QTARK(Aoa- RADP)S (Cambridge Peptides), 1:20,000 anti-His6-Eu3+ cryptate (HTRF donor; PerkinElmer AD0402) and 1:500 streptavidin-XL665 (HTRF acceptor; PerkinElmer 610SAXLB) in assay buffer (25 mM HEPES pH 7.0, 20 mM NaCl, 0.05% bovine serum albumin and 0.05% Tween-20, the latter also to reduce aggregation artefacts). Assay reagents were dispensed manually into plates using an electronic multichannel pipette. Mac1 and peptide were pre-incubated for 30 min at room temperature before HTRF reagents were added. Fluorescence was measured after a 1-h incubation at room temperature using a Perkin Elmer EnVision 2105-0010 Dual Detector Multimode microplate reader with dual emission protocol (A = excitation of 320 nm and emission of 665 nm, and B = excitation of 320 nm and emission of 620 nm). Compounds were tested in triplicate in a 14-point dose response. Raw data were processed to give an HTRF ratio (channel A/B Ă 10,000), which was used to generate IC50 curves using nonlinear regression using GraphPad Prism v10.0.2 (GraphPad Software).
hERG channel inhibition was evaluated using a Thallium Flux assay on HEK293 cells (ATCC CRL-1573) stably expressing the hERG potassium channel. Cells were seeded at a density of 8,000 cells per well in 384-well poly-D-lysine-coated plates and incubated for 24 h under standard conditions (37 °C at 5% CO2). The next day, a thallium-sensitive dye was added to the cells, followed by a 1-h incubation to ensure dye uptake. Test compounds were added to achieve a final concentration of 30 ÎŒM in 0.5% DMSO, and the cells were incubated for an additional 30 min at room temperature. Subsequently, a stimulation buffer containing thallium was added, and fluorescence measurements were taken using a FLIPR Tetra system. Data were collected every 3 s for 3 min (excitation of 470â495 nm and emission of 515â575 nm). Fluorescence intensity over time was analysed to calculate the area under the curve from which percentage inhibition was determined versus haloperidol at 100 ÎŒM (positive control) and vehicle (DMSO).
Wild-type human α2AAR was cloned into a pVL1392 vector with an N-terminal FLAG tag. The construct was expressed in Spodoptera frugiperda (Sf9) insect cells using the BestBac system. Cells at a density of 4 Ă 106 cells per millilitre were infected with virus and incubated for 48 h at 27 °C. The receptor was solubilized and purified by FLAG affinity chromatography and size-exclusion chromatography in the presence of 10 ÎŒM compound â4905. Monomeric peak fractions were concentrated and used for G protein complex formation. GαoGÎČ1Îł2 heterotrimeric G proteins were expressed in Hi5 insect cells (Expression Systems, 94-002S) and purified using Ni2+ affinity following detergent solubilization and dephosphorylation. The final α2AARâGαoGÎČ1Îł2âscFv16 complex72 was assembled in the presence of â4905 and purified by size-exclusion chromatography. Cryo-EM grids were prepared using UltrAufoil R1.2/1.3 300-mesh grids and vitrified in liquid ethane. Data were collected on a Titan Krios G3 electron microscope equipped with a K3 direct electron detector. Image processing was performed using cryoSPARC, yielding a final reconstruction at approximately 2.8 Ă resolution. Model building and refinement were performed with Protein Data Bank (PDB) 7EJ8 as a starting model using ChimeraX73, Phenix74 and Coot75. The final structures have been deposited in the PDB with accession codes 9PLO and 9PLN. Structural representations of proteinâligand complexes were generated using the PyMOL Molecular Graphics System v2.5.5 (Schrödinger).
Wild-type Mac1 protein (P43 construct, residues 3â169) was expressed in Escherichia coli BL21(DE3) as an N-terminal His6-tagged construct and purified by Ni2+-affinity chromatography71. The His-tag was cleaved with TEV protease, followed by size-exclusion chromatography (Superdex 75) in 20 mM Tris-HCl (pH 7.5), 150 mM NaCl and 1 mM dithiothreitol. Purified protein was concentrated to 40 mg mlâ1 for crystallization and stored at â80 °C.
Crystals were obtained by sitting-drop vapour diffusion in 28% PEG 3000 and 100 mM CHES (pH 9.5). Compounds (100 mM in DMSO) were added to crystal drops using an Echo 650 acoustic dispenser76 to a final concentration of 10 mM. Crystals were incubated at room temperature for 2â4 h and vitrified in liquid nitrogen without additional cryoprotection. X-ray diffraction data were collected at 100 K at the Advanced Light Source (beamline 8.3.1) using an X-ray wavelength of 0.88557 Ă , and diffraction data were processed using XDS77 and Aimless78. Structures were determined to resolutions ranging from 0.97 to 1.02 Ă . Ligands with low occupancy or conformational disorder were modelled using PanDDA79 and Coot75 and refined using phenix.refine80 as previously described81. The structure of â3184 was determined using a racemic preparation of the compound (â9037). X-ray data collection and refinement statistics are shown in Extended Data Tables 1â4, as are the 32 PDB IDs. Structural representations of proteinâligand complexes were generated using the PyMOL Molecular Graphics System v2.5.5 (Schrödinger).
Microsomal stability of compounds was evaluated using pooled mouse liver microsomes (M3000/lot #2010026, XenoTech) to estimate their metabolic stability and predict hepatic clearance. Each compound was incubated at 2 ÎŒM in a reaction mixture containing 0.42 mg mlâ1 microsomal protein, phosphate buffer (100 mM, pH 7.4), MgCl2 (3.3 mM), NADPH (3 mM), glucose-6-phosphate (5.3 mM) and glucose-6-phosphate dehydrogenase (0.67 units per millilitre). Reactions were conducted at 37 °C in 96-well plates with shaking at 100 rpm. Samples were collected at five time points (0, 7, 15, 25 and 40 min), and reactions were quenched by adding five volumes of acetonitrile containing an internal standard. After centrifugation at 5,500 rpm for 5 min, the supernatants were analysed via high-performance liquid chromatographyâtandem mass spectrometry (HPLCâMS/MS). The elimination rate constant (kel), half-life (t1/2) and intrinsic clearance (Clint) were calculated by plotting the natural logarithm of the remaining parent compound versus time. Stability was compared with reference standards such as imipramine and propranolol.
Plasma protein binding was measured using equilibrium dialysis with a 14-kDa molecular weight cut-off membrane in a 96-well HTD96b dialyser. Mouse plasma containing 1 ÎŒM test compound (0.005% DMSO and 1% acetonitrile) was placed in one chamber, and phosphate-buffered saline (PBS, pH 7.4) in the opposing chamber. The assembled plates were incubated at 37 °C with 5% CO2 and approximately 95% humidity, shaking at 250 rpm for 5 h to reach equilibrium. After incubation, aliquots from each chamber were mixed with equal volumes of the blank opposite matrix and processed with acetonitrile containing internal standard. Supernatants obtained after centrifugation were analysed by HPLCâMS/MS. The percentage of compound bound to plasma proteins was calculated using the peak area ratio in buffer to plasma compartments. Recovery and stability standards were included to ensure accuracy and reliability. Verapamil served as a reference control. Most compounds showed moderate binding, typically ranging between 75% and 85%.
Plasma stability was assessed in non-sterile mouse plasma (Li-heparin treated) at 1 ÎŒM concentration (final DMSO content was 0.005%). Incubations were carried out in aliquots of 60 ÎŒl each (two per time point) at 37 °C under 5% CO2 and high humidity (approximately 95%). The reactions were quenched with 240 ”l of 90% acetonitrile containing an internal standard at 0, 20, 40, 60 and 120 min, followed by centrifugation at 6,000 rpm for 5 min. Supernatants were analysed via HPLCâMS/MS to determine the percentage of parent compound remaining at each time point. Data were plotted to calculate compound half-lives (t1/2). Reference compounds, verapamil and propantheline, were used as high and low stability controls, respectively. This assay is critical for identifying compounds susceptible to degradation by plasma esterases or hydrolytic enzymes, helping inform PK optimization strategies during lead selection.
Aqueous thermodynamic solubility was determined in PBS (pH 7.4) for 247 compounds using a shake-flask method followed by UV absorbance quantification. Dry powder compounds were dissolved in PBS to a theoretical concentration of 4 mM and incubated in duplicate at 25 °C for 4 h and 24 h with shaking. After incubation, samples were filtered using HTS 96-well filter plates. The filtrates were diluted twofold in acetonitrile with 4% DMSO for UV analysis. The incubation samples for charged molecules were additionally diluted tenfold with 50% acetonitrileâPBS with 2% final DMSO. Calibration curves (0â200 ÎŒM) were prepared in 50% acetonitrileâPBS (2% final DMSO). Absorbance was measured between 230 and 550 nm using a SpectraMax Plus microplate reader. Compound-specific absorbance maxima were used to calculate concentration using SoftMax Pro and Excel. The assay dynamic range was approximately 2â400 ÎŒM (about 20â4,000 ÎŒM for charged molecules), with values near the upper limit treated as semi-quantitative. Ondansetron was used as a reference compound. This method reflects equilibrium solubility under physiologically relevant conditions and helps to rank compounds for formulation feasibility.
Passive bloodâbrain barrier (BBB) permeability was estimated using PAMPA (PAMPA-BBB) with a phospholipid-coated membrane simulating the brain endothelium. Test compounds (50 ÎŒM in Prisma HT buffer, pH 7.4, with 0.5% DMSO) were added to donor wells, whereas brain sink buffer was added to the acceptor wells. The donor and acceptor chambers were separated by a 0.45-”m filter membrane coated with brain polar lipids. Plates were incubated without agitation at room temperature for 4 h. Post-incubation, samples from both chambers, as well as a standard solution, were diluted with acetonitrile containing an internal standard. Apparent permeability coefficients (log Papp) were calculated based on peak area ratio. Clozapine and chlorpromazine were used as high-permeability controls, whereas ranitidine represented low permeability. The assay provides a high-throughput, non-cell-based method for estimating central nervous system exposure potential.
Male C57BL/6J mice 7â8 weeks of age were obtained from The Jackson Laboratory (JAX strain #000664) and housed five per cage under a standard 12â12-h lightâdark cycle at 20.7 °C (69.3 °F) and 58% humidity. All animal procedures were approved by the University of California, San Francisco Institutional Animal Care and Use Committee (protocol #AN208219). Animals were randomly assigned to treatment and control groups. For behavioural experiments, mice were initially placed together in a cage and allowed to move freely for a few minutes; each mouse was then randomly selected, injected with compound or vehicle, and placed in a separate cylinder before testing. The experimenter assessing behaviour was different from the experimenter administering injections and placing mice into the cylinders; experimenters were therefore blinded to treatment. Behavioural experiments were conducted in two independent cohorts. Mice were habituated individually in Plexiglas enclosures for 1 h before testing. Compounds were administered subcutaneously 30 min before behavioural assessment, and, where applicable, the α2AAR antagonist atipamezole (2 mg kgâ1, intraperitoneally) was given 15 min before compound injection. Tail-flick latency was measured by immersing the distal third of the tail in a 50 °C water bath and recording the withdrawal time. For the neuropathic pain model, SNI was performed under isoflurane anaesthesia by ligating and transecting two of the three branches of the sciatic nerve, sparing the sural nerve. Mechanical thresholds were assessed 7â14 days post-surgery using von Frey filaments and the upâdown method, and values were normalized to the baseline of each animal. Thermal nociception was evaluated using a 55 °C hotplate, and the latency to nocifensive behaviour (paw lick or jump) was recorded with a cut-off of 45 s to prevent tissue damage. Behavioural data are summarized as group size (n), mean and s.d. (Supplementary Table 13). Statistical analyses were performed in GraphPad Prism. For the SNI experiment (Fig. 5f), we used two-way ANOVA followed by Tukeyâs multiple-comparisons test. The hotplate assay (Fig. 5g) was analysed using the Friedman test followed by Dunnâs multiple-comparisons test. For the tail-flick assays, the 4905 doseâresponse (Fig. 5c) was analysed by two-way ANOVA followed by Dunnettâs multiple-comparisons test against vehicle (control), whereas the PS75 and 3629 dose groups (Fig. 5d,e) were analysed using the KruskalâWallis test followed by Dunnâs multiple-comparisons test against vehicle. The KruskalâWallis and Friedman tests were nonparametric; Prism calculated two-sided, multiplicity-adjusted P values for the Dunn comparisons (Supplementary Table 13). For sample size, we did not perform formal a priori power calculations. Instead, group sizes were guided by our previous experience with these assays and by precedent in the literature. This approach may limit sensitivity to small effects. Raw data for the animal assays underlying Fig. 5câg are provided in the Source Data file.
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
All primary data are available through this article and its Extended Data and Supplementary Information. The cryo-EM structures and maps have been deposited in the PDB and Electron Microscopy Data Bank, respectively, under the accession codes 9PLO and EMD-71719 (the â4905-bound α2AAR complex with G proteins and scFv16), and PDB 9PLN and EMD-71718 (the â4905-bound locally refined α2AAR structure). The Mac1 crystal structures and supporting structure factors have been deposited in the PDB under the accession codes 14AB, 7IIW, 7IIX, 7IIY, 7IIZ, 7IJ0, 7IJ1, 7IJ2, 7IJ3, 7IJ4, 7IJ5, 7IJ6, 7IJ7, 7IJ8, 7IJ9, 7IJA, 7IJB, 7IJC, 7IJD, 7IJE, 7IJF, 7IJG, 7IJH, 7IJI, 7IJJ, 7IJK, 7IJL, 14AM, 14AN, 14AO, 14AP and 7IJM. Source data are provided with this paper.
ZINC tools used in the selection of the small-perturbation analogues are openly available (https://zinc22.docking.org).
Hughes, J. P., Rees, S., Kalindjian, S. B. & Philpott, K. L. Principles of early drug discovery. Br. J. Pharmacol. 162, 1239â1249 https://doi.org/10.1111/j.1476-5381.2010.01127.x (2011).
Article CAS PubMed PubMed Central Google Scholar
Bleicher, K. H., Bohm, H. J., Muller, K. & Alanine, A. I. Hit and lead generation: beyond high-throughput screening. Nat. Rev. Drug Discov. 2, 369â378 https://doi.org/10.1038/nrd1086 (2003).
Gaulton, A. et al. ChEMBL: a large-scale bioactivity database for drug discovery. Nucleic Acids Res. 40, D1100âD1107 https://doi.org/10.1093/nar/gkr777 (2012).
Zdrazil, B. et al. The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Res. 52, D1180âD1192 https://doi.org/10.1093/nar/gkad1004 (2024).
Article CAS PubMed PubMed Central Google Scholar
Paul, S. M. et al. How to improve R&D productivity: the pharmaceutical industryâs grand challenge. Nat. Rev. Drug Discov. 9, 203â214 https://doi.org/10.1038/nrd3078 (2010).
Lipinski, C. A. Lead- and drug-like compounds: the rule-of-five revolution. Drug Discov. Today Technol. 1, 337â341 https://doi.org/10.1016/j.ddtec.2004.11.007 (2004).
Stumpfe, D., Hu, H. & Bajorath, J. Evolving concept of activity cliffs. ACS Omega 4, 14360â14368 https://doi.org/10.1021/acsomega.9b02221 (2019).
Article CAS PubMed PubMed Central Google Scholar
Leeson, P. D. & Springthorpe, B. The influence of drug-like concepts on decision-making in medicinal chemistry. Nat. Rev. Drug Discov. 6, 881â890 https://doi.org/10.1038/nrd2445 (2007).
Wunberg, T. et al. Improving the hit-to-lead process: data-driven assessment of drug-like and lead-like screening hits. Drug Discov. Today 11, 175â180 https://doi.org/10.1016/S1359-6446(05)03700-1 (2006).
Jorgensen, W. L. Efficient drug lead discovery and optimization. Acc. Chem. Res. 42, 724â733 https://doi.org/10.1021/ar800236t (2009).
Article CAS PubMed PubMed Central Google Scholar
Topliss, J. G. Utilization of operational schemes for analog synthesis in drug design. J. Med. Chem. 15, 1006â1011 https://doi.org/10.1021/jm00280a002 (1972).
Hansch, C. & Fujita, T. Ï-Ï-Ï Analysis: a method for the correlation of biological activity and chemical structure. J. Am. Chem. Soc. 86, 1616â1626 (1964).
Dong, J. et al. ADMETlab: a platform for systematic ADMET evaluation based on a comprehensively collected ADMET database. J. Cheminform. 10, 29 https://doi.org/10.1186/s13321-018-0283-x (2018).
Article CAS PubMed PubMed Central Google Scholar
Wang, L. et al. Accurate and reliable prediction of relative ligand binding potency in prospective drug discovery by way of a modern free-energy calculation protocol and force field. J. Am. Chem. Soc. 137, 2695â2703 https://doi.org/10.1021/ja512751q (2015).
Klimovich, P. V., Shirts, M. R. & Mobley, D. L. Guidelines for the analysis of free energy calculations. J. Comput. Aided Mol. Des. 29, 397â411 https://doi.org/10.1007/s10822-015-9840-9 (2015).
Article ADS CAS PubMed PubMed Central Google Scholar
Hajduk, P. J. & Sauer, D. R. Statistical analysis of the effects of common chemical substituents on ligand potency. J. Med. Chem. 51, 553â564 https://doi.org/10.1021/jm070838y (2008).
Kramer, C., Kalliokoski, T., Gedeck, P. & Vulpetti, A. The experimental uncertainty of heterogeneous public data. J. Med. Chem. 55, 5165â5173 https://doi.org/10.1021/jm300131x (2012).
Booker, T. R., Jackson, B. C. & Keightley, P. D. Detecting positive selection in the genome. BMC Biol. 15, 98 https://doi.org/10.1186/s12915-017-0434-y (2017).
Article CAS PubMed PubMed Central Google Scholar
Pritchard, J. K. & Cox, N. J. The allelic architecture of human disease genes: common diseaseâcommon variant⊠or not? Hum. Mol. Genet. 11, 2417â2423 https://doi.org/10.1093/hmg/11.20.2417 (2002).
Manolio, T. A. et al. Finding the missing heritability of complex diseases. Nature 461, 747â753 https://doi.org/10.1038/nature08494 (2009).
Article ADS CAS PubMed PubMed Central Google Scholar
Fowler, D. M. & Fields, S. Deep mutational scanning: a new style of protein science. Nat. Methods 11, 801â807 https://doi.org/10.1038/Nmeth.3027 (2014).
Article CAS PubMed PubMed Central Google Scholar
Wells, J. A. Systematic mutational analyses of protein protein interfaces. Method Enzymol. 202, 390â411 (1991).
Pennington, L. D., Aquila, B. M., Choi, Y., Valiulin, R. A. & Muegge, I. Positional analogue scanning: an effective strategy for multiparameter optimization in drug design. J. Med. Chem. 63, 8956â8976 https://doi.org/10.1021/acs.jmedchem.9b02092 (2020).
Muegge, I. & Hu, Y. In silico positional analogue scanning with Amber GPU-TI. J. Chem. Inf. Model. 62, 4448â4459 https://doi.org/10.1021/acs.jcim.2c00860 (2022).
Schönherr, H. & Cernak, T. Profound methyl effects in drug discovery and a call for new CâH methylation reactions. Angew. Chem. Int. Ed. 52, 12256â12267 (2013).
Cramer, J., Sager, C. P. & Ernst, B. Hydroxyl groups in synthetic and natural-product-derived therapeutics: a perspective on a common functional group. J. Med. Chem. 62, 8915â8930 (2019).
Chiodi, D. & Ishihara, Y. âMagic chloroâ: profound effects of the chlorine atom in drug discovery. J. Med. Chem. 66, 5305â5331 (2023).
Gillis, E. P., Eastman, K. J., Hill, M. D., Donnelly, D. J. & Meanwell, N. A. Applications of fluorine in medicinal chemistry. J. Med. Chem. 58, 8315â8359 (2015).
Hopkins, A. L., Groom, C. R. & Alex, A. Ligand efficiency: a useful metric for lead selection. Drug Discov. Today 9, 430â431 https://doi.org/10.1016/S1359-6446(04)03069-7 (2004).
Zhu, T. et al. Hit identification and optimization in virtual screening: practical recommendations based on a critical literature analysis. J. Med. Chem. 56, 6560â6572 https://doi.org/10.1021/jm301916b (2013).
Article CAS PubMed PubMed Central Google Scholar
Kuntz, I. D., Chen, K., Sharp, K. A. & Kollman, P. A. The maximal affinity of ligands. Proc. Natl Acad. Sci. USA 96, 9997â10002 https://doi.org/10.1073/pnas.96.18.9997 (1999).
Article ADS CAS PubMed PubMed Central Google Scholar
Halgren, T. A. Identifying and characterizing binding sites and assessing druggability. J. Chem. Inf. Model. 49, 377â389 https://doi.org/10.1021/ci800324m (2009).
Thomas, M., Bender, A. & de Graaf, C. Integrating structure-based approaches in generative molecular design. Curr. Opin. Struct. Biol. 79, 102559 https://doi.org/10.1016/j.sbi.2023.102559 (2023).
Flowers, J. et al. Expanding automated multiconformer ligand modeling to macrocycles and fragments. eLife 14, RP103797 https://doi.org/10.7554/eLife.103797 (2025).
Article CAS PubMed PubMed Central Google Scholar
Di, L., Kerns, E. H. & Carter, G. T. Drug-like property concepts in pharmaceutical design. Curr. Pharm. Des. 15, 2184â2194 https://doi.org/10.2174/138161209788682479 (2009).
Morgan, P. et al. Can the flow of medicines be improved? Fundamental pharmacokinetic and pharmacological principles toward improving phase II survival. Drug Discov. Today 17, 419â424 https://doi.org/10.1016/j.drudis.2011.12.020 (2012).
Houston, J. B. Utility of in vitro drug metabolism data in predicting in vivo metabolic clearance. Biochem. Pharmacol. 47, 1469â1479 https://doi.org/10.1016/0006-2952(94)90520-7 (1994).
Iwatsubo, T. et al. Prediction of in vivo drug metabolism in the human liver from in vitro metabolism data. Pharmacol. Ther. 73, 147â171 https://doi.org/10.1016/s0163-7258(96)00184-2 (1997).
Amidon, G. L., Lennernas, H., Shah, V. P. & Crison, J. R. A theoretical basis for a biopharmaceutic drug classification: the correlation of in vitro drug product dissolution and in vivo bioavailability. Pharm. Res. 12, 413â420 https://doi.org/10.1023/a:1016212804288 (1995).
Hopkins, A. L., Keseru, G. M., Leeson, P. D., Rees, D. C. & Reynolds, C. H. The role of ligand efficiency metrics in drug discovery. Nat. Rev. Drug Discov. 13, 105â121 https://doi.org/10.1038/nrd4163 (2014).
Young, R. J. & Leeson, P. D. Mapping the efficiency and physicochemical trajectories of successful optimizations. J. Med. Chem. 61, 6421â6467 https://doi.org/10.1021/acs.jmedchem.8b00180 (2018).
Rocklin, G. J., Mobley, D. L., Dill, K. A. & Hunenberger, P. H. Calculating the binding free energies of charged species based on explicit-solvent simulations employing lattice-sum methods: an accurate correction scheme for electrostatic finite-size effects. J. Chem. Phys. 139, 184103 https://doi.org/10.1063/1.4826261 (2013).
Article ADS CAS PubMed PubMed Central Google Scholar
Kim, H. et al. Artificial intelligence in drug discovery: a comprehensive review of data-driven and machine learning approaches. Biotechnol. Bioprocess Eng. 25, 895â930 https://doi.org/10.1007/s12257-020-0049-y (2020).
Kumar, A., Kini, S. G. & Rathi, E. A recent appraisal of artificial intelligence and in silico ADMET prediction in the early stages of drug discovery. Mini Rev. Med. Chem. 21, 2788â2800 https://doi.org/10.2174/1389557521666210401091147 (2021).
Fu, L. et al. ADMETlab 3.0: an updated comprehensive online ADMET prediction platform enhanced with broader coverage, improved performance, API functionality and decision support. Nucleic Acids Res. 52, W422âW431 https://doi.org/10.1093/nar/gkae236 (2024).
Article PubMed PubMed Central Google Scholar
Myung, Y., de Sa, A. G. C. & Ascher, D. B. Deep-PK: deep learning for small molecule pharmacokinetic and toxicity prediction. Nucleic Acids Res. 52, W469âW475 https://doi.org/10.1093/nar/gkae254 (2024).
Article PubMed PubMed Central Google Scholar
Swanson, K. et al. ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries. Bioinformatics https://doi.org/10.1093/bioinformatics/btae416 (2024).
Bergstrom, C. A., Wassvik, C. M., Johansson, K. & Hubatsch, I. Poorly soluble marketed drugs display solvation limited solubility. J. Med. Chem. 50, 5858â5862 https://doi.org/10.1021/jm0706416 (2007).
Obach, R. S. Prediction of human clearance of twenty-nine drugs from hepatic microsomal intrinsic clearance data: an examination of in vitro half-life approach and nonspecific binding to microsomes. Drug Metab. Dispos. 27, 1350â1359 (1999).
Friden, M. et al. Structure-brain exposure relationships in rat and human using a novel data set of unbound drug concentrations in brain interstitial and cerebrospinal fluids. J. Med. Chem. 52, 6233â6243 https://doi.org/10.1021/jm901036q (2009).
Fink, E. A. et al. Structure-based discovery of nonopioid analgesics acting through the α2A-adrenergic receptor. Science 377, eabn7065 https://doi.org/10.1126/science.abn7065 (2022).
Article CAS PubMed PubMed Central Google Scholar
Leung, C. S., Leung, S. S., Tirado-Rives, J. & Jorgensen, W. L. Methyl effects on protein-ligand binding. J. Med. Chem. 55, 4489â4500 https://doi.org/10.1021/jm3003697 (2012).
Article CAS PubMed PubMed Central Google Scholar
Sindt, F., Bret, G. & Rognan, D. On the difficulty to rescore hits from ultralarge docking screens. J. Chem. Inf. Model. 65, 5553â5566 https://doi.org/10.1021/acs.jcim.5c00730 (2025).
Lyu, J. et al. Ultra-large library docking for discovering new chemotypes. Nature 566, 224â229 https://doi.org/10.1038/s41586-019-0917-9 (2019).
Article ADS CAS PubMed PubMed Central Google Scholar
Moroz, Y., Chuprina, A. & Mykytenko, D. Enamine REAL DataBase â an instrumental and practical vehicle for charting new regions of the relevant drug discovery chemical space. In 251st American Chemical Society National Meeting & Exposition (ACS, 2016).
Verteramo, M. L. et al. Interplay of halogen bonding and solvation in protein-ligand binding. iScience 27, 109636 https://doi.org/10.1016/j.isci.2024.109636 (2024).
Article ADS CAS PubMed PubMed Central Google Scholar
Krimmer, S. G., Betz, M., Heine, A. & Klebe, G. Methyl, ethyl, propyl, butyl: futile but not for water, as the correlation of structure and thermodynamic signature shows in a congeneric series of thermolysin inhibitors. ChemMedChem. 9, 833â846 https://doi.org/10.1002/cmdc.201400013 (2014).
Wermuth, C. G. The Practice of Medicinal Chemistry (Elsevier Science & Technology, 2008).
Gahbauer, S. et al. Iterative computational design and crystallographic screening identifies potent inhibitors targeting the Nsp3 macrodomain of SARS-CoV-2. Proc. Natl Acad. Sci. USA 120, e2212931120 (2023).
Article CAS PubMed PubMed Central Google Scholar
Liu, F. et al. The impact of library size and scale of testing on virtual screening. Nat. Chem. Biol. 21, 1039â1045 https://doi.org/10.1038/s41589-024-01797-w (2025).
Clopper, C. J. & Pearson, E. S. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika 26, 404â413 (1934).
Virtanen, P. et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat. Methods 17, 261â272 (2020).
Article CAS PubMed PubMed Central Google Scholar
Lu, C. et al. OPLS4: improving force field accuracy on challenging regimes of chemical space. J. Chem. Theory Comput. 17, 4291â4300 https://doi.org/10.1021/acs.jctc.1c00302 (2021).
Balius, T. E. et al. Testing inhomogeneous solvation theory in structure-based ligand discovery. Proc. Natl Acad. Sci. USA 114, E6839âE6846 (2017).
Article CAS PubMed PubMed Central Google Scholar
Singh, I. et al. Structure-based discovery of conformationally selective inhibitors of the serotonin transporter. Cell 186, 2160â2175.e17 (2023).
Article CAS PubMed PubMed Central Google Scholar
Vigneron, S. F. et al. Docking 14 million virtual isoquinuclidines against the ÎŒ and Îș opioid receptors reveals dual antagonistsâinverse agonists with reduced withdrawal effects. ACS Central Sci. 11, 770â790 (2025).
Hua, T. et al. Activation and signaling mechanism revealed by cannabinoid receptor-Gi complex structures. Cell 180, 655â665.e18 (2020).
Article CAS PubMed PubMed Central Google Scholar
Hua, T. et al. Crystal structures of agonist-bound human cannabinoid receptor CB1. Nature 547, 468â471 (2017).
Article ADS CAS PubMed PubMed Central Google Scholar
Manglik, A. et al. Structure-based discovery of opioid analgesics with reduced side effects. Nature 537, 185â190 (2016).
Article ADS CAS PubMed PubMed Central Google Scholar
Yung-Chi, C. & Prusoff, W. H. Relationship between the inhibition constant (Ki) and the concentration of inhibitor which causes 50 per cent inhibition (I50) of an enzymatic reaction. Biochem. Pharmacol. 22, 3099â3108 (1973).
Schuller, M. et al. Fragment binding to the Nsp3 macrodomain of SARS-CoV-2 identified through crystallographic screening and computational docking. Sci. Adv. 7, eabf8711 https://doi.org/10.1126/sciadv.abf8711 (2021).
Article ADS PubMed PubMed Central Google Scholar
Maeda, S. et al. Development of an antibody fragment that stabilizes GPCR/G-protein complexes. Nat. Commun. 9, 3712 https://doi.org/10.1038/s41467-018-06002-w (2018).
Article ADS CAS PubMed PubMed Central Google Scholar
Goddard, T. D. et al. UCSF ChimeraX: meeting modern challenges in visualization and analysis. Protein Sci. 27, 14â25 https://doi.org/10.1002/pro.3235 (2018).
Liebschner, D. et al. Macromolecular structure determination using X-rays, neutrons and electrons: recent developments in Phenix. Acta Crystallogr. D Struct. Biol. 75, 861â877 https://doi.org/10.1107/S2059798319011471 (2019).
Article ADS CAS PubMed PubMed Central Google Scholar
Emsley, P., Lohkamp, B., Scott, W. G. & Cowtan, K. Features and development of Coot. Acta Crystallogr. D Biol. Crystallogr. 66, 486â501 https://doi.org/10.1107/S0907444910007493 (2010).
Article ADS CAS PubMed PubMed Central Google Scholar
Collins, P. M. et al. Gentle, fast and effective crystal soaking by acoustic dispensing. Acta Crystallogr. D Struct. Biol. 73, 246â255 https://doi.org/10.1107/S205979831700331X (2017).
Article ADS CAS PubMed PubMed Central Google Scholar
Kabsch, W. XDS. Acta Crystallogr. D Biol. Crystallogr. 66, 125â132 https://doi.org/10.1107/S0907444909047337 (2010).
Article ADS CAS PubMed PubMed Central Google Scholar
Evans, P. R. & Murshudov, G. N. How good are my data and what is the resolution? Acta Crystallogr. D Biol. Crystallogr. 69, 1204â1214 https://doi.org/10.1107/S0907444913000061 (2013).
Article ADS CAS PubMed PubMed Central Google Scholar
Pearce, N. M. et al. A multi-crystal method for extracting obscured crystallographic states from conventionally uninterpretable electron density. Nat. Commun. 8, 15123 https://doi.org/10.1038/ncomms15123 (2017).
Article ADS PubMed PubMed Central Google Scholar
Afonine, P. V. et al. Towards automated crystallographic structure refinement with phenix.refine. Acta Crystallogr. D Biol. Crystallogr. 68, 352â367 https://doi.org/10.1107/S0907444912001308 (2012).
Article ADS CAS PubMed PubMed Central Google Scholar
Correy, G. J. et al. Exploration of structure-activity relationships for the SARS-CoV-2 macrodomain from shape-based fragment linking and active learning. Sci. Adv. 11, eads7187 https://doi.org/10.1126/sciadv.ads7187 (2025).
Article CAS PubMed PubMed Central Google Scholar
We thank Y. Xiong for help with the FEP studies; and the UCSF Cryo-EM facility staff for training and technical assistance.
This work was supported by US NIH R35GM122481 (to B.K.S.), US DARPA ABC grant HR001123S0038 and US ARPA-H grant 1AY1AX000035 (principal investigator J.S.F.). O.M. was partially supported by US NIH postdoctoral fellowship F32GM154469. The UCSF cryo-EM equipment is partially supported by NIH grants S10OD020054, S10OD021741 and S10OD026881, and by the Howard Hughes Medical Institute.
These authors contributed equally: Xinyu Xu, Olivier Mailhot, Galen J. Correy, Xi-Ping Huang
Department of Pharmaceutical Chemistry, University of California San Francisco, San Francisco, CA, USA
Xinyu Xu, Olivier Mailhot, Karthik Srinivasan, Moira M. Rachman, Fangyu Liu, Katie L. Holland, Yujin Wu, Aashish Manglik & Brian K. Shoichet
Department of Bioengineering and Therapeutic Sciences, University of California San Francisco, San Francisco, CA, USA
Galen J. Correy, Kara Zielinski & James S. Fraser
Department of Pharmacology, NIMH Psychoactive Drug Screening Program, School of Medicine, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA
Xi-Ping Huang, Jing Wang & Bryan L. Roth
Department of Anatomy, University of California San Francisco, San Francisco, CA, USA
Yuliia Holota, Yuliia Kuziv & Yurii S. Moroz
Center for Drug Discovery, Department of Pharmaceutical Sciences, Northeastern University, Boston, MA, USA
Christos Iliopoulos-Tsoutsouvas & Alexandros Makriyannis
Department of Medicinal Chemistry, College of Pharmacy, University of Utah, Salt Lake City, UT, USA
Helen Diller Family Comprehensive Cancer Center, University of California, San Francisco, CA, USA
Yagmur U. Doruk, Morgan E. Diolaiti, Maisie G. V. Stevens & Alan Ashworth
Department of Chemistry and Pharmacy, Medicinal Chemistry, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg, Erlangen, Germany
Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
B.K.S., X.X. and O.M. designed the project. X.X. and O.M. designed the analogues with help from N.D.L. α2 Receptor studies were performed by X.X. and H.H., with guidance from P.G. α2A-related behavioural analyses were conducted by J.M.B., with guidance from A.I.B. α2 Structural studies were performed by K.S. under the supervision of A. Manglik. SERT uptake assays were carried out by X.-P.H. and J.W., with guidance from B.L.R. and ligand choice from Y.W. Mac1 biochemical assays were conducted by Y.U.D., M.G.V.S., M.E.D. and K.Z., and were supervised by A.A. Mac1 crystallography was performed by G.J.C. with guidance from J.S.F. CB2 receptor assays were performed by X.X. and C.I.-T., with guidance from A. Makriyannis and ligand advice from M.M.R. AmpC assays were conducted by X.X., F.L. and K.L.H. Y.S.M. coordinated and supervised small-molecule synthesis. All in vitro ADME and safety assays were supported by Y.H. and Y.K. through Bienta. FEP calculations and analyses were performed by D.S., guided by Y.Z. and R.A. O.M. performed statistical analyses for both ChEMBL and experimental data. X.X., O.M. and B.K.S. prepared the manuscript. B.K.S. supervised the project. All authors reviewed and approved the final manuscript.
Correspondence to Bryan L. Roth, James S. Fraser or Brian K. Shoichet.
B.K.S. is co-founder of Epiodyne, BlueDolphin and Deep Apple Therapeutics; serves on the Scientific Advisory Board (SAB) for Schrödinger, Vilya Therapeutics and Frontier Discovery; and is on the Scientific Resource Board (SRB) of Genentech. B.L.R. is founder of Onsero Therapeutics. J.S.F. is a consultant to and a shareholder of Vilya Therapeutics and Relay Therapeutics. D.S., Y.Z. and R.A. are employed by Schrödinger Inc. Y.H., Y.K. and Y.S.M. are employed by Enamine Ltd. Y.S.M. is employed by Chemspace LLC. The remaining authors declare no competing interests.
Nature thanks Derek Lowe, Celine Valant and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.
Publisherâs note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
a, Cumulative frequency of single non-hydrogen atom substitutions (CH3, F, Cl, Br, N and OH) in ChEMBL that improve activity, stratified by that of the parent compound. âPotentâ denotes parents with activity â€32 nM, âmidâ 32 nMâ1 ÎŒM, and âweakâ 1 ÎŒMâ1 mM. b, For each target in this study, the percentage of analogs that improve or decrease activity by â„3-fold or â„10-fold versus their parent. Open bars and points show observed percentages; black error bars show two-sided 95% confidence intervals estimated by percentile bootstrap (20,000 resamples). Sample sizes were n = 77, 33, 39, 82, 16 and 10 analogs for alpha2AAR, CB2, Mac1, SERT, ÎŒOR and AmpC, respectively. Individual binary outcomes are overlaid for AmpC (n = 10). Fold changes were rounded to the nearest integer before thresholding (for example, 9.6â9.9 counted as 10-fold). Compounds yielding no measurable curves were counted as â„10-fold decreases. c, Parent-level summary of analog effects in this study (rounded fold-change).
a, 2D structures of the parent compounds, with their measured affinity or potency and calculated cLogP values. b, Molecular Weight (MW) and Affinity or potency values for the parent compounds. c, Structural Similarity matrix showing the relationships between parent compounds based on molecular fingerprint comparisons.
a, Change in affinity or potency (ÎpActivity = pActivity(analog) â pActivity(parent), where pActivity is pKi, pIC50 or pEC50) plotted against ÎcLogP (analog â parent) for all analogâparent pairs (n = 220). Diagonal isoclines indicate ÎLipE values of 0, 0.5, 1.0 and 1.5, where LipE = pActivity â cLogP; a point above a given isocline has a larger gain in lipophilic efficiency than the indicated value. b, ÎpKi or pEC50 (analog â parent) against ÎcLogP for all analogâparent pairs with â„3Ă rounded potency/affinity improvement; â„10Ă improvements are highlighted in red. c, Scatter plot of ÎpKi or pEC50 (log M) versus ÎLE (analog â parent), where LE = 1.37 Ă pKi or pEC50/Nheavy (n = 220 analogâparent pairs). Pearson productâmoment and Spearman rank correlations were assessed using two-sided tests (R = 0.970, P = 3.62 Ă 10 â 136; Ï = 0.977, P = 1.91 Ă 10 â 148). d, Scatter plot of ÎpKi or pEC50 (log M) versus ÎLipE (analog â parent), where LipE = pKi or pEC50 â cLogP (n = 220 analogâparent pairs). Pearson productâmoment and Spearman rank correlations were assessed using two-sided tests (R = 0.895, P = 2.71 Ă 10 â 78; Ï = 0.854, P = 6.68 Ă 10 â 64). e, Binding free-energy change versus heavy-atom change for improved analogs. The y-axis shows ÎÎG (kcal/mol), computed from the affinity metric (Ki/IC50/EC50) as ÎG = RT ln(K) (298 K) and ÎÎG = ÎG_analog â ÎG_parent; negative ÎÎG indicates improved binding. The x-axis is ÎN (analog â parent; heavy-atom count difference). Points with rounded fold â„10 are highlighted in red; rounded 3â9 are shown in blue. f, Affinity or potency gains versus changes in lipophilic efficiency. Scatter plot of ÎpKi or pEC50 (analog â parent) against ÎLipE (analog â parent) for all analogâparent pairs with â„3Ă rounded potency/affinity improvement; â„10Ă improvements are highlighted in red.
a, Chemical structure of the â3453 parent and overlay of X-ray crystal structures of 12 analogs. b, Alignment of â0676 (21 ÎŒM) and â9249 (53 ÎŒM) with the â3453 parent, showing the different binding poses of the ortho-chloro and ortho-methyl analogs. c, Comparison of â6343 (2.1 ÎŒM), â1304 (1.9 ÎŒM) and â0676 (21 ÎŒM) shows that repositioning the aryl chloride disrupts packing with I131 and F132, reducing potency. d, Para-methyl (â9870, 7.6 ÎŒM) and para-chloro (â6343, 2.1 ÎŒM) substitutions have different effects despite similar sterics. The fluoro analog â6404 (5.5 ÎŒM) binds more weakly than the chloro analog â6343, consistent with weaker non-covalent interactions. e, Nitrogen substitutions (C â N) in â3169, â6627 and â8675 have different effects relative to their shared parent (8.9 ÎŒM): â8675 (meta-N, 4.5 ÎŒM) improves potency by approximately twofold, whereas â6627 (8.7 ÎŒM) and â3169 (10 ÎŒM) show little or no improvement. f, Overlay of the X-ray crystal structures of 12 ligand-bound complexes from the â3453 series, showing that single-atom modifications can cause substantial ligand movement or multiple binding poses. g, Examples of binding-pose shifts: â9249, â3176 and â3194 adopt markedly different poses relative to the â3453 parent, and two conformations are observed for â3194.
a, Observed percentages of single-atom substitutions with â„3-fold decreases in PK properties. Open bars and points show observed percentages; black error bars show two-sided 95% confidence intervals estimated by percentile bootstrap (10,000 resamples). Exact numbers of measured analogs (Cl, F, N, CH3 and OH, in that order) were: microsomal stability, n = 47, 35, 43, 61 and 32; permeability, n = 47, 35, 43, 61 and 32; hERG IC50, n = 36, 28, 26, 52 and 17; plasma stability, n = 47, 35, 43, 61 and 32; fraction unbound, n = 43, 31, 39, 56 and 28; and solubility, n = 46, 35, 43, 60 and 32. b, Counts of analogs with measured PK parameters and >3-fold or >10-fold losses.
Supplementary Figures 1â9, Supplementary Methods, and titles and legends for Supplementary Tables
Xu, X., Mailhot, O., Correy, G.J. et al. Development of a random background to understand ligand optimization. Nature (2026). https://doi.org/10.1038/s41586-026-11013-5
DOI: https://doi.org/10.1038/s41586-026-11013-5
Anyone you share the following link with will be able to read this content:
Sorry, a shareable link is not currently available for this article.
Provided by the Springer Nature SharedIt content-sharing initiative
Close Sign up for the Nature Briefing: Translational Research newsletter â top stories in biotechnology, drug discovery and pharma.
Comments (0)
No comments yet. Be the first to share your opinion!