Blog
September 17, 2026
Beacon-2: Completing the reward function for chemical optimization in drug discovery
Today we’re introducing Beacon-2, a major update to our Beacon family of models, which have been state-of-the-art in ADMET prediction since their release in May 2025. Beacon-2 extends beyond ADMET, providing a unified system of models that predicts a molecule’s efficacious human dose directly from its chemical structure. Dose is the primary reward function of pre-clinical discovery, and we believe that making it computable via Beacon-2 is a step change for AI in drug discovery. Now, instead of optimizing individual properties in isolation, agentic systems can drive the drug discovery process itself.
In this blog post, we explain why dose makes sense as a reward function for chemical optimization in drug discovery; describe how Beacon-2 works; show how it performs on open benchmarks and across 20 anonymized partner programs; and share a public case study in which Indy, our medicinal chemistry agent, used Beacon-2 to autonomously reduce the predicted dose of a SARS-CoV‑2 Mac1 inhibitor by a factor of 17, producing more drug-like molecules than the same agent optimizing potency alone.
Dose as a reward function
To improve iteratively, AI systems need a reward function they can evaluate quickly and often: win rate in chess or Go, test suite pass rate for coding, proof compilation in formal math. Defining an objective a computer can measure is often the hard part. Once it exists, progress becomes tractable.
For preclinical compound optimization in drug discovery, human dose is the most natural reward function. Simply put, how many milligrams of the compound would a person have to take daily for it to have its intended therapeutic effect? Dose is highlighted again and again1-4 as the value that best captures a compound’s therapeutic potential, because it integrates both potency and its pharmacokinetics. A low dose drug has lower off-target toxicity risk,5,6 better patient compliance,7 and lower cost of goods.
However, the frontier of AI for drug discovery has focused disproportionately on one input to dose: potency. After the initial development of protein-ligand co-folding with AlphaFold 3,8 research groups have sprinted to develop accurate potency prediction methods.9-12 But optimizing towards potency is not enough. It’s well known in the field of reinforcement learning that optimizing over an incomplete reward function can lead to reward hacking and agent misalignment,13-16 and in fact the trap of optimizing too aggressively towards potency has been known in drug discovery for years.17 Mitigations such as multi-parameter optimization (MPO) scores that add other molecular properties alongside potency, or hard property cutoffs, risk optimizing towards the wrong goal or eliminating promising candidates.2,18,19 Only dose provides the properly aligned reward function for progressing drug programs to the clinic.
Fluconazole and itraconazole, two marketed antifungals with billions of dollars in sales, are a case in point. As we discussed in our previous post on dose projection, both compounds would be discarded under the wrong reward function: fluconazole for its mediocre potency, and itraconazole for being highly susceptible to liver metabolism. Only when integrating these compounds’ properties together into dose does it become clear that both support 100-200 mg daily dosing, rationalizing the clinical success of both drugs.
How Beacon-2 makes dose computable
Where Beacon-1 provided a suite of best-in-class models for individual molecular properties, Beacon-2 is a single system that predicts human dose end-to-end. It has three configurable modules:
- The ADMET-PK module predicts individual chemical properties, built on the industry-spanning data consortium and state-of-the-art ML architecture that have led Inductive to win all three OpenADMET blind challenges. This includes in vitro properties like permeability and hepatocyte stability, and in vivo pre-clinical pharmacokinetic properties for programs where this data is routinely collected.
- The potency module provides structurally informed estimates of compound potency via either of two approaches: automated free energy perturbation (FEP) using in-house infrastructure and auto-posing workflows; or Inductive’s fine-tunable co-folding affinity model built for high-throughput prediction, currently in alpha.
- The PK/PD module combines the outputs of the ADMET-PK module and potency module with physiologically-informed mechanistic models of pharmacokinetics. These methods are configurable per program based on the program’s measured data, clearance paths and exposure-response hypothesis, and are developed by Inductive’s team of experienced drug hunters.
Beacon-2 builds on a well-validated scientific foundation of projecting clinical human dose using pre-clinical in vitro data. For example, a recent paper from Genentech20 showed that first-in-human dose could be predicted accurately using a suite of measured in vitro ADME assays in conjunction with PK/PD modeling (see Page (2016)21 for a similar analysis from AstraZeneca).

But while those approaches require a compound to be synthesized and extensively tested before dose can be predicted, Beacon-2 computes dose entirely in silico. This is where Inductive’s state-of-the-art ADMET modeling capabilities and industry-spanning data consortium provide a unique unlock. Off-the-shelf ADMET models are not accurate enough to drive dose prediction and optimization, but Inductive’s program-finetuned models provide highly accurate predictions of the inputs to dose estimation (see, for example, this OpenADMET post22 comparing off-the-shelf models to our winning submission). As we show below, dose predictions from Inductive’s models line up well with those from measured in vitro data, making the reward function of pre-clinical drug discovery truly computable in silico, not just after a compound is synthesized and assayed. This compresses the compound optimization loop from months to minutes.
Applying Beacon-2: the OpenADMET blind challenge and live drug discovery programs
We’ve validated Beacon-2 both on open benchmarking datasets and on live partner programs.
To validate with public data we used the ExpansionRx OpenADMET blind challenge dataset, which includes over 7,000 compounds assayed by ExpansionRx for programs targeting RNA-mediated diseases. Participants competed to predict nine ADMET properties on a blinded set of compounds. Inductive’s Beacon-1 models were the most predictive out of 370+ competitors, ahead of big pharma and frontier AI teams like Merck, NVIDIA, and EMD Serono.
Using Beacon-2, we calculated predicted dose for the 325 competition test set compounds with complete measured ADME data. Because potency data was not included in the competition, we configured Beacon-2 with an assumed 10 nM reference potency value. We then compared this fully blind in silico prediction with predictions made by applying our PK/PD module to the measured in vitro data from the unblinded test set.
We observed excellent concordance between predicted doses from Beacon-2 and those using ExpansionRx’s measured ADME data. Spearman rank correlation is 0.79, with average fold error of 2.4x. That is high enough to drive real optimization and prioritization decisions, and is in line with the best rank correlations reported for state-of-the-art potency prediction methods like free-energy perturbation.23

To further evaluate Beacon-2, we’ve tracked its prospective performance in 20 programs running on Inductive’s Compass platform. These are programs where Inductive’s models aren’t just being benchmarked, but are being used week to week to decide which compounds are made and tested. They vary widely in their chemical space, target indication, stage in the lead optimization process, and scale of program data. They also vary in their ADMET-PK and PK/PD module configurations, with some for example incorporating extra-hepatic clearance modeled via whole-blood stability prediction. These dose projections also used an assumed 10 nM reference potency.
Across these diverse programs, Beacon-2 provides actionable dose prediction, with an average Spearman’s ρ of 0.61. In the types of comparisons that matter most, where two compounds in a program have a meaningful, 3x-or-more difference in dose using measured inputs, Beacon-2 correctly selects the lower-dose design over 81% of the time. This level of accuracy allows scientists using Beacon-2 dose projection to iterate rapidly towards development candidates.

Agentically optimizing towards low dose with Beacon-2
What excites us most about Beacon-2 is that it gives agents a reward function they can iterate against autonomously. To showcase this, we handed our medicinal chemistry agent, Indy, the recently disclosed compound AVI-645124 from a team at UCSF, targeting SARS-CoV‑2 Mac1.
Despite its unusual hydrazine functional group, AVI-6451 shows 28 nM IC50 potency and promising mouse PK. However, antivirals often require IC90 coverage, and against this challenging PD target, Beacon-2 predicts a poor efficacious human dose of several grams per day for QD (once-daily) dosing. This is far too high to take AVI-6451 into the clinic. Given the existing structure-activity relationship from the program, Indy flagged replacements of AVI-6451’s cyclopropyl and fluoro groups as the most promising directions for improving the molecule’s properties. We gave Indy access to Beacon-2 and asked it to drive down the predicted efficacious dose.
For this experiment, Indy had access to Beacon-2 equipped with its potency module, including both FEP-based and co-folding-based potency prediction. Inductive’s fine-tunable co-folding affinity model, currently in alpha, was trained on all available data from the Mac1 program. In temporal evaluation on existing Mac1 data, this fine-tuned model reached Spearman’s ρ of 0.72, compared to 0.10 for the open source, non-fine-tuned Boltz-2.

In parallel, FEP was configured using the AVI-6451 crystal structure, our auto-posing workflow and scalable FEP system. Beacon-2 used co-folding affinity for early high-throughput evaluation of compounds, then ran FEP on promising compounds, using the resulting dG values to produce the final potency prediction. We gave Indy a budget of 5 FEP rounds with Beacon-2, each containing 10 compounds, with the goal of optimizing dose.
As a comparator, we ran a second instance of Indy from the same starting compound with the same budget, optimizing on potency alone and with no consideration of predicted dose.
.png)
The contrast is stark. As shown in the figure above, when optimizing towards dose with Beacon-2, Indy makes steady progress both in improving potency and in reducing dose. This results in 12 compounds withpredicted daily doses similar to currently marketed antivirals like Paxlovid. The best compound, IB-47, was found in the final round with roughly a 17-fold reduction from AVI-6451.
Optimizing towards potency, in contrast, improves potency at a rapid clip while dose stagnates or even decays over the course of the 5 rounds. Only 3 compounds with predicted viable doses are identified, and the compounds with the best potency tend to have poor dose. The compound predicted to be the most potent, IB-93, has a predicted dose 17-fold worse than the original compound, driven by poor permeability and metabolic stability. This is reward hacking, encapsulated in one molecule.

To check that compounds prioritized with Beacon-2 are genuinely attractive candidates and not artifacts of Indy exploiting flaws in the dose model, we took the top 10 compounds from each run and evaluated them three further ways. All three favor dose-driven optimization:
- Druglikeness. The top Beacon-2-optimized compounds have a mean QED (quantitative estimate of druglikeness) of 0.64, against 0.52 for the top potency-optimized compounds. For reference, compounds AstraZeneca chemists rated "attractive" averaged 0.67 QED, while "unattractive" ones averaged 0.49.25
- Blinded chemist preference. We gave three experienced medicinal chemists 10 blinded head-to-head pairs, one compound from each run, and asked which was more likely to become a drug (following the approach of Choung et al. (2023)26). They picked the Beacon-2 compound in 87% of pairs.
- Alternative Beacon-2 PK configuration. We recomputed every dose with Beacon-2 configured using our ADMET-PK models of in vivo rat unbound clearance and unbound volume of distribution (via allometric scaling), in place of in vitro human microsomal clearance and the Oie-Tozer27,28 Vd model. This Beacon-2 configuration is designed to be used on programs where in vitro–in vivo correlation doesn't hold, or that have extensive in vivo data collection. The conclusion is unchanged: the Beacon-2-optimized compounds averaged 20x lower dose than the reference (up from 13x), while the potency-optimized compounds averaged 5x (up from 4x).
Projected doses ultimately need wet lab validation, and we are working towards letting agents equipped with Beacon-2 iterate towards low dose with the lab in the loop. But even with in silico tools alone, this case study shows how having the right reward function equips agents to find compounds that are worth taking to patients in the clinic.
Looking forward
The field of AI for drug discovery has an abundance of models, and increasingly, AI agents that can use them. What is needed is the right reward function for agents and models to build towards. Efficacious human dose is that reward function, and Beacon-2 makes it computable at high speed and scale. Beacon-2 is already running in dozens of programs, helping chemists design towards development candidates. If you're running a lead optimization program, or building drug discovery agents that need an objective worth optimizing towards, let's talk.
References
- Maurer, T. S.; Smith, D.; Beaumont, K.; Di, L. Dose Predictions for Drug Design. J. Med. Chem. 2020, 63 (12), 6423–6435.
- Pan, Y. Misconceptions among Medicinal Chemists Regarding ADME and PK Optimization. Med. Chem. Res. 2026, 35 (4), 668–685.
- Van Rompaey, D.; Ray Chaudhuri, S.; Ahmad, M.; Cisar, J.; Van Den Bergh, A.; Ash, J.; Wu, Z.; Bryan, M. C.; Edwards, J. P.; DesJarlais, R.; Wegner, J. K.; Ceulemans, H.; Mitra, K.; Polidori, D. Toward Dose Prediction at Point of Design. J. Med. Chem. 2024, 67 (24), 22282–22290.
- Lucas, A. J.; Sproston, J. L.; Barton, P.; Riley, R. J. Estimating Human ADME Properties, Pharmacokinetic Parameters and Likely Clinical Dose in Drug Discovery. Expert Opin. Drug Discov. 2019, 14 (12), 1313–1327.
- Lammert, C.; Einarsson, S.; Saha, C.; Niklasson, A.; Bjornsson, E.; Chalasani, N. Relationship between Daily Dose of Oral Medications and Idiosyncratic Drug-Induced Liver Injury: Search for Signals. Hepatology 2008, 47 (6), 2003–2009.
- Martin, M. T.; Koza-Taylor, P.; Di, L.; Watt, E. D.; Keefer, C.; Smaltz, D.; Cook, J.; Jackson, J. P. Early Drug-Induced Liver Injury Risk Screening: "Free," as Good as It Gets. Toxicol. Sci. 2022, 188 (2), 208–218.
- Nachega, J. B.; Parienti, J.-J.; Uthman, O. A.; Gross, R.; Dowdy, D. W.; Sax, P. E.; Gallant, J. E.; Mugavero, M. J.; Mills, E. J.; Giordano, T. P. Lower Pill Burden and Once-Daily Antiretroviral Treatment Regimens for HIV Infection: A Meta-Analysis of Randomized Controlled Trials. Clin. Infect. Dis. 2014, 58 (9), 1297–1307.
- Abramson, J.; Adler, J.; Dunger, J.; Evans, R.; Green, T.; Pritzel, A.; Ronneberger, O.; Willmore, L.; Ballard, A. J.; Bambrick, J.; et al. Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3. Nature 2024, 630 (8016), 493–500.
- Genesis Research Team. Pearl: A Foundation Model for Placing Every Atom in the Right Location. arXiv 2025, arXiv:2510.24670.
- Rossi, M.; Pederson, R.; Wang-Henderson, M.; Kaufman, B.; Williams, E. C.; Underkoffler, C.; Howell, O. L.; Layer, A.; Thaler, S.; Mardirossian, N.; Parkhill, J. A. TerraBind: Fast and Accurate Binding Affinity Prediction through Coarse Structural Representations. arXiv 2026, arXiv:2602.07735.
- Passaro, S.; Corso, G.; Wohlwend, J.; Reveiz, M.; Thaler, S.; Somnath, V. R.; et al. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. bioRxiv 2025, 2025.06.14.659707.
- Shenoy, N.; Errington, D.; Bengio, E.; Kapuśniak, K.; Klaeser, K.; Pang, Y. T.; Radenkovic, V.; Tossou, P.; Bois, T.; Wedlake, A.; Di Giovanni, F. Nesso-1: Accelerating Open-Source Binding Affinity Predictions. bioRxiv 2026, 2026.08.01.742196.
- Skalse, J.; Howe, N. H. R.; Krasheninnikov, D.; Krueger, D. Defining and Characterizing Reward Hacking. arXiv 2022, arXiv:2209.13085.
- Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; Mané, D. Concrete Problems in AI Safety. arXiv 2016, arXiv:1606.06565.
- Pan, A.; Bhatia, K.; Steinhardt, J. The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. arXiv 2022, arXiv:2201.03544.
- Zhuang, S.; Hadfield-Menell, D. Consequences of Misaligned AI. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020); Curran Associates, 2020; pp 15763–15773.
- Hann, M. M. Molecular Obesity, Potency and Other Addictions in Drug Discovery. Med. Chem. Commun. 2011, 2 (5), 349–355.
- Miller, R. R.; Madeira, M.; Wood, H. B.; Geissler, W. M.; Raab, C. E.; Martin, I. J. Integrating the Impact of Lipophilicity on Potency and Pharmacokinetic Parameters Enables the Use of Diverse Chemical Space during Small Molecule Drug Optimization. J. Med. Chem. 2020, 63 (21), 12156–12170.
- Webborn, P. J. H.; Beaumont, K.; Martin, I. J.; Smith, D. A. Free Drug Concepts: A Lingering Problem in Drug Discovery. J. Med. Chem. 2025, 68 (7), 6850–6856.
- Wang, Y.; Wang, W.; Broccatelli, F.; Kenny, J. R.; Wright, M. R.; Sorenson, J.; Desai, P. Human-Focused Multiparameter Optimization Scores for Rank Ordering Compounds during Early Drug Discovery: Validation of PBPK Models Based on Clinical PK Data. J. Med. Chem. 2025, 68 (16), 17960–17970.
- Page, K. M. Validation of Early Human Dose Prediction: A Key Metric for Compound Progression in Drug Discovery. Mol. Pharmaceutics 2016, 13 (2), 609–620.
- Castellanos, M.; MacDermott-Opeskin, H.; Ainsley, J.; Walters, P. Lessons Learned from the OpenADMET-ExpansionRx Blind Challenge: Can We Trust Zero-Shot ADMET Predictions? OpenADMET Blog, 2026. https://openadmet.ghost.io/zero-shot-expansiorx-admet-predictions/ (accessed 2026-09-15).
- Schindler, C. E. M.; Baumann, H.; Blum, A.; Böse, D.; Buchstaller, H.-P.; Burgdorf, L.; Cappel, D.; Chekler, E.; Czodrowski, P.; Dorsch, D.; et al. Large-Scale Assessment of Binding Free Energy Calculations in Active Drug Discovery Projects. J. Chem. Inf. Model. 2020, 60 (11), 5457–5474.
- Jaishankar, P.; Correy, G. J.; Matsui, Y.; Togo, T.; Rachman, M. M.; Stevens, M. G. V.; Hantz, E. R.; Zheng, J.; Diolaiti, M. E.; Montano, M.; et al. Discovery of AVI-6451, a Potent and Selective Inhibitor of the SARS-CoV-2 ADP-Ribosylhydrolase Mac1 with Oral Efficacy In Vivo. J. Med. Chem. 2026, 69 (1), 553–573.
- Bickerton, G. R.; Paolini, G. V.; Besnard, J.; Muresan, S.; Hopkins, A. L. Quantifying the Chemical Beauty of Drugs. Nat. Chem. 2012, 4 (2), 90–98.
- Choung, O.-H.; Vianello, R.; Segler, M.; Stiefl, N.; Jiménez-Luna, J. Extracting Medicinal Chemistry Intuition via Preference Machine Learning. Nat. Commun. 2023, 14, 6651.
- Øie, S.; Tozer, T. N. Effect of Altered Plasma Protein Binding on Apparent Volume of Distribution. J. Pharm. Sci. 1979, 68 (9), 1203–1205.
- Lombardo, F.; Obach, R. S.; Shalaeva, M. Y.; Gao, F. Prediction of Human Volume of Distribution Values for Neutral and Basic Drugs. 2. Extended Data Set and Leave-Class-Out Statistics. J. Med. Chem. 2004, 47 (5), 1242–1250.
- QC analytical reports (1H/13C/19F/2D NMR, chiral SFC)
- Propose, visualize, and triage forward and retrosyntheses
- Monitor the synthesis queue and flag stalled targets
- QC assay data as it lands and flag suspicious results
- Prioritize repeats and nominate compounds for downstream assays
- Analyze SAR and identify data gaps
- Evaluate IVIVC and suggest diagnostic experiments
- Design, enumerate, and triage analogs
- Launch and analyze physics-based workflows like docking and FEP
- Project human PK and efficacious dose
- Answer questions about a program’s history (what was learned and when)
- Track team progress against the target candidate profile
- Get new scientists up to speed on a program
- Make meeting-ready slides and figures
Prompt: I’m attaching two files with activity data, one with new data from today and one with historical data. Can you look through the new data and identify any QC issues that I should be aware of? Create a csv where every compound + assay run pair is labeled with any relevant issues.
Prompt: Do a nitrogen walk around the core of this molecule: O=C(NCC1=CC=C(C)C(F)=C1)C2=CC=C3C=C(N4[C@H](C)CCCC4)C=CN32
Prompt: Do an OH walk around this whole molecule: O=C(NCC1=CC=C(C)C(F)=C1)C2=CC=C3C=C(N4[C@H](C)CCCC4)C=CN32
Prompt: Replace the core of this molecule with this set of 5,6-ring systems (file attached): O=C(NCC1=CC=C(C)C(F)=C1)C2=CC=C3C=C(N4[C@H](C)CCCC4)C=CN32
[Related posts]
September 14, 2026
Blog
Meet Indy: an AI chemistry assistant designed by medicinal chemists to do medicinal chemistry
Indy, our medicinal chemistry assistant, supports drug discovery programs in improving the rigor of day-to-day analyses and decisions, by QC-ing assay data and monitoring the synthesis queue, to launching FEP, projecting human dose, and making meeting-ready SAR and synthesis slides.
August 1, 2024
Blog
Are local or global models better? Why not both?
A long-standing debate in cheminformatics is whether global property-prediction models perform better or worse than local QSAR models. We describe a result from our publication with Nested Therapeutics in which we show that the best of both worlds is to train a model on global data and then fine-tune it on local data.
June 7, 2024
Blog
Approaching AlphaFold 3 docking accuracy in 100 lines of code
AlphaFold 3 (AF3) is an exciting leap forward in our ability to predict the structure and properties of biomolecular systems. We explore how AF3 small-molecule docking compares to existing techniques and find that the story is more nuanced than headlines suggest. We conclude with thoughts on where AF3 may ultimately be most useful.
November 20, 2023
Blog
Get with the program: building ADME datasets that drive impact
Research advances often don't translate to practical impact. In ML for small molecule drug discovery, this is at least partly because benchmark datasets don't capture important components of drug programs. We show how this can happen in ADME prediction and provide a path forward for building more realistic benchmarks from existing public data.