Predicting Small-Molecule Permeability and Efflux Online: GNN-MTL, PyPermM, and Molecular Descriptors on Neurosnap

Written by Danial Gharaie Amirabadi | Published 2026-10-7

Human P-glycoprotein in a POPC bilayer with the eight case-study drugs

Figure 1. Human P-glycoprotein (PDB 6QEX) after a 1 ns all-atom simulation in a POPC bilayer, run with the OpenMM Molecular Dynamics service. The near half of the membrane is cut away. Paclitaxel (gold, inside the transporter) is placed from the cryo-EM pose and was not part of the simulation. The eight drugs are placed by hand for illustration, with carbons coloured by reference class: blue = high passive permeability, orange = low, green = known P-gp substrate.

Target binding is only useful if a drug reaches its target. Intestinal absorption, brain exposure and access to intracellular targets depend partly on membrane permeability and on efflux transporters such as P-glycoprotein (P-gp).

Caco-2 and MDCK cell assays measure permeability and directional transport; cell-free PAMPA measures passive permeability. On Neurosnap, three complementary approaches help prioritise compounds for these experiments:

We run all of them on eight reference drugs whose behaviour is well known, so you can see where the tools agree, where they disagree, and what to do about it.

Why permeability and efflux are different questions

Passive transcellular permeability depends on partitioning into a membrane and crossing its hydrophobic core. Lipophilicity can favour partitioning, while charge and exposed polarity impose a desolvation cost. Lipinski’s rule of five considers molecular weight, cLogP and hydrogen-bond donors and acceptors; Veber’s complementary heuristics use TPSA ≤140 Ų and ≤10 rotatable bonds (Lipinski et al., 1997; Veber et al., 2002).

Efflux transporters move substrates back out of cells. P-gp (ABCB1) is a major example: a lipophilic compound can cross membranes readily yet have limited brain exposure because P-gp removes it.

The standard way to separate the two effects is the bidirectional Transwell assay (Figure 2). Cells grow as a monolayer on a porous filter. You dose one side and measure how much appears on the other, then repeat in the opposite direction:

Bidirectional Transwell assay

Figure 2. The bidirectional Transwell assay behind Caco-2 and MDCK permeability and efflux numbers. The efflux assay is shown at pH 7.4 on both sides, without transporter inhibitors. The separate intrinsic-permeability assay in the GNN paper uses a 6.5/7.4 pH gradient with efflux-transporter inhibitors.

Caco-2 monolayers model intestinal permeability and express several transporters. MDCK-MDR1 cells overexpress human P-gp; the GNN paper also uses a more sensitive NIH-MDCK-MDR1 line. PAMPA contains an artificial lipid membrane and no transporters.

The ML model: multitask GNN on AstraZeneca data

Neurosnap runs the public four-task GNN-MTL checkpoint from Ivers Ohlsson et al. (ACS Omega, 2025), trained with Chemprop on harmonised AstraZeneca data from a single laboratory.

Two features distinguish the approach:

  1. Multitask learning. A shared molecular representation predicts four related endpoints, allowing information from one assay to improve predictions for another.
  2. Physicochemical features. Adding pKa and LogD improved the authors’ models. The public Chemprop v2 checkpoint used on Neurosnap requires no extra descriptors and accepts molecular structures directly.

The Caco-2 permeability training endpoint was measured with efflux-transporter inhibitors and a pH gradient (6.5/7.4). The ER endpoints were measured separately, without inhibitors, at pH 7.4 on both sides. Thus, the predicted Papp is an estimate of intrinsic permeability under that protocol; it is not the A→B denominator of the separately predicted ER.

The four outputs are Caco-2 ER, Caco-2 Papp, MDCK ER and NIH-MDCK ER. External macrocycle tests in the paper gave modest rank correlations; those evaluated model variants are not all identical to the public checkpoint.

The physics-based view: PyPermM

PyPermM estimates the free-energy cost of moving a molecule through a lipid bilayer. It implements the PerMM method (Lomize et al., 2019), integrating an insertion-energy profile across a 30 Å DOPC hydrocarbon core and applying empirical calibrations for different membrane systems.

Submit SMILES, SDF or CCD molecules. SMILES and CCD inputs receive a reproducible RDKit 3D conformer; 3D SDF coordinates are retained. In the tested PyPermM 0.0.4 implementation, propranolol and ibuprofen gave identical results at pH 2, 7.4 and 12: ionisable atom types were not assigned. These predictions therefore represent neutral species rather than a working pH-dependent ionisation correction. Outputs include:

The profile is the most informative output, because it shows why a molecule gets through or doesn't:

PyPermM supports C, N, O, S, F, Cl, Br and I, accepts up to 10 single-component molecules per job (2 to 80 heavy atoms each, no salts), and only models passive diffusion. It does not know about transporters.

Molecular descriptors: the quick first filter

Molecular Descriptors returns molecular weight, cLogP, TPSA, hydrogen-bond counts, rotatable bonds, formal charge and QED for up to 500 molecules per job. For a larger feature set, use the Mordred Molecular Descriptor Calculator.

Case study: eight drugs, three kinds of tool

Our references span high permeability (caffeine, propranolol, metoprolol), low Caco-2 permeability (atenolol, cimetidine, chlorothiazide) and established P-gp substrates (loperamide, digoxin). These groups are not mutually exclusive. Atenolol’s low Caco-2 permeability should not be confused with its moderate human absorption classification in ICH M9 (references 13–18).

The eight case-study molecules

Figure 3. The case-study molecules, drawn from an MMFF-optimised conformer (RDKit) and rendered at a common scale in Blender with Molecular Nodes. Labels show descriptors from the Molecular Descriptors service.

We submitted the eight reference drugs as a single batch to each prediction service. Open the shared case-study jobs to inspect the inputs, settings and results directly on Neurosnap. Digoxin retains its specified stereochemistry; the beta-blocker inputs do not specify an individual enantiomer.

Keep the stereochemistry in your SMILES: digoxin’s predicted Caco-2 ER is 37 with the full isomeric SMILES and 47 with stereocentres removed.

Step 1: descriptors

Drug MW TPSA (Ų) cLogP HBD QED
Caffeine 194 62 -1.0 0 0.54
Propranolol 259 41 2.6 2 0.84
Metoprolol 267 51 1.6 2 0.71
Atenolol 266 85 0.5 3 0.64
Cimetidine 252 89 0.6 3 0.24
Chlorothiazide 296 119 0.1 2 0.76
Loperamide 477 44 5.1 1 0.52
Digoxin 781 203 2.2 6 0.16

The high-permeability references have TPSA below 65 Ų and at most two donors; the low-permeability group has TPSA of 85–119 Ų. Digoxin lies beyond conventional rule-of-five space. Loperamide’s low TPSA and high cLogP suggest favourable membrane partitioning, making its strong efflux especially instructive.

Two cases show the limit of rules of thumb. Caffeine has the lowest cLogP in the set and is still highly permeable, because it is small and has no donors. Chlorothiazide's TPSA of 119 Ų is under Veber's 140 Ų limit, and it is still a poorly permeable drug.

Step 2: permeability and efflux from the GNN-MTL model

Human P-gp has a large drug-binding cavity lined by hydrophobic, aromatic and polar residues. Figure 4 shows paclitaxel in that cavity (Alam et al., 2019; PDB 6QEX), illustrating the structural context of its broad substrate recognition.

Paclitaxel in the P-glycoprotein drug-binding cavity

Figure 4. The drug-binding cavity of human P-gp with paclitaxel bound (PDB 6QEX). Rendered with Molecular Nodes in Blender. The structure is oriented in the membrane using the OPM database (Lomize et al., 2012).

Here are the GNN-MTL predictions after converting from log₁₀. Caco-2 Papp is in 10⁻⁶ cm/s, and the ERs are ratios:

Drug Caco-2 Papp Caco-2 ER MDCK ER NIH-MDCK ER
Caffeine 102 0.32 0.53 0.89
Propranolol 47 0.22 0.90 1.9
Metoprolol 31 0.33 1.3 3.0
Atenolol 0.94 0.56 0.43 0.26
Cimetidine 1.0 11 8.1 4.0
Chlorothiazide 0.12 4.9 0.08 0.05
Loperamide 15 3.0 6.1 41
Digoxin 3.4 37 31 22

Predicted Caco-2 permeability vs efflux ratio

Figure 5. Predicted Caco-2 Papp against Caco-2 efflux ratio for the eight reference drugs. Both axes are logarithmic. The shaded region is above the common efflux flag of ER = 2.

Reading Figure 5:

Step 3: passive diffusion with PyPermM

Figure 6 places illustrative copies of two drugs at different depths in the OpenMM bilayer beside their PyPermM profiles. The placements are not simulated trajectories, and the profiles come from PyPermM’s separate DOPC model.

Drugs at different depths in the bilayer next to their PyPermM insertion-energy profiles

Figure 6. Propranolol (left, high permeability) and chlorothiazide (right, low permeability) at seven depths across the bilayer from the OpenMM run, next to each molecule's PyPermM insertion-energy profile. Dots mark the depths at which the copies are drawn. PyPermM uses its own DOPC membrane model, so the two depth scales are aligned at the bilayer centre only.

Propranolol's profile has two wells of about -4.5 kcal/mol near the lipid-water interface and only a shallow rise (to about -2.5) at the bilayer centre. Its central energy remains below the water reference, although moving from the interfacial minimum to the centre still requires climbing about 2 kcal/mol. Chlorothiazide goes the other way. It barely binds at the interface (about -1 kcal/mol) and faces a barrier of roughly +12 kcal/mol at the centre. That difference is what drives its Caco-2 log P of -5.9 against propranolol's -3.1.

PyPermM insertion-energy profiles for all eight drugs

Figure 7. PyPermM insertion-energy profiles for all eight drugs (all modelled as neutral species), ordered by predicted Caco-2 permeability. The shaded band is the model's hydrocarbon core (±15 Å).
Drug Caco-2 log P BBB log P Binding energy (kcal/mol)
Loperamide -2.71 -1.84 -7.2
Propranolol -3.07 -2.34 -4.5
Metoprolol -3.61 -3.08 -2.5
Atenolol -3.96 -3.56 -2.0
Cimetidine -4.24 -3.94 -1.7
Digoxin -4.91 -4.87 -3.6
Caffeine -5.13 -5.17 -0.3
Chlorothiazide -5.86 -6.17 -1.1

We omit PyPermM’s PAMPA column from quantitative interpretation: loperamide and propranolol imply roughly 35 and 1.7 cm/s, implausibly large apparent assay permeabilities. All membrane calibrations derive from the same integral and are not independent evidence.

The coarse split agrees with the ML model and the rules of thumb: the lipophilic compounds (loperamide, propranolol, metoprolol) on top and the polar ones below. The details differ. For example, PyPermM puts digoxin below atenolol and cimetidine, while the GNN-MTL puts it above them. Two results stand out.

Caffeine is the outlier. Experimentally highly permeable, it ranks first in GNN-MTL but seventh in PyPermM. Its predicted central barrier is about +8 kcal/mol. This reveals a limitation of this run; the comparison alone cannot distinguish atom-typing, conformational or calibration errors.

Digoxin ranks sixth of eight. It has a favourable binding energy (-3.6), probably because of its large surface area, but a high central barrier (about +7 kcal/mol). PyPermM models passive diffusion only, so none of digoxin's efflux shows up here. That information has to come from the GNN-MTL model.

Step 4: cross-check with ADMET-AI and Admetica

Two more services give additional estimates for the same eight SMILES. ADMET-AI (Chemprop-RDKit trained on Therapeutics Data Commons datasets) returns a Caco-2 log Papp and a PAMPA probability. Admetica returns a Caco-2 log Papp and P-gp substrate and inhibitor probabilities. Figure 8 puts every tool's output side by side, with each column shaded by rank so the agreement is visible at a glance.

Cross-tool comparison

Figure 8. Outputs from descriptors, GNN-MTL, PyPermM, ADMET-AI and Admetica (GNN-MTL values converted from log₁₀) for the eight drugs. Cell shade is the rank within each column, darker meaning more favourable for net uptake: higher permeability, or lower TPSA and lower efflux. The right-edge colours show the reference class, as in Figure 5.

The picture is useful, and not entirely tidy:

A practical workflow

  1. Filter by descriptors. Run the whole library through Molecular Descriptors. Flag compounds with very high TPSA, many donors or MW well above 500 for closer inspection; use thresholds appropriate to the chemical series and delivery route.
  2. Rank with GNN-MTL. Run the survivors through Permeability and Efflux Prediction (up to 250 molecules per job). Prioritise compounds with high predicted Papp (a common convention treats above about 10 × 10⁻⁶ cm/s as high and below 1 as low) and low predicted efflux in the endpoints relevant to your project. Treat ER ≈ 2 as a screening flag, not a universal pass/fail boundary.
  3. Check the mechanism with PyPermM. For the shortlist, look at the insertion-energy profile, not just the number. A deep well with no barrier is a different story from a modest barrier, and comparing matched analogues can help test whether a structural change lowers the predicted barrier. The profile alone does not assign the barrier to a particular functional group.
  4. Cross-check the disagreements. Where the tools disagree by a large margin, run ADMET-AI or Admetica and look at the molecule: ionisation state, size, flexibility.
  5. Measure the few that matter. Keep your assay budget for compounds where the tools disagree or where the project depends on the answer.

Conclusion

Descriptors provide a fast first filter, GNN-MTL estimates assay permeability and efflux, and PyPermM adds a physical view of membrane insertion. Together, the eight-drug results highlight both useful agreement and informative disagreements. All tools are available through Neurosnap’s web interface or API.

All of the tools in this post are available on Neurosnap with no installation, through the web interface or the API.

Case-study jobs

Explore the inputs, settings and outputs in these public Neurosnap jobs:

Sources

  1. P. Ivers Ohlsson, G. M. Ghiandoni, S. Winiwarter, R. Mercado, V. Subramanian, "Prediction of Permeability and Efflux Using Multitask Learning," ACS Omega 10(45), 54148-54159 (2025). doi:10.1021/acsomega.5c04861. Model checkpoint: doi:10.5281/zenodo.16948542.
  2. A. L. Lomize, J. M. Hage, K. Schnitzer, et al., "PerMM: A Web Tool and Database for Analysis of Passive Membrane Permeability and Translocation Pathways of Bioactive Molecules," J. Chem. Inf. Model. 59(7), 3094-3099 (2019). doi:10.1021/acs.jcim.9b00225.
  3. C. Wagen, "PyPermM: permeability modeling made easy." github.com/rowansci/pypermm.
  4. K. Swanson, P. Walther, J. Leitz, et al., "ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries," Bioinformatics 40(7), btae416 (2024). doi:10.1093/bioinformatics/btae416.
  5. Datagrok, "Admetica: Datagrok repository for ADMET property evaluation." github.com/datagrok-ai/admetica. K. Yang, K. Swanson, W. Jin, et al., "Analyzing Learned Molecular Representations for Property Prediction," J. Chem. Inf. Model. 59(8), 3370-3388 (2019). doi:10.1021/acs.jcim.9b00237.
  6. K. Amani and D. Gharaie Amirabadi, "Neurosnap SDK Package" (2024). github.com/NeurosnapInc/neurosnap.
  7. P. Eastman et al., "OpenMM 7: Rapid development of high performance algorithms for molecular dynamics," PLOS Comput. Biol. (2017). doi:10.1371/journal.pcbi.1005659.
  8. A. Alam, J. Kowal, E. Broude, I. Roninson, K. P. Locher, "Structural insight into substrate and inhibitor discrimination by human P-glycoprotein," Science 363, 753-756 (2019). doi:10.1126/science.aav7102. PDB 6QEX.
  9. M. A. Lomize, I. D. Pogozheva, H. Joo, H. I. Mosberg, A. L. Lomize, "OPM database and PPM web server: resources for positioning of proteins in membranes," Nucleic Acids Res. 40, D370-D376 (2012). doi:10.1093/nar/gkr703
  10. C. A. Lipinski, F. Lombardo, B. W. Dominy, P. J. Feeney, "Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings," Adv. Drug Deliv. Rev. 23, 3-25 (1997). Reprinted in 64 (suppl.), 4-17 (2012), doi:10.1016/j.addr.2012.09.019.
  11. D. F. Veber, S. R. Johnson, H.-Y. Cheng, et al., "Molecular properties that influence the oral bioavailability of drug candidates," J. Med. Chem. 45, 2615-2623 (2002). doi:10.1021/jm020017n
  12. I. J. Hidalgo, T. J. Raub, R. T. Borchardt, "Characterization of the human colon carcinoma cell line (Caco-2) as a model system for intestinal epithelial permeability," Gastroenterology 96, 736-749 (1989). doi:10.1016/0016-5085(89)90897-4
  13. U.S. FDA, M9 Biopharmaceutics Classification System-Based Biowaivers, Guidance for Industry (May 2021), Annex I, Table 2. Full guidance.
  14. C. Wandel, R. Kim, M. Wood, A. Wood, "Interaction of morphine, fentanyl, sufentanil, alfentanil, and loperamide with the efflux drug transporter P-glycoprotein," Anesthesiology 96, 913-920 (2002). PubMed 11964599.
  15. T. Taur and R. Rodriguez-Proteau, "Effects of dietary flavonoids on the transport of cimetidine via P-glycoprotein and cationic transporters in Caco-2 and LLC-PK1 cell models," Xenobiotica 38, 1536-1550 (2008). doi:10.1080/00498250802499467.
  16. E. Beéry, Z. Rajnai, T. Abonyi, I. Makai, et al., "ABCG2 modulates chlorothiazide permeability: in vitro characterization of its interactions," Drug Metab. Pharmacokinet. 27, 349-353 (2012). PubMed 22790065.
  17. I. Bachmakov, U. Werner, B. Endress, D. Auge, et al., "Characterization of beta-adrenoceptor antagonists as substrates and inhibitors of the drug transporter P-glycoprotein," Fundam. Clin. Pharmacol. 20, 273-282 (2006). PubMed 16671962.
  18. L. Smetanova, V. Stetinova, D. Kholova, J. Kvetina, J. Smetana, Z. Svoboda, "Caco-2 cells and Biopharmaceutics Classification System (BCS) for prediction of transepithelial transport of xenobiotics (model drug: caffeine)," Neuro Endocrinol. Lett. 30(Suppl 1), 101–105 (2009). Journal article.

Explore more posts

RoseTTAFold3 Online: All-Atom Protein-DNA Structure Prediction

By Danial Gharaie Amirabadi

Revolutionizing Medicine: The Remarkable Stories of Imatinib and Oseltamivir

By Amélie Lagacé-O'Connor

Dimensionality Reduction Algorithms in Biology: UMAP, t-SNE, PCA, and Beyond

By Danial Gharaie Amirabadi

Achieving AlphaFold3-Level Accuracy with Open-Source Boltz-1

By Danial Gharaie Amirabadi

From Density to Atoms: Deep Learning Tools Advancing Cryo-EM

By Danial Gharaie Amirabadi

How to Use RFpeptides Online for Macrocyclic Peptide Design

By Danial Gharaie Amirabadi

Making Scientific Research
Faster & Easier

Register for free — upgrade anytime.

Interested in getting a license? Contact Sales.

Try Free