2026.07.28
Automated SAR Cliff Analysis
An activity cliff occurs when a tiny structural change between otherwise similar molecules produces a disproportionately large difference in biological activity. Activity cliffs are a double-edged sword in drug discovery: a hidden cliff can ruin an analog series, yet it can also reveal the key pharmacophoric triggers that drive potency against targets such as the Pregnane X Receptor (PXR).

Traditionally, identifying activity cliffs requires manual or semi-automated analysis — often relying on custom scripts to compute fingerprints and calculate the Structure-Activity Landscape Index (SALI). With the new Insilico Medicine’s Cheminformatics Engine MolTools MCP, the entire analysis runs in natural language from just two prompts. In this case study, we demonstrate MolTools MCP in combination with the Grok 4.5 model inside the Cursor IDE. We apply it to the PXR dataset gathered by our data team to identify SAR cliffs and analyze its diversity.
A Transparent Engine:
Inspecting SAR Capabilities
Start with a general question: "What analytical methods and visualization tools do you provide for identifying activity cliffs and exploring lead series similarity?" MolTools then lays out its capabilities, each with a clear description and adjustable parameters.
The success of any SAR cliff analysis depends heavily on data quality. You can ask MolTools MCP about data preparation directly: "Should I preprocess and deduplicate my dataset before running a SALI activity cliff analysis, or can I upload a raw file directly?" — and it will recommend the right next step.
Otherwise, you can follow the automated data-preparation workflow we described in our earlier blog post.
Identify Top-30 Divergent Pairs
An activity cliff is usually identified by comparison with the most similar structures. In a simple case, you can narrow this down to analyzing structural pairs, where two molecules are very similar but possess different biological activity. MolTools MCP can retrieve the top-30 structural pairs with a single prompt: "Please find SAR cliffs from the PXR_data_deduplicated.csv dataset. Show me the table with top-30 most divergent structural pairs."
The MolTools MCP engine processes all unique compound pairs by calculating the similarity of the two structures in each pair. A pair is flagged when the similarity metric falls in the 0.9–1.0 range but the activity difference is large relative to the activity distribution across the dataset. The MCP returns an interactive widget and a structured table isolating the top-30 pairs — in each pair, one structure is likely the activity cliff.
Similarity Analysis
Once individual cliffs are identified, assessing overall dataset diversity gives crucial context for the SAR landscape. To gauge library diversity, we prompt: "Generate similarity distribution histograms in both interactive canvas and downloadable PNG formats." The MolTools MCP computes the pairwise Morgan fingerprint matrix and delivers the results in two formats:

  1. Interactive Workspace Canvas (pxr-similarity-histogram): a view with hoverable bin counts, quantile statistics, and contextual interpretation notes.
  2. Downloadable Static PNGs: high-resolution figures for publications or slides, including full-count plots with mean/median indicators (PXR_tanimoto_similarity_histogram.png) and custom MolTools-themed bar charts (PXR_tanimoto_similarity_histogram_moltools.png).
The engine aggregates all pairwise Tanimoto scores across the library, rendering an interactive histogram on your workspace canvas while exporting a high-resolution PNG to your project directory.
Quantitative Landscape Insights
The engine automatically aggregates full distribution quantiles to give medicinal chemists a complete profile of the library:

  • Mean Tanimoto: 0.139 | Median Tanimoto: 0.127 (5th–95th percentile: 0.08–0.21)
  • High-Similarity Tail: Only 0.29% of pairs have a Tanimoto similarity ≥ 0.70, and just 0.04% reach ≥ 0.85.
  • Singletons: 399 compounds have no structural neighbors with a Tanimoto similarity ≥ 0.70.

The agent also interprets the results: the PXR library is highly diverse overall, with a low median Tanimoto similarity of ~0.13. The thin high-similarity tail (≥ 0.70), representing just 0.29% of all pairs, contains near-duplicates and same-ligand assay variants.
The Takeaway: Automating SAR cliff identification removes the manual scripting bottleneck in SAR analysis. The new Insilico Medicine’s Cheminformatics Engine MolTools MCP provides a transparent, agent-driven pipeline that highlights critical activity cliffs, computes global library metrics, and exports publication-ready charts — all while guarding against false positives caused by assay variance or unstandardized data. By replacing custom RDKit scripts with natural-language prompts, researchers can map structure-activity landscapes rapidly and reproducibly before committing to wet-lab synthesis.

Ready to try it yourself? Reach out via petrina@insilicomedicine.com
Stay tuned, follow us on social media!