Duplicate entries in a reference library quietly inflate coverage: the same warhead listed twice produces two matches for one chemical entity. So, standardization and deduplication run inside the extraction call. On our set of 981 PROTACs, 996 warheads, and 236 E3 ligands, that left 949, 981, and 228 distinct structures: 30 duplicate PROTAC rows collapsed plus 2 standardization failures, 6 duplicate warheads collapsed plus 9 standardization failures, and 4 duplicate ligands collapsed plus 4 standardization failures — each count itemized per file in the run summary. You can turn off the auto-clean stage if you want raw-input coverage numbers on purpose.
Once you are ready, proceed with the prompt:
"I have attached the files. Map the components and extract the linkers."
Matching 949 distinct PROTACs against 981 warheads and 228 ligands took about four minutes. The strict pass resolved 906 of them. The 43 it could not resolve were escalated to the MCS fallback, which recovered 2 more; the relaxed recovery pass then pulled 12 more out of what remained — each of those 12 comes back flagged for manual review, never presented as equivalent to a strict match. That left 29 PROTACs unresolved. Along the way, each accepted cut is grown out to the first heteroatom or nearest ring, so linkers end at a sensible point for a wet-lab synthesis instead of mid-motif.