Ligands#
Ligand detection, preparation, parameter I/O, and fragmentation are available
from the public tmol.ligand package.
Ligand typing, preparation, fragmentation, and serialization.
- class tmol.ligand.FragmentConnection(fragment_id: int, partner_fragment_id: int, connection_name: str, partner_connection_name: str, atom_name: str, partner_atom_name: str, bond_type: str, in_ring: bool = False)[source]#
Bases:
objectOne directed side of a cut bond.
- class tmol.ligand.LigandFragmentDefinition(ligand_name: str, atom_to_fragment: Mapping[str, int], fragment_ids: tuple[int, ...], fragment_preparations: tuple[LigandPreparation, ...], connections: tuple[FragmentConnection, ...])[source]#
Bases:
objectStructure-independent definition of one fragmented ligand type.
- class tmol.ligand.LigandPreparation(residue_type: RawResidueType, partial_charges: dict[str, float], cartbonded_params: CartRes, atom_type_elements: dict[str, str] | None = None, adds_patches: tuple = (), variant_partial_charges: dict[str, dict[str, float]] | None = None, connection_params: tuple[ConnectionCartRes, ...] = (), additional_cartbonded_params: dict[str, CartRes] | None = None, baseline_sha256: str | None = None)[source]#
Bases:
objectThe unified abstraction both ligand-pipeline paths converge on.
A
LigandPreparationis everything tmol needs to inject one ligand into aParameterDatabase: the residue type definition, partial charges, cartbonded parameters, and (optionally) the element mapping for any new atom-type names introduced.Both pipeline entry points produce this same struct:
AtomArray / SMILES path —
tmol.ligand.prepare_single_ligand()types the (already protonated, already charged) SMILES-derived molecule, builds the residue, and extracts cartbonded params, returning oneLigandPreparationper ligand.Params-file path —
tmol.ligand._params_file.load_params_file()parses a.tmolYAML and returnslist[LigandPreparation]describing the residues defined in that file.
Either list is then handed to
inject_ligand_preparations(), the single chokepoint that extends theParameterDatabase. Tests can equally roundtripAtomArray → LigandPreparation → .tmol → LigandPreparationand expect bit-equivalent injection.
- exception tmol.ligand.LigandPreparationError[source]#
Bases:
RuntimeErrorA detected ligand could not be prepared, registered, or retained.
Raised by
prepare_ligands()(and theprepare_ligands=TrueIO paths) whenstrict_ligands=Trueand a non-standard residue is skipped or fails preparation, instead of silently dropping it. Passstrict_ligands=Falseto downgrade these failures to warnings.
Bases:
RuntimeErrorRaised when an OB-fallback helper is called but
openbabelis missing.
- tmol.ligand.apply_fragment_connections(pose_stack, mapping: FragmentedLigandPoseMapping)[source]#
Install fragment cut bonds and rebuild all inter-block bond separations.
- tmol.ligand.build_split_block_mapping(pose_stack, resolved_mapping: FragmentedLigandPoseMapping)[source]#
Build a SplitBlockMapping from a fragmented PoseStack.
Ensures the original (unfragmented) block types are present in the PackedBlockTypes of the returned PoseStack, then records for each fragment block: its pose/block indices, its group within that pose, the index of the original block type, and the per-atom mapping split_atom → orig_atom.
Fragment block-type names are expected to follow the convention
"ORIGNAME.FRAGMENT_ID"(e.g."LIG.0").Returns the (possibly PBT-extended) PoseStack and the SplitBlockMapping.
- tmol.ligand.chem_comp_types_from_cif(cif_path) dict[source]#
Read the _chem_comp.type of every component declared by a CIF file.
A file that does not number its residues along a polymer sequence can still declare their types, which is what says whether they are polymer-linking.
- tmol.ligand.detect_nonstandard_residues(atom_array: AtomArray, canonical_ordering: CanonicalOrdering, chem_comp_types: dict | None = None) list[NonStandardResidueInfo][source]#
Detect residues in an AtomArray that are not in tmol’s database.
Any residue whose 3-letter code is not in the canonical ordering is returned for preparation, regardless of whether it is a ligand, modified amino acid, or modified nucleotide.
- Parameters:
atom_array – Biotite AtomArray from a CIF or PDB file.
canonical_ordering – The current tmol CanonicalOrdering, which defines known residue types.
chem_comp_types –
{comp_id: type}declared by the input file, consulted where it does not number residues along a sequence.
- Returns:
A list of NonStandardResidueInfo objects, one per unique unknown residue name.
- tmol.ligand.expand_fragmented_ligands(structure: AtomArray | AtomArrayStack, definitions: Sequence[LigandFragmentDefinition]) tuple[AtomArray | AtomArrayStack, FragmentedLigandPoseMapping][source]#
Replace each annotated ligand residue with contiguous fragment residues.
- tmol.ligand.fragment_ids_from_atom_array(atom_array: AtomArray) ndarray | None[source]#
Return validated fragment IDs, or
Nonewhen no split is requested.
- tmol.ligand.get_chem_comp_type(res_name: str, chem_comp_types: dict | None = None) str | None[source]#
The chemical component type the input file declares for a residue.
- Parameters:
res_name – Three-letter residue code.
chem_comp_types – Types declared by the input file.
- Returns:
The type string (e.g. “NON-POLYMER”, “L-PEPTIDE LINKING”), or None where the file declares none.
- tmol.ligand.is_polymer_linking_component_type(component_type: str | None) bool[source]#
Whether a declared chemical-component type denotes a polymer residue.
Returns True for modified amino acids, nucleotides, and saccharides (e.g. “L-PEPTIDE LINKING”, “DNA LINKING”, “D-SACCHARIDE”), which are expected to be covalently attached to a chain. Returns False for “NON-POLYMER” small-molecule ligands and where nothing is declared.
- tmol.ligand.polymer_entity_residues(atom_array: AtomArray) frozenset[str] | None[source]#
Names of the residues the input file places in a polymer entity.
An mmCIF says which of its entities are polymers, and that is what separates a chain from a ligand, a glycan or solvent – including for a terminal cap, which is typed NON-POLYMER despite being part of the chain.
tmol.io.atom_array_from_cif()records it per atom.Failing that,
label_seq_idis read as a proxy: a polymer entity numbers its residues along its sequence. It is only a proxy – a generated file may number a ligand that belongs to no sequence – so it is used only where the entities themselves are not recorded. Returns None when neither is.
- tmol.ligand.inject_ligand_preparations(param_db: ParameterDatabase, preparations: list[LigandPreparation], *, strict_atom_types: bool = False) ParameterDatabase[source]#
Inject a batch of
LigandPreparationrecords into a database.The single chokepoint both pipeline paths use — given a list of prepared ligands (regardless of whether they came from a
.tmolfile or an AtomArray), this function aggregates their residue types, atom types, charges, and cartbonded params and evolves the inputParameterDatabaseviatmol.database.inject_residue_params().Base definitions whose name already exists in
param_dbare skipped. Their additional patches, variant charges and connection records are still installed; repeating an identical bundle returns the original database. A preparation withbaseline_sha256instead replaces an exact residue after checking its original parameters (or an already installed result).- Parameters:
param_db – Base database (not modified).
preparations – One
LigandPreparationper ligand to register.strict_atom_types – If True, raise when an atom type’s element cannot be resolved from any preparation’s
atom_type_elements— otherwise fall back to a name-based heuristic and emit a warning.
- Returns:
A new frozen
ParameterDatabaseextended with all provided preparations.
- tmol.ligand.inject_params_file(param_db: ParameterDatabase, path: str | Path, *, strict_atom_types: bool = False) ParameterDatabase[source]#
Load a single
.tmolfile and inject it into a ParameterDatabase.
- tmol.ligand.inject_params_files(param_db: ParameterDatabase, paths: list[str | Path], *, strict_atom_types: bool = False) ParameterDatabase[source]#
Load multiple
.tmolfiles and inject them in one shot.
- tmol.ligand.ligand_smiles_from_atom_array(atom_array: AtomArray, *, res_name: str | None = None, with_atom_map: bool = False, keep_hydrogens: bool = False) str[source]#
Derive a canonical SMILES for a ligand AtomArray from its bond table.
The SMILES is derived purely from the input atoms and their explicit bonds (never a residue-code / CCD-template lookup, never geometry-based bond perception). Geometry-based bond-order corrections are still applied for motifs the input encodes inconsistently (carboxylates).
- Parameters:
atom_array – The ligand sub-array (heavy + optional hydrogen atoms).
res_name – Residue code, used only for log/error messages.
with_atom_map – Tag heavy atoms with source-index map numbers for CIF atom naming downstream.
keep_hydrogens – Preserve the supplied hydrogen counts, including zero, when the input already has its intended protonation state.
- Returns:
A canonical SMILES string.
- Raises:
ValueError – If the AtomArray carries no bond table (bond orders must be supplied by the input; a bonds-absent ligand such as a plain PDB cannot be prepared without guessing chemistry), or if no SMILES could be derived from the bonds present.
- tmol.ligand.load_params_file(path: str | Path) list[LigandPreparation][source]#
Load a tmol params YAML file as a list of
LigandPreparation.The returned list is the same abstraction the AtomArray pipeline produces (see
tmol.ligand.prepare_single_ligand()), so the caller can pass it directly totmol.ligand._registry.inject_ligand_preparations()regardless of which input form (file or AtomArray) it came from.The
.tmolschema is the nestedchemical:/elec:/cartbonded:shape — files using the legacy flat schema (top-levelresidues:etc.) raise aValueErrorpointing at the migration.
- tmol.ligand.nonstandard_residue_info_from_mol2(mol2_path: str | Path, res_name: str | None = None) NonStandardResidueInfo[source]#
Construct
NonStandardResidueInfofrom a ligand Mol2 file.Low-level reader retained for the DUD-80 SMILES->params parity harness (it reads both the OpenBabel-generated and Rosetta ground-truth mol2 files). Preserves Tripos aromatic flags, atom-type subtypes, and per-atom partial charges, avoiding lossy rdkit<->biotite round-trips.
- tmol.ligand.nonstandard_residue_info_from_mol2_block(mol2_block: str, res_name: str | None = None) NonStandardResidueInfo[source]#
Construct
NonStandardResidueInfofrom an in-memory mol2 string.In-memory analogue of
nonstandard_residue_info_from_mol2()— parses a TRIPOS mol2 block directly, with no temp-file write/read. Preferred for high-throughput SMILES batches (seenonstandard_residue_info_from_smiles_via_mol2()).
- tmol.ligand.nonstandard_residue_info_from_smiles_via_mol2(smiles: str, res_name: str | None = None, *, ph: float = 7.4, protonate: bool = True, seed: int | None = None) NonStandardResidueInfo[source]#
Construct
NonStandardResidueInfofrom a SMILES via the mol2 route.Implements the canonical ligand-prep protocol end to end:
normalize bare radical oxygens (
[O]->[O-]) so source carboxylate/sulfonate notation has a well-defined charge,optionally pKa-protonate the SMILES with Dimorphite-DL (
protonate),generate a 3D mol2 with MMFF94 partial charges via OpenBabel (kept in memory as a string — no temp file), then
read that mol2 with
nonstandard_residue_info_from_mol2_block().
This never builds a biotite atom-array from an RDKit embedding and never recomputes MMFF on a reconstructed graph — the OpenBabel MMFF94 charges flow through untouched (
skip_protonation/ authoritative charges are set by the mol2 reader), so fused-ring aromatics keep correct charges.- Parameters:
smiles – Ligand SMILES string.
res_name – Residue name (default inferred /
"L_1").ph – Target pH for the Dimorphite protonation step.
protonate – When
True(default) run Dimorphite onsmilesfirst; setFalseto pin an already-protonated SMILES verbatim.seed – Fixed RNG seed for reproducible 3D coordinates;
Noneis random.
- Raises:
OpenBabelUnavailableError – If the
openbabelpackage is missing (this path requires it for the SMILES -> mol2 conversion).ValueError – If OpenBabel cannot build a charged mol2 for
smiles.
- tmol.ligand.prepare_ligand_from_cif(cif_path: str, *, param_db: ParameterDatabase | None = None, ph: float = 7.4, seed: int | None = None, strict_atom_types: bool = False, res_name: str | None = None) tuple[ParameterDatabase, CanonicalOrdering][source]#
Prepare a single ligand from a CIF file and inject it into a database.
Runs the same full pipeline as
prepare_ligand_from_smiles(); the only CIF-specific step is the front end. A SMILES is derived from the CIF ligand’s bond table (supplemented from the CCD, never geometry perception) and run through the SMILES -> mol2 -> params path (protonation, 3D conformer, MMFF94 charges). The prepared residue’s heavy-atom names are then mapped back to the CIF atom names via the atom-order map carried through the round-trip.- Parameters:
cif_path – One ligand residue in a text, compressed or binary CIF file; the first model is selected. Use
prepare_ligands()for arrays containing multiple residues.param_db – Base database (not modified); defaults to the tmol default.
ph – Target pH for protonation.
seed – Reproducible conformer seed, shared with the MOL2/SMILES entry paths.
strict_atom_types – Fail on unknown atom-type element mappings.
res_name – Optional residue name override.
- Returns:
A
(ParameterDatabase, CanonicalOrdering)with the ligand injected.
- tmol.ligand.prepare_ligand_from_mol2(mol2_path: str, *, param_db: ParameterDatabase | None = None, ph: float = 7.4, mode: str = 'auto', seed: int | None = None, strict_atom_types: bool = False, res_name: str | None = None) tuple[ParameterDatabase, CanonicalOrdering][source]#
Prepare a single ligand from a Tripos mol2 file and inject it.
By default, preserve complete molecules with supported partial charges; otherwise run shared protonation and charge generation. Use
regenerateto apply the requested pH even to a neutralized, fully hydrogenated input.- Parameters:
mol2_path – Path to the ligand MOL2 or SDF file.
param_db – Base database (not modified); defaults to the tmol default.
ph – Target pH for protonation.
mode –
"keep"requires prepared input and retains its protonation and charges;"auto"(default) regenerates only unprepared input;"regenerate"always runs pH-dependent preparation. Complete hydrogens and supported charges do not establish the intended pH.seed – Reproducible conformer seed when charges must be generated.
strict_atom_types – Fail on unknown atom-type element mappings.
res_name – Optional residue name override.
- Returns:
A
(ParameterDatabase, CanonicalOrdering)with the ligand injected.
- tmol.ligand.prepare_ligand_from_smiles(smiles: str, *, param_db: ParameterDatabase | None = None, ph: float = 7.4, strict_atom_types: bool = False, res_name: str | None = None, protonate: bool = True, seed: int | None = None) tuple[ParameterDatabase, CanonicalOrdering][source]#
Prepare a single ligand from a SMILES string and inject it into a database.
Follows the canonical ligand-prep protocol: Dimorphite-DL pKa-protonates the SMILES at
ph, OpenBabel generates a 3D mol2 with MMFF94 partial charges, and that mol2 is read verbatim (atom names, coordinates, charges, and bond orders preserved). The MMFF94 charges flow through untouched — there is no biotite atom-array round-trip or MMFF recompute.- Parameters:
protonate – When
True(default) Dimorphite protonatessmilesfirst; setFalseto pin an already-protonated SMILES verbatim.seed – Fixed RNG seed for reproducible 3D coordinates;
Noneis random.
- tmol.ligand.prepare_ligands(atom_array: AtomArray, param_db: ParameterDatabase | None = None, ph: float = 7.4, strict_atom_types: bool = False, params_files: list[str] | None = None, params_output: str | None = None, strict_ligands: bool = True, return_fragment_definitions: bool = False, return_cut_partners: bool = False, chem_comp_types: dict[str, str] | None = None, seed: int | None = None, coordinating_atoms: dict[str, frozenset[str]] | None = None) tuple[source]#
Detect, prepare, and register all non-standard residues.
Scans the input AtomArray for residues not in the ParameterDatabase, runs each through the unified SMILES→OpenBabel mol2→typing→residue-build pipeline, and returns a new ParameterDatabase with the ligand data injected. Ligand/glycan attachment chemistry uses the connected capped model for local typing, bonded terms and atom construction, preserving prepared residue charges. Lengths and angles use the same Cartesian constants as ordinary ligand parameters. Supplied explicit connection records take precedence and preserve both endpoint residue types.
- Parameters:
atom_array – A biotite AtomArray from a CIF or PDB file.
param_db – The base ParameterDatabase (not modified). If None, the default database is used.
ph – pH AtomWorks protonates residues without hydrogens at; each ligand keeps the protonation state its hydrogens give it.
strict_atom_types – If True, fail when unknown atom-type element mappings are encountered during registration.
params_files – Optional list of tmol YAML params file paths to inject before detection. Residues defined in these files skip residue generation. Missing attachment records are generated; existing records are reused without regeneration.
params_output – Optional path to write all prepared ligand data to a tmol YAML params file for later reuse.
strict_ligands – If True (default), raise
LigandPreparationErrorwhen a detected non-standard residue is skipped (metal-containing or covalently linked) or fails preparation, instead of silently dropping it. If False, such residues are logged as warnings and skipped, leaving them to be filtered out during pose construction.return_cut_partners – Internal option. If True, append the residue names whose covalent bonds were cut. A ligand skipped under
strict_ligands=Falseloses its bonds here, and a caller that goes on to build a pose must know that, or it will reject the dangling partner it was told to drop.return_fragment_definitions – Internal/context-building option. If True, include definitions derived from
tmol_fragment_idannotations as the third return value.chem_comp_types –
{comp_id: type}read from the input file’s_chem_comptable (seetmol.ligand.chem_comp_types_from_cif()), consulted where the file does not number the residue along a polymer sequence. A polymer-linking type routes the residue toprepare_polymer_residue()instead of the free-molecule ligand path.seed – Fixed RNG seed for the 3D conformer each residue is built from.
Noneis random, which makes the prepared residue types differ between runs.coordinating_atoms – {residue name: atoms bonded to a metal}, which a residue’s charges cannot account for. Each hydrogen state the copies of a ligand present gets a type (see
_state_variants).
- Returns:
A (ParameterDatabase, CanonicalOrdering) tuple. When
return_fragment_definitionsis true, a third element containing the structure-independent ligand fragment definitions is returned. The returned ParameterDatabase is a new instance with all detected ligands injected; the inputparam_dbis not modified.- Raises:
LigandPreparationError – If
strict_ligandsand any detected ligand cannot be prepared and registered.
- tmol.ligand.prepare_polymer_residue(atom_array, canonical_ordering, param_db: ParameterDatabase | None = None, *, ph: float = 7.4, profile=None, connection_atoms=None, seed: int | None = None) LigandPreparation[source]#
Prepare one polymer residue: cap it, run the ligand pipeline, restore the polymer connections.
atom_array is the residue as it appears in a structure, with its polymer connections open; profile defaults to whichever backbone its atoms match. connection_atoms names the atoms that bond to neighbouring residues, which is what distinguishes a residue linked through its carbonyl from one linked through a sidechain carbon.
- tmol.ligand.prepare_ligands_from_smiles(smiles, *, param_db: ParameterDatabase | None = None, ph: float = 7.4, strict_atom_types: bool = False, protonate: bool = True, seed: int | None = None) tuple[ParameterDatabase, dict][source]#
Prepare one residue type per SMILES, naming them L_1, L_2, …
Returns the extended database and a {smiles: residue name} mapping.
- tmol.ligand.prepare_single_ligand(ligand_info: NonStandardResidueInfo, name_source: NonStandardResidueInfo | None = None, assign_ring_chis: bool = False, generate_heavy_chi_samples: bool = False) LigandPreparation[source]#
Build a
LigandPreparationfrom a SMILES-derived ligand.This is the final, naming-and-typing step of the unified pipeline. Its input must already be fully resolved chemistry: explicit hydrogens at the desired protonation state and authoritative per-atom partial charges (the OpenBabel MMFF94 charges produced by the SMILES -> mol2 step). Protonation and charge generation are not done here – they happen upstream in
tmol.ligand._detect.nonstandard_residue_info_from_smiles_via_mol2().Charges are mapped onto atoms by stable RDKit index (source atom order), so they are independent of the atom renaming below and never recomputed.
Returns a
LigandPreparation– the same structtmol.ligand._params_file.load_params_file()produces for each residue defined in a.tmolfile, so the AtomArray-driven path and the params-file path converge on a single abstraction thatinject_ligand_preparations()consumes.- Parameters:
ligand_info – A SMILES-derived ligand (
skip_protonation=Truewith authoritativepartial_charges). Raw CIF/atom-array ligands must be routed throughprepare_ligands()/prepare_ligand_from_cif().name_source – Optional ligand whose atom names the prepared residue should adopt (mapped to the prepared heavy atoms via the atom-order map). On the unified CIF path this is the original CIF ligand. Defaults to
ligand_info.
- Raises:
ValueError – If
ligand_infolacks explicit hydrogens / authoritative charges (there is no charge-generation fallback).
- tmol.ligand.recombine_fragmented_ligands(structure: AtomArray | AtomArrayStack, pose_stack) AtomArray | AtomArrayStack[source]#
Restore original residue identities on exported ligand fragments.
Uses
pose_stack.split_block_mappingto identify which residues in structure are split-block fragments and what their original PDB identity should be. Atoms are matched by the residue label stored in the pose’s PDBInfo (which is whatbiotite_from_pose_stackwrites tores_id).
- tmol.ligand.unsplit_pose_stack(pose_stack)[source]#
Reconstruct an unsplit PoseStack from one containing split (fragment) blocks.
Each group of split blocks sharing a
(pose_ind, group_ind)in the PoseStack’ssplit_block_mappingis collapsed back into a single block of the corresponding original block type. Atom coordinates are gathered from the fragment blocks using the per-entrysplit_to_orig_atom_indsarrays; atoms present in the original type but absent from all fragments are left at zero.Inter-residue connections between fragment blocks are removed (they become intra-block interactions in the original). External connections from a fragment block to a non-fragment block are transferred to the reconstructed original block.
Returns a new PoseStack with
split_block_mapping=None.
- tmol.ligand.write_params_file(preparation: LigandPreparation | list[LigandPreparation], path: str | Path) None[source]#
Write a ligand
LigandPreparationas a tmol.tmolfile.- Parameters:
preparation – A
LigandPreparation(itsresidue_type/partial_charges/cartbonded_paramsare used), or a list of them.path – Output file, holding every supplied residue.
- tmol.ligand.write_params_from_mol2(mol2_path: str | Path, out_path: str | Path, *, res_name: str | None = None, ph: float = 7.4, mode: str = 'auto', seed: int | None = None) None[source]#
Build params from a mol2 file and write a tmol
.tmolfile.- Parameters:
mol2_path – Input Tripos MOL2 or MDL SDF; see
prepare_ligand_from_mol2().out_path – Output file path (see
write_params_file()).res_name – Optional residue name override.
ph – Target pH for protonation.
mode –
"keep","auto"(default), or"regenerate", as inprepare_ligand_from_mol2().seed – Reproducible conformer seed when charges must be generated.
Public constants#
|
str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str |
|
int([x]) -> integer int(x, base=10) -> integer |
|
int([x]) -> integer int(x, base=10) -> integer |
|
str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str |