Input and Output#
The public tmol.io API converts common structure representations to and
from tmol.pose.PoseStack.
Once the residue types are resolved, tmol.io.pose_stack_from_atom37()
builds a pose from tensors alone, for canonical and noncanonical chemistry
alike: residue identity comes from res_types and any further covalent
chemistry from covalent_bonds. Its slot layout is built by
tmol.io.atom37_slot_map_for_ordering(). The direct AtomWorks adapter is
the canonical-amino-acid special case of that path. To read chemistry from a
structure instead, tmol.io.pose_stack_from_atom37_and_topology() derives
the same information from general Biotite topology, including nucleic acids and
ligands. Repeated diffusion,
guidance, and search workloads should bind their fixed topology once with
tmol.io.prepare_atom37_pose_builder(); its returned callable accepts
each coordinate batch and preserves ordinary finite authored protein hydrogens
by default. Generated ligand types may rebuild hydrogens unless
trust_hydrogen_names=True.
Use tmol.io.atom_array_from_file() or
tmol.io.pose_stack_from_file() for PDB, CIF, compressed CIF and binary CIF.
The reader options are model and assembly_id. Reading preserves
supplied names, coordinates, hydrogens and covalent bonds; declared unresolved
heavy atoms carry NaN coordinates. Chemistry the file leaves out is completed
from the component dictionary. Missing chemistry is allowed at this stage.
Known residues need only names and coordinates. Pass an existing param_db or
load prepared ligand_params_files to reuse ligand parameters with a coordinate-only
PDB/CIF. With prepare_ligands=True, pose preparation generates only missing
parameters and reports the residue name when required bond orders are absent.
Use return_context=True to retain the resulting parameter database for scoring
and repeated construction. Direct AtomArray inputs follow the same contract.
# Reuse known chemistry; the PDB needs only matching names and coordinates.
pose = pose_stack_from_file(
"complex.pdb", device, ligand_params_files=["ligand.tmol"],
)
# Prepare missing chemistry once, then reuse context.parameter_database.
pose, context = pose_stack_from_file(
"complex.cif", device, prepare_ligands=True, return_context=True,
)
tmol.io.pose_stack_from_pdb() accepts paths, text and lists of lines through
the same file path. File and AtomArray inputs preserve matching supplied coordinates
for ordinary protein residues. Pose construction builds missing hydrogens but
preserves finite authored protein hydrogen coordinates by default. Pass
no_optH=False to request OptH packing. Generated ligand types may rebuild
hydrogens whose names changed during parameter generation; pass
trust_hydrogen_names=True only when those names match the prepared database.
Low-level PDB atom-record/DataFrame utilities retain their canonical-only contract.
Structure conversion between external formats and TMol poses.
- tmol.io.add_metal_coordination(canonical_ordering: CanonicalOrdering, pose_stack: PoseStack, pose: int, metal: int, donor: int, atom: str, site: int | None = None) PoseStack[source]#
Bond
atomofdonorto an open site ofmetal(by default the one whose virtual faces it best); the donor atom loses its hydrogens.
- tmol.io.assemble_input(*parts: AtomArray) AtomArray[source]#
Join separately-read parts with their bonds, annotations and templates; a chain ID claimed by an earlier part is moved to a free one.
- Parameters:
*parts – Structures to join, in the order they should appear.
- Returns:
The joined structure.
- Raises:
ValueError – If no parts are given, or chain IDs cannot be made unique.
Examples
>>> complex = assemble_input(protein, ligand) >>> sorted(set(complex.chain_id)) ['A', 'L']
- tmol.io.atom_array_from_file(cif_path, *, model: int = 1, assembly_id: str | None = None)#
Read a PDB, CIF, compressed CIF, binary CIF, MOL2 or SDF into an AtomArray.
Preserve supplied names, coordinates, hydrogens and covalent connections. Declared unresolved heavy atoms have NaN coordinates. Chemistry the file leaves out is supplemented from the component dictionary.
Missing chemical metadata is allowed here. Pose construction uses existing parameters for known residues and requires additional chemistry only when preparing an unknown residue. Pose construction controls hydrogen optimization and whether to trust regenerated ligand hydrogen names.
- tmol.io.atom_array_from_mol2(mol2_path: str | Path, *, res_name: str | None = None, chain_id: str = 'L', res_id: int = 1) AtomArray[source]#
Read a Tripos MOL2 (or MDL SDF/MOL) ligand as an AtomArray with the file’s bond orders, formal charges and aromatic flags, registered as its own component template.
- Parameters:
mol2_path – Path to the MOL2, SDF or MOL file.
res_name – Component code to give the ligand. Taken from the file’s substructure record when None.
chain_id – Chain to place the ligand on.
res_id – Residue number within that chain.
- Returns:
A single-residue AtomArray with a populated bond list.
Examples
>>> ligand = atom_array_from_mol2("ligand.mol2") >>> ligand.bonds.get_bond_count() > 0 True
- tmol.io.cif_from_atom_array(atom_array: AtomArray, *, path: str | Path | None = None, entry_id: str = 'assembled') str | Path[source]#
Write a structure as a CIF carrying its own component chemistry: known components from the dictionary, others (a MOL2 ligand) from the structure itself.
- Parameters:
atom_array – Structure to write.
path – Where to write. Returns the CIF text instead when None.
entry_id – Data block name.
- Returns:
The written path, or the CIF text when
pathis None.
Examples
>>> text = cif_from_atom_array(assemble_input(protein, ligand)) >>> "_chem_comp_bond" in text True
- exception tmol.io.Atom37MappingError[source]#
Bases:
ValueErrorAn AtomArray cannot be routed unambiguously into an Atom37 tensor.
- class tmol.io.CanonicalForm(chain_id: Tensor[slice(None, None, None), slice(None, None, None)], res_types: Tensor[slice(None, None, None), slice(None, None, None)], coords: Tensor[slice(None, None, None), slice(None, None, None), slice(None, None, None), 3], res_labels: NDArray[slice(None, None, None), slice(None, None, None)], residue_insertion_codes: NDArray[slice(None, None, None), slice(None, None, None)], chain_labels: NDArray[slice(None, None, None), slice(None, None, None)], atom_occupancy: NDArray[slice(None, None, None), slice(None, None, None), slice(None, None, None)] | None, atom_b_factor: NDArray[slice(None, None, None), slice(None, None, None), slice(None, None, None)] | None, disulfides: Tensor[slice(None, None, None), 3] | None, res_not_connected: Tensor[slice(None, None, None), slice(None, None, None), 2] | None, cyclic_bonds: Tensor[slice(None, None, None), 3] | None = None, covalent_bonds: Tensor[slice(None, None, None), 5] | None = None, metal_sites: Tensor[slice(None, None, None), 3] | None = None, metal_coordination: Tensor[slice(None, None, None), 5] | None = None, metal_origins: NDArray[slice(None, None, None), slice(None, None, None)] | None = None, residue_annotations: ndarray | None = None, protonation_variants: Tensor[slice(None, None, None), slice(None, None, None)] | None = None)[source]#
Bases:
objectThis class holds the data that describe a (stack of) structure(s) in a poised, ready-to-use state.
This datastructure holds the information necessary to determine the chemical identities of the residues in the structure(s), which may be under-determined from tmol’s perspective by the source of the structure (e.g. OpenFold does not explicitly model termini). The atoms that are present are represented with non-NaN coordinates in the coords array; the order in which those atoms appear is given by a particular CanonicalOrdering object.
The datastructure also holds convenience information such as author-provided residue labels (ints), chain labels (strings) & insertion codes (strings) as well as the occupancy and B-factor of each atom. These are not strictly necessary but are often useful when processing structures.
- class tmol.io.CanonicalOrdering(max_n_canonical_atoms: int, restype_io_equiv_classes: Tuple[str, ...], restypes_ordered_atom_names: Mapping[str, Tuple[str, ...]], restypes_atom_elements: Mapping[str, Mapping[str, str]], restypes_atom_index_mapping: Mapping[str, Mapping[str, int]], restypes_mainchain_atoms: Mapping[str, Tuple[str, ...] | None], restypes_required_mainchain_atoms: Mapping[str, Tuple[str, ...] | None], restypes_default_termini_mapping: Mapping[str, Tuple[str | None, str | None]], down_termini_patches: Tuple[str, ...], up_termini_patches: Tuple[str, ...], termini_patch_added_atoms: Mapping[Tuple[str, str], Tuple[str, ...]], termini_only_atoms: Mapping[Tuple[str, str], Tuple[str, ...]], termini_only_atoms_by_class: Mapping[str, Tuple[str, ...]], cys_inds: CysSpecialCaseIndices, his_inds: HisSpecialCaseIndices, polymer_conn_inds: PolymerConnectionIndices, name3_aliases: Mapping[str, str] = NOTHING, restypes_virtual_atoms: Mapping[str, FrozenSet[str]] = NOTHING)[source]#
Bases:
objectThe canonical ordering class describes the integer ordering of residue types and for atoms within those residue types for the collection of available residue types defined by a PatchedChemicalDatabase.
The canonical ordering class’s purpose is to enable creation of a “canonical form” dictionary that describes a molecular system in the way that tmol expects in order to construct a PoseStack.
There is no “canonical form” dictionary is simply a dictionary holding the at-least-three-but-as-many-as-eight arguments to tmol.io._pose_stack_construction.pose_stack_from_canonical_form after the first two. That is, it must contain “chain_id”, “res_types” and “coords” entries.
When constructing a PoseStack, there are multiple residue types for each “equivalence class” (think 3-letter code); e.g. for “CYS” there’s the standard middle-of-a-polypeptide-chain CYS, the standard middle-of-a-polypeptide-chain disulfide-forming CYS, and then for those two, four variants for the N-, C-, and both-N- and-C terminal forms; eight total options for a single “CYS” three-letter code. tmol collects all of the various forms of a single equivalence class and creates a list of all atom names across all the residue types for it. You can then provide tmol the set of atoms that are present at a given position by giving a non-NaN coordinate for that entry in an [n-poses x max-n-res x max-ats-per-res x 3] tensor of coordinates. Atoms with NaN coordinates are taken as possibly present in the residue type; tmol will decide the best fit for which residue type to use at each position. If an atom is provided to tmol and it is not present for a given residue type, then that residue type will be disqualified from consideration. Thus an important part of telling tmol which atoms are present is mapping from an atom name to an index for that atom. The CanonicalOrdering object is where that mapping is encoded. It also handles the mapping from alternate-atom-name to canonical-form-atom index; e.g. in PDBv2, glycine’s two hydrogens were named “HA1” and “HA2”, but in PDBv3, they are named “1HA” and “2HA.” So that we can parse PDB files written in PDBv2 and PDBv3, we have an idea of an “alias” for an atom; see the restypes_atom_index_mapping data member.
Four data members are especially useful:
max_n_canonical_atomsrestype_io_equiv_classesrestypes_ordered_atom_namesrestypes_atom_index_mapping
The remaining data members are primarily for internal tmol functionality.
- max_n_canonical_atoms: the largest number of distinct atom names among all
variants of a single residue type (equivalence class) across all residue types
- restype_io_equiv_classes:
essentially the list of 3-letter codes for the residue types that are readable; use the index function (e.g. co.restype_io_equiv_classes.index(“TRP”)) to obtain the integer meant to represent each restype
- restypes_ordered_atom_names:
the ordered list of the names of each atom for every allowed residue type; does not include the alternate names for atoms. Atoms should be given to tmol in this order; e.g. by putting the coordinate of the ith atom in the ith entry of the coordinate tensor (e.g. coords[p, r, i] for pose p, residue r)
- restypes_atom_index_mapping:
mapping for each name3 from atom name and atom name alias to the index of that atom for every allowed residue type in the restypes_ordered_atom_names list; this is probably more useful than the restypes_ordered_atom_names list, especially if you are using the PDBv2 naming convention (as Rosetta3 does) instead of the PDBv3 convention.
- class tmol.io.PoseBuildContext(canonical_ordering: CanonicalOrdering, packed_block_types: PackedBlockTypes, parameter_database: ParameterDatabase, restype_set: ResidueTypeSet, fragment_definitions: tuple[LigandFragmentDefinition, ...]=(), ligand_names: dict[str, str]=<factory>, cut_covalent_partners: frozenset[str] = frozenset({}))[source]#
Bases:
objectImmutable, structure-independent construction context.
Holds only the pieces that depend on the parameter database / ligand set (not on any particular input), so it can be built once and reused across many inputs that share the same ligand(s).
- class tmol.io.PreparedAtom37PoseBuilder(context: PoseBuildContext, canonical_template: CanonicalForm, mapped_token_id: Tensor, mapped_slot: Tensor, mapped_residue: Tensor, mapped_atom: Tensor, max_token_id: int, required_mainchain_entries: tuple[tuple[int, int, str], ...], fragment_mapping: FragmentedLigandPoseMapping | None = None, topology_cache_safe: bool = True, pose_topologies: dict[int, _PreparedAtom37PoseTopology] = NOTHING)[source]#
Bases:
objectBind immutable Biotite topology for repeated Atom37 pose construction.
Calls accept float32 coordinates shaped
[n_poses, n_tokens, 37, 3]on the context’s device. The first call for a batch size prepares its fixed pose topology; up to four recently used batch sizes are cached. Inputs with coordinate-dependent atom presence or ambiguous histidine hydrogens use the normal uncached construction path.The builder owns a mutable topology cache and is not safe for concurrent calls. Use one builder per calling thread when pose construction overlaps.
- tmol.io.atom_records_from_coords(pbt: PackedBlockTypes, chain_ind_for_block: Tensor[slice(None, None, None), slice(None, None, None)], block_types64: Tensor[slice(None, None, None), slice(None, None, None)], pose_like_coords: Tensor[slice(None, None, None), slice(None, None, None), 3], block_coord_offset: Tensor[slice(None, None, None), slice(None, None, None)], residue_labels: NDArray[slice(None, None, None), slice(None, None, None)] | None, residue_insertion_codes: NDArray[slice(None, None, None), slice(None, None, None)] | None, chain_labels: NDArray[slice(None, None, None), slice(None, None, None)] | None, atom_occupancy: NDArray[slice(None, None, None), slice(None, None, None)], atom_b_factor: NDArray[slice(None, None, None), slice(None, None, None)]) NDArray[source]#
Create a numpy array holding the atom records needed to write a PDB file from the coordinates and block types of a stack of structures, laid out in pose-stack form.
- tmol.io.atom_records_from_pose_stack(pose_stack: PoseStack, merge_fragments: bool = True) NDArray[source]#
Create a numpy array holding the atom records needed to write a PDB file from a PoseStack.
Fragmented ligands use their original residue identity by default. Pass
merge_fragments=Falseto retain the separate fragment residue numbers.
- tmol.io.atomworks_from_pose_stack(pose_stack: PoseStack) tuple[source]#
Convert a PoseStack back to atomworks UNIFIED_ATOM37_ENCODING tensors.
- Parameters:
pose_stack (PoseStack) – The PoseStack to convert. Must contain only standard amino acids.
- Returns:
coords (Tensor, shape [n_poses, max_n_res, 37, 3]) – Atom coordinates in the atomworks atom37 layout. Absent atoms are 0.
residue_type (Tensor[int64], shape [n_poses, max_n_res]) – Atomworks token indices (1..20 for real residues, 0 for padding).
chain_iid (Tensor[int64], shape [n_poses, max_n_res]) – Chain identifiers.
- tmol.io.biotite_from_canonical_form(cf: CanonicalForm, co: CanonicalOrdering | None = None, include_virtual_atoms: bool = False) AtomArray | AtomArrayStack[source]#
Export coordinates and chemical/author labels using a shared atom layout.
Multi-model arrays require identical residue and atom annotations. Their atom layout is the union of resolved atoms; absent coordinates remain NaN. Missing author labels default to internal chain IDs and one-based residue positions. Virtual atoms are left out unless
include_virtual_atoms. This host-array export detaches coordinates from autograd. A canonical form carries no residue types, so the result has no bonds; biotite_from_pose_stack adds them.
- tmol.io.biotite_from_pose_stack(pose_stack: PoseStack, co: CanonicalOrdering | None = None, merge_fragments: bool = True, include_virtual_atoms: bool = False) AtomArray | AtomArrayStack[source]#
Convert PoseStack back to Biotite structure.
- Parameters:
pose_stack – Pose stack to convert.
co – Canonical ordering used for conversion. Provide the ordering that was used when ligands or custom residue types are present.
merge_fragments – Restore fragmented ligands to their original residue identity. Set to False to keep fragment residues separate.
include_virtual_atoms – Also write virtual atoms, such as a metal’s site virtuals. A metal split out of a component is then left as its own residue, which its virtuals belong to; otherwise it is put back in the component it came from.
- Returns:
Biotite AtomArray for single-pose or AtomArrayStack for multi-pose, with every bond the poses’ residue types and connections declare: bond orders within residues, and chemical bonds between them, metal coordination typed COORDINATION.
- tmol.io.build_context_from_biotite(biotite_structure: AtomArray | AtomArrayStack, torch_device: device, param_db: ParameterDatabase | None = None, prepare_ligands: bool = False, ligand_ph: float = 7.4, strict_atom_types: bool = False, strict_ligands: bool = True, ligand_params_files: list[str] | None = None, chem_comp_types: dict | None = None, ligand_seed: int | None = None) PoseBuildContext[source]#
Build the structure-independent construction context.
The returned context holds only database/ligand-derived pieces (canonical ordering, residue-type set, packed block types, parameter database); it does not depend on the input structure’s coordinates and can be reused across structures sharing the same ligand(s).
biotite_structureis used only to detect and prepare ligands (whenprepare_ligands=True).- Parameters:
biotite_structure – Input AtomArray or AtomArrayStack. Used only for ligand detection/preparation when
prepare_ligands=True.torch_device – Target torch device.
param_db – Optional parameter database. When provided, canonical ordering, residue types, and packed block types are built from this database. If prepare_ligands=True, it is extended with ligand data. If None, defaults are used.
prepare_ligands – If True, detect and prepare non-standard residues (via
tmol.ligand, which uses RDKit for atom typing and residue-type construction).ligand_ph – pH AtomWorks protonates residues lacking hydrogens at before ligands are prepared (default 7.4, only used when prepare_ligands=True).
strict_atom_types – If True, unknown ligand atom types raise errors instead of using a fallback element heuristic.
strict_ligands – If True (default), raise when a detected ligand cannot be prepared and registered (instead of silently dropping it during pose construction). Pass False to fall back to warn-and-skip. Only used when prepare_ligands=True.
ligand_params_files – Optional list of tmol YAML params file paths. Residues defined in these files skip the RDKit/OB pipeline.
chem_comp_types –
{comp_id: type}from the input file’s_chem_comptable (seetmol.ligand.chem_comp_types_from_cif), which says whether a residue belongs to a polymer where the file does not number it along a sequence. Only used when prepare_ligands=True.ligand_seed – Fixed RNG seed for the conformer each prepared residue is built from, making preparation reproducible. Only used when prepare_ligands=True.
- Returns:
PoseBuildContext containing canonical ordering, packed block types, parameter database, and residue type set.
- tmol.io.canonical_form_from_atomworks(coords: Tensor, residue_type: Tensor, chain_iid: Tensor) CanonicalForm[source]#
Build a CanonicalForm from atomworks UNIFIED_ATOM37_ENCODING tensors.
- Parameters:
- Return type:
- tmol.io.canonical_form_from_biotite(biotite_structure: AtomArray | AtomArrayStack, torch_device: device, co: CanonicalOrdering | None = None, missing_density_distance_threshold: float = 2.4, atom37_coords: Tensor | None = None, _filter_missing_mainchain: bool = True, _cut_covalent_partners: frozenset[str] = frozenset({})) CanonicalForm[source]#
Convert a Biotite AtomArray or AtomArrayStack to a CanonicalForm.
This function bridges between Biotite’s data structures and tmol’s internal representation by converting atom and residue information from string-based identifiers to tmol’s canonical integer-based indexing system.
- Parameters:
biotite_structure – A Biotite AtomArray (single structure) or AtomArrayStack (multiple structures) containing the molecular data. Must contain atom coordinates, residue names, atom names, chain IDs, and optionally B-factors and occupancy values.
torch_device – PyTorch device (e.g., torch.device(‘cuda’) or torch.device(‘cpu’)) where the resulting tensors should be allocated.
co – A CanonicalForm in case you want to use a non-default database (and thus may need a different mapping)
missing_density_distance_threshold – Maximum distance in angstroms for treating a polymer gap as missing density rather than a chain break.
atom37_coords – Optional autograd-tracked coordinate tensor of shape [n_poses, n_tokens, 37, 3]. When provided, coordinates are sourced from this tensor (routed by the
token_idandatom37_slotannotations onbiotite_structure); all-NaN mapped triplets are treated as missing while unmapped finite reference atoms remain context. The geometry-based missing-density check is skipped so the topology stays fixed and gradients flow. Seepose_stack_from_atom37_and_topology().
- Returns:
- A data structure containing:
chain_id: Tensor mapping residues to chain indices
res_types: Tensor mapping residues to tmol residue type indices
coords: 4D tensor of atomic coordinates (poses x residues x atoms x 3)
res_labels: Original residue sequence numbers from the structure
residue_insertion_codes: PDB insertion codes for residues
chain_labels: Original chain identifiers from the structure
atom_occupancy: Optional tensor of atom occupancy values
atom_b_factor: Optional tensor of atom B-factor values
disulfides: Explicit cysteine sulfur bonds, including unresolved SG
res_not_connected: Tensor describing whether two consecutive residues should be treated as chemically bonded.
- Return type:
- tmol.io.canonical_form_from_pdb(canonical_ordering: CanonicalOrdering, pdb_lines_or_fname: str | List, device: device, *, residue_start: int | None = None, residue_end: int | None = None, res_not_connected: Tensor[slice(None, None, None), slice(None, None, None), 2] | None = None) CanonicalForm[source]#
Create a canonical form from either the contents of a PDB file as one long string or a list of individual lines from the file or by providing the name/path of a PDB file
pdb_lines_or_fname must either be a list of the lines in a PDB file or a string representing a file
- tmol.io.canonical_form_from_pose_stack(canonical_ordering: CanonicalOrdering, pose_stack: PoseStack, chain_id=None)[source]#
Convert a pose stack to canonical residue and atom tensors.
- Parameters:
canonical_ordering – Residue and atom ordering for the output tensors.
pose_stack – Poses to deconstruct.
chain_id – Optional integer chain identifiers shaped
[pose, residue].
- Returns:
Canonical form containing coordinates, residue metadata, and connectivity.
- tmol.io.canonical_ordering_for_atomworks() CanonicalOrdering#
Construct the CanonicalOrdering for the protein subset used by the atomworks UNIFIED_ATOM37_ENCODING.
- tmol.io.canonical_ordering_for_biotite() CanonicalOrdering#
Construct the CanonicalOrdering object to use for Biotite. This wont be used as a typical CanonicalOrdering object, since we aren’t mapping from int-to-int, and instead are going from string-to-int.
- tmol.io.create_pose_stack_from_sequences(seqs, packed_block_types: PackedBlockTypes | None = None, device: device | None = None, param_db=None, termini: bool = True, context: PoseBuildContext | None = None, return_context: bool = False)[source]#
Construct a PoseStack with zero coordinates from sequence strings.
See tmol.pose._sequence for the grammar. Returns (PoseStack, PoseBuildContext) when return_context is set; the context carries the database extended with any ligands the sequence names.
- tmol.io.default_canonical_ordering() CanonicalOrdering#
Create a CanonicalOrdering object from the default set of residue types
- tmol.io.default_packed_block_types(device: device) PackedBlockTypes[source]#
Create a PackedBlockTypes object from the default set of residue types
- tmol.io.extended_pose_stack_from_sequences(seqs, device: device | None = None, param_db=None, termini: bool = True, context=None, return_context: bool = False)[source]#
Build a PoseStack from sequences with ideal geometry and extended backbone torsions.
See tmol.pose._sequence for the grammar. Returns (PoseStack, PoseBuildContext) when return_context is set.
- tmol.io.packed_block_types_for_atomworks(device: device) PackedBlockTypes#
Construct the PackedBlockTypes for the protein subset used by the atomworks UNIFIED_ATOM37_ENCODING.
- tmol.io.packed_block_types_for_biotite(device: device) PackedBlockTypes#
Construct the PackedBlockTypes (PBT) object that will used for Biotite. We’ll use the defaults since anything might show up in a Biotite AtomArray. Some things may show up in the AtomArrays that are not handled by this PBT, but that is work for the future.
- tmol.io.atom37_slot_map_for_ordering(canonical_ordering: CanonicalOrdering, atom_names_by_slot: Mapping[str, Sequence[str]], device: device) Tensor[source]#
Map each residue type’s Atom37 slots to canonical atom indices.
The mapping is a property of the residue-type table and the slot layout, not of any structure, so it is built once per (ordering, layout) pair.
Atoms that a terminus patch adds are deliberately left unmapped. They exist only on terminal residue variants, and offering one on an interior residue would disqualify that residue’s block types; tmol applies termini itself from the chain boundaries. Which atoms those are is read from the chemical database rather than named here.
- Parameters:
canonical_ordering – Residue and atom ordering the
res_typesindex into.atom_names_by_slot – Atom name occupying each slot, per residue name3. An empty name marks an unused slot.
device – Device for the returned tensor.
- Returns:
[n_residue_types, n_slots]canonical atom index per slot,-1where the slot is unused or the atom is not part of that residue type.
Examples
>>> slot_map = atom37_slot_map_for_ordering(co, layout, torch.device("cpu")) >>> slot_map.shape[1] 37
- tmol.io.canonical_form_from_atom37(atom37_coords: Tensor, res_types: Tensor, chain_id: Tensor, canonical_ordering: CanonicalOrdering, *, slot_map: Tensor | None = None, disulfides: Tensor | None = None, cyclic_bonds: Tensor | None = None, covalent_bonds: Tensor | None = None, res_not_connected: Tensor | None = None) CanonicalForm[source]#
Scatter Atom37 coordinates into a canonical form.
- Parameters:
atom37_coords –
[n_poses, n_tokens, n_slots, 3], autograd-tracked.res_types –
[n_poses, n_tokens]index intocanonical_ordering;-1marks padding.chain_id –
[n_poses, n_tokens]chain index per token.canonical_ordering – Ordering the
res_typesindex into.slot_map –
[n_residue_types, n_slots]fromatom37_slot_map_for_ordering(). Defaults to the standard Atom37 layout, which covers the canonical amino acids; supply one built from an extended layout for other chemistry.disulfides –
[n, 3]explicit(pose, res1, res2)rows.cyclic_bonds –
[n, 3]explicit head-to-tail closures.covalent_bonds –
[n, 5]explicit(pose, res1, atom1, res2, atom2)links in canonical ordering.res_not_connected –
[n_poses, n_tokens, 2]suppressed polymer connections.
- Returns:
A canonical form whose coordinates carry gradients back to
atom37_coords.
- tmol.io.pose_stack_from_atom37(atom37_coords: Tensor, res_types: Tensor, chain_id: Tensor, context: PoseBuildContext, *, slot_map: Tensor | None = None, disulfides: Tensor | None = None, cyclic_bonds: Tensor | None = None, covalent_bonds: Tensor | None = None, res_not_connected: Tensor | None = None, no_optH: bool = True, **kwargs: object) PoseStack[source]#
Build a differentiable PoseStack from Atom37 tensors alone.
One path for canonical and noncanonical poses. Residue identity comes from
res_typesindexed into the context’s canonical ordering, and any covalent chemistry the polymer connections do not already describe comes fromcovalent_bonds; noAtomArrayis consulted.- Parameters:
atom37_coords –
[n_poses, n_tokens, n_slots, 3], autograd-tracked.res_types –
[n_poses, n_tokens]index intocontext.canonical_ordering;-1marks padding.chain_id –
[n_poses, n_tokens]chain index per token.context – Chemistry resolved once, from any of the
build_context_from_*helpers.slot_map – Atom37 slot layout; see
atom37_slot_map_for_ordering().disulfides – Explicit
(pose, res1, res2)rows.cyclic_bonds – Explicit head-to-tail closures.
covalent_bonds – Explicit
(pose, res1, atom1, res2, atom2)links.res_not_connected – Suppressed polymer connections.
no_optH – Preserve finite input hydrogens rather than optimizing them.
- Returns:
A pose whose coordinates carry gradients back to
atom37_coords.
Examples
>>> context = build_context_from_biotite(array, device) >>> pose = pose_stack_from_atom37(coords, res_types, chain_id, context) >>> pose.coords.requires_grad True
- tmol.io.pose_stack_from_canonical_aa_atom37(coords: Tensor, residue_type: Tensor, chain_iid: Tensor, **kwargs) PoseStack[source]#
Build a PoseStack from atomworks UNIFIED_ATOM37_ENCODING tensors.
This is the tensor-only Atom37 entry point. Residue identity comes from
residue_typealone, so nothing but tensors is required – but for the same reason it is limited to the canonical amino acids with the canonical n- and c-termini patches.For anything else – noncanonical residues, ligands, nucleic acids, or covalent links – residue identity cannot be read off a 1..20 token, and
pose_stack_from_atom37_and_topology()is the entry point; it takes the chemistry from a supplied topology instead.Both build a
CanonicalFormand finish in the same constructor, so the choice here is only about how residue identity is supplied, not about how the pose is built.- Parameters:
coords (Tensor, shape [batch, n_res, 37, 3]) – Atom coordinates in the atomworks atom37 layout.
residue_type (Tensor[int64], shape [batch, n_res]) – Atomworks token indices. Must be in 1..20 (standard protein only).
chain_iid (Tensor[int64], shape [batch, n_res]) – Chain identifiers (integer IDs, not string labels).
**kwargs – Additional arguments passed to
pose_stack_from_canonical_form.
- Return type:
- Raises:
ValueError – If any
residue_typevalue is outside 1..20 (protein-only).
- tmol.io.pose_stack_from_file(cif_path, device, *, model: int = 1, assembly_id: str | None = None, residue_start: int | None = None, residue_end: int | None = None, **kwargs)#
Construct a PoseStack from a PDB, CIF or binary/compressed CIF structure.
Reads the structure with
atom_array_from_cif(), so the residues carry every atom their chemistry declares and the unresolved ones arrive at NaN for pose construction to rebuild. Further keyword arguments are passed totmol.io.pose_stack_from_biotite().
- tmol.io.pose_stack_from_canonical_form_and_context(cf: CanonicalForm, context: PoseBuildContext, *, no_optH: bool, atom37_coords: Tensor | None, fragment_mapping=None, return_context: bool = False, packer_seed: int | None = None, **kwargs: object) PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#
Build a pose from a canonical form and a reusable build context.
This is the single construction path. Every entry point – canonical amino-acid Atom37 tensors, an Atom37 batch bound to a fixed topology, a parsed CIF/PDB structure – reduces its input to a
CanonicalFormplus aPoseBuildContextand finishes here. The entry points differ only in how they derive those two objects, not in how a pose is built from them.The context carries the canonical ordering and packed block types, so noncanonical residues, ligands, and covalent links are handled by the same code as standard amino acids: nothing here inspects an
AtomArray. Callers that already hold a canonical form and a context – notably repeated guidance or search steps over one fixed topology – should call this directly rather than re-deriving topology per batch.- Parameters:
cf – Canonical-form tensors describing residue identity and coordinates.
context – Structure-independent chemistry resolved once.
no_optH – Preserve finite input hydrogens instead of optimizing them.
atom37_coords – When supplied, the autograd-tracked source of
cf’s coordinates; retained so gradients survive hydrogen rebuilding.fragment_mapping – Mapping produced when fragmented ligands were expanded.
return_context – Also return the context used.
packer_seed – Seed of the packer, as for
pose_stack_from_biotite.
For ordinary structure inputs, coincident bonded heavy atoms raise before packing; hydrogens coincident with their parent are rebuilt without changing protonation. Differentiable
atom37_coordsretain their supplied geometry.- Returns:
The constructed pose, optionally with atom mappings or the context.
- tmol.io.prepare_atom37_pose_builder(biotite_structure: AtomArray | AtomArrayStack, context: PoseBuildContext) PreparedAtom37PoseBuilder[source]#
Prepare a callable for repeatedly binding Atom37 coordinates to topology.
This is the campaign-oriented counterpart to
pose_stack_from_atom37_and_topology(): immutable residue identity, connectivity, fragmentation, and Atom37 routing are resolved once. Calling the returned builder with a coordinate tensor constructs a differentiable pose while retaining TMol’s usual missing-atom behavior. The returned builder preserves finite input hydrogens by default; passopt_h=Trueto optimize them after construction.
- tmol.io.remove_metal_coordination(canonical_ordering: CanonicalOrdering, pose_stack: PoseStack, pose: int, metal: int, site: int) PoseStack[source]#
Open one filled site of a metal; the donor keeps its hydrogens.
- tmol.io.pose_stack_from_atom37_and_topology(atom37_coords: Tensor, biotite_structure: AtomArray | AtomArrayStack, context: PoseBuildContext, no_optH: bool = True, **kwargs) PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#
Build a differentiable PoseStack from atom37 coordinates and a topology.
This is the topology-supplied Atom37 entry point. Unlike
pose_stack_from_canonical_aa_atom37(), which reads residue identity from a 1..20 token tensor and so covers only canonical amino acids, this supports any chemistry shared by AtomWorks and the supplied TMol context (including ordinary ligands and nucleic acids): the chemical topology is taken frombiotite_structurewhile mapped coordinates come from the autograd-trackedatom37_coordstensor.The topology is only read to derive residue identity and the Atom37 slot map. Once derived, construction is pure tensor work in the shared
pose_stack_from_canonical_form_and_context(). When the same topology is scored repeatedly, useprepare_atom37_pose_builder()to pay that derivation once; the returned builder never touches theAtomArrayagain. Unmapped finite reference atoms remain context. This is the entry point for differentiable scoring/guidance over atomized inputs, where the same fixed topology is scored repeatedly as coordinates move.Build
contextonce withbuild_context_from_biotite()(withprepare_ligands=Truewhen ligands are present). For repeated diffusion or search steps, bind the topology once withprepare_atom37_pose_builder()and call the returned builder with each coordinate batch. Mapped reference coordinates are ignored, so one reference topology can be reused as those coordinates change. It must carry two integer annotations that map each atom into the atom37 tensor:token_id(the token axis) andatom37_slot(the 0..36 slot).Topology is derived from chemical identity alone –
missing_densitybreaks and automatic disulfide detection (both coordinate-dependent) are disabled – so the block types, termini, and atom count stay fixed as coordinates change. This adapter does not classify or allowlist residue types or elements: newly supported PTMs, ions, and metals work through the same API once they are represented by the supplied context and canonical ordering.- Parameters:
atom37_coords (Tensor, shape [n_poses, n_tokens, 37, 3]) – Autograd-tracked coordinates in the atomworks atom37 layout.
biotite_structure (biotite AtomArray) – Reference topology carrying
token_idandatom37_slotannotations.context (PoseBuildContext) – Structure-independent context from
build_context_from_biotite().no_optH (bool) – Preserve finite input hydrogens and leave newly built hydrogens at ideal positions when True (default). Pass False to run TMol’s hydrogen optimization pipeline.
**kwargs – Additional arguments forwarded to
pose_stack_from_biotite.
- Returns:
Whose
coordscarry gradients back toatom37_coords.- Return type:
- tmol.io.pose_stack_from_biotite(biotite_structure: AtomArray | AtomArrayStack, torch_device: device, param_db: ParameterDatabase | None = None, missing_density_distance_threshold: float = 2.4, no_optH: bool = True, prepare_ligands: bool = False, ligand_ph: float = 7.4, strict_atom_types: bool = False, strict_ligands: bool = True, ligand_params_files: list[str] | None = None, chem_comp_types: dict | None = None, ligand_seed: int | None = None, packer_seed: int | None = None, return_context: bool = False, context: PoseBuildContext | None = None, atom37_coords: Tensor | None = None, **kwargs: object) PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#
Build a PoseStack from the output generated by Biotite.
To score many structures that share the same ligand(s) efficiently, build the (expensive, structure-independent) context once and reuse it:
context = build_context_from_biotite(struct0, dev, prepare_ligands=True) for struct in structures: pose_stack = pose_stack_from_biotite(struct, dev, context=context)
Reusing a context skips rebuilding the parameter database, canonical ordering, residue-type set, and packed block types; only the per-structure canonical form is recomputed (see the
contextarg).Missing non-polymer atoms use prepared conformer geometry and resolved coordinate anchors. If only an attachment’s endpoints are resolved, its first declared torsion sample and the next resolved partner atom can orient the missing component. These are starting conformers for scoring/packing, not recovered experimental coordinates. Supplied heavy-atom coordinates stay unchanged; insufficient or degenerate references still raise.
- Parameters:
biotite_structure – A Biotite AtomArray or AtomArrayStack.
torch_device – Target PyTorch device.
param_db – Optional ParameterDatabase. When provided, conversion and pose construction use this database. If prepare_ligands=True, it is extended with ligand data. Mutually exclusive with
context.missing_density_distance_threshold – Distance threshold in Angstroms. Adjacent polymer residues whose connection atoms exceed this distance are disconnected only where topology was inferred. A supplied bond across a gap raises; set to 0 to retain it. Default is 2.4.
no_optH – Residues the input gives no hydrogens take AtomWorks’ protonation state: database residues then get hydrogens built by tmol, other residues those AtomWorks places. When True (default), preserve finite input hydrogen coordinates, except coincident hydrogen-parent pairs, which are rebuilt with a warning, and build only missing hydrogens and heavy-atom sidechains. When False, residues with complete heavy atoms are packed with OptHSampler to optimize hydrogen positions and NHQ flips, while residues with missing heavy atoms are rebuilt with DunbrackChiSampler. Generated ligand types may still rebuild hydrogens whose names changed during parameter generation; pass
trust_hydrogen_names=Trueonly when those names are known to match the prepared database.prepare_ligands – If True, detect and prepare non-standard residues (see
build_context_from_biotitefor details).ligand_ph – pH AtomWorks protonates residues lacking hydrogens at (default 7.4).
strict_atom_types – If True, unknown ligand atom types raise errors instead of using a fallback element heuristic.
strict_ligands – If True (default), raise when a detected ligand cannot be prepared and registered, instead of silently dropping it. Pass False to warn-and-skip. Only used when prepare_ligands=True.
ligand_params_files – Optional list of tmol YAML params file paths.
chem_comp_types –
{comp_id: type}from the input file’s_chem_comptable, which says whether a residue belongs to a polymer where the file does not number it along a sequence. Only used when prepare_ligands=True.ligand_seed – Fixed RNG seed for the conformer each prepared residue is built from, making preparation reproducible. Only used when prepare_ligands=True.
packer_seed – Seed of the packer that places hydrogens and builds missing side chains; unseeded, it continues torch’s global random state.
return_context – If True, return
(pose_stack, PoseBuildContext).context – Reusable context from
build_context_from_biotite. It must be ontorch_deviceand is mutually exclusive withparam_dbandprepare_ligands=True.atom37_coords – Optional coordinates shaped
[pose, token, 37, xyz]. When supplied, mapped coordinates are read from this tensor using the input structure’s integertoken_idandatom37_slotannotations. Wholly finite mapped triplets are authoritative; all-NaN mapped triplets are missing and do not fall back to the Biotite coordinates. Unmapped finite Biotite atoms remain context, and absent leaf atoms are completed normally. Partial-NaN or infinite mapped triplets raiseAtom37MappingError. The resulting pose coordinates remain connected to this tensor for autograd. Geometry-based missing-density and additional-disulfide detection are disabled so topology is fixed.**kwargs – Additional arguments passed to pose_stack_from_canonical_form.
- Returns:
PoseStack when no optional values requested and return_context is False.
(PoseStack, PoseBuildContext)when return_context is True.(PoseStack, dict)when optional return values were requested via kwargs. Fragmented poses expose their block mapping aspose_stack.split_block_mapping.
- tmol.io.pose_stack_from_cif(cif_path, device, *, model: int = 1, assembly_id: str | None = None, residue_start: int | None = None, residue_end: int | None = None, **kwargs)[source]#
Construct a PoseStack from a PDB, CIF or binary/compressed CIF structure.
Reads the structure with
atom_array_from_cif(), so the residues carry every atom their chemistry declares and the unresolved ones arrive at NaN for pose construction to rebuild. Further keyword arguments are passed totmol.io.pose_stack_from_biotite().
- tmol.io.atom_array_from_cif(cif_path, *, model: int = 1, assembly_id: str | None = None)[source]#
Read a PDB, CIF, compressed CIF, binary CIF, MOL2 or SDF into an AtomArray.
Preserve supplied names, coordinates, hydrogens and covalent connections. Declared unresolved heavy atoms have NaN coordinates. Chemistry the file leaves out is supplemented from the component dictionary.
Missing chemical metadata is allowed here. Pose construction uses existing parameters for known residues and requires additional chemistry only when preparing an unknown residue. Pose construction controls hydrogen optimization and whether to trust regenerated ligand hydrogen names.
- tmol.io.component_chemistry_from_cif(cif_path) dict[source]#
The complete chemistry a CIF declares for each component it names.
A structure records only the atoms it resolved, but it declares the whole component in
chem_comp_atomandchem_comp_bond– names, elements and bond orders, independent of what the density showed. That declaration is what residue-type generation needs, since a residue built from resolved atoms alone is a different molecule: an unresolved sidechain tip becomes a real methyl once the pipeline protonates it.Returns
{comp_id: AtomArray}with no coordinates – the pipeline generates its own conformer – and only for components the file bonds: connectivity cannot be guessed, so a multi-atom component with no authored bonds is left out for the caller to resolve another way.
- tmol.io.pose_stack_from_pdb(pdb_lines_or_fname: str | list | Path, device: device, *, residue_start: int | None = None, residue_end: int | None = None, res_not_connected: Tensor[slice(None, None, None), slice(None, None, None), 2] | None = None, **kwargs) PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#
Read PDB through AtomWorks and the shared annotated-array constructor.
Accept a path, PDB text or a list of lines. Residue slicing uses a half-open index range. Additional keywords follow
pose_stack_from_file(), includingprepare_ligandsandreturn_context. Supplied hydrogens are retained and hydrogen optimization is disabled unless explicitly requested.
- tmol.io.pose_stack_to_pdb_string(pose_stack: Any) str[source]#
Convert a
PoseStackinto PDB text suitable for molecular viewers.
- tmol.io.selection_gallery(atom_array: Any, selections: Mapping[str, object], *, notes: Mapping[str, str] | None = None, width: int = 720, height: int = 420, highlight_color: str = '#e83e8c')[source]#
Return one interactive viewer for several labeled AtomArray selections.
Selection values may be boolean atom masks. Query strings are also accepted when the supplied AtomArray provides a callable
aa.mask(query)method. Selection results are resolved in Python and exact PDB atom serials are baked into the HTML. This avoids viewer-side query-language differences and remains exact when atom names or residue identifiers are duplicated. Clicking a label restyles the same model and animates the camera to the selected atoms, following the AtomWorks selection-gallery interaction.- Parameters:
atom_array – Structure displayed by the shared viewer.
selections – Display labels mapped to boolean atom masks or query strings.
notes – Optional explanatory text keyed by selection label.
width – Viewer width in pixels.
height – Viewer height in pixels.
highlight_color – Color used for the active selection.
- tmol.io.switchable_view(structures: Mapping[str, Any], *, notes: Mapping[str, str] | None = None, width: int = 720, height: int = 420)[source]#
Return HTML that switches one 3Dmol viewer among labeled structures.
- Parameters:
structures – Ordered mapping of display labels to structures accepted by
view().notes – Optional mapping of structure labels to short explanatory text.
width – Viewer width in pixels.
height – Viewer height in pixels.
- tmol.io.view(model: Any, *, width: int = 720, height: int = 420, style: Literal['cartoon', 'stick'] = 'cartoon', background_color: str = 'white', cartoon_color: str = 'spectrum', show_sidechains: bool = True, show_heteroatoms: bool = True, show_hover: bool = True, zoom_to: dict | None = None, highlighted: object | None = None, highlight_color: str = '#e83e8c')[source]#
Create a draggable py3Dmol viewer for a molecular structure.
modelmay be aPoseStack, a BiotiteAtomArrayorAtomArrayStack, PDB text, or a PDB path.highlightedis an optional boolean mask over the atoms in the first model; highlighted atoms are shown as thicker sticks and spheres. The return value remains a realpy3Dmol.viewobject for compatibility with existing notebooks.
- tmol.io.write_pose_stack_pdb(pose_stack: PoseStack, fname_out: str, merge_fragments: bool = True, **kwargs)[source]#
Write a PDB-formatted file to disk given an input PoseStack. Optionally, additional arguments may be passed to the inner function “atom_records_from_pose_stack.” Fragmented ligands use their original residue identity by default; pass
merge_fragments=Falseto keep fragment residues separate.