Input and Output#

The public tmol.io API converts common structure representations to and from tmol.pose.PoseStack.

Once the residue types are resolved, tmol.io.pose_stack_from_atom37() builds a pose from tensors alone, for canonical and noncanonical chemistry alike: residue identity comes from res_types and any further covalent chemistry from covalent_bonds. Its slot layout is built by tmol.io.atom37_slot_map_for_ordering(). The direct AtomWorks adapter is the canonical-amino-acid special case of that path. To read chemistry from a structure instead, tmol.io.pose_stack_from_atom37_and_topology() derives the same information from general Biotite topology, including nucleic acids and ligands. Repeated diffusion, guidance, and search workloads should bind their fixed topology once with tmol.io.prepare_atom37_pose_builder(); its returned callable accepts each coordinate batch and preserves ordinary finite authored protein hydrogens by default. Generated ligand types may rebuild hydrogens unless trust_hydrogen_names=True.

Use tmol.io.atom_array_from_file() or tmol.io.pose_stack_from_file() for PDB, CIF, compressed CIF and binary CIF. The reader options are model and assembly_id. Reading preserves supplied names, coordinates, hydrogens and covalent bonds; declared unresolved heavy atoms carry NaN coordinates. Chemistry the file leaves out is completed from the component dictionary. Missing chemistry is allowed at this stage.

Known residues need only names and coordinates. Pass an existing param_db or load prepared ligand_params_files to reuse ligand parameters with a coordinate-only PDB/CIF. With prepare_ligands=True, pose preparation generates only missing parameters and reports the residue name when required bond orders are absent. Use return_context=True to retain the resulting parameter database for scoring and repeated construction. Direct AtomArray inputs follow the same contract.

# Reuse known chemistry; the PDB needs only matching names and coordinates.
pose = pose_stack_from_file(
    "complex.pdb", device, ligand_params_files=["ligand.tmol"],
)
# Prepare missing chemistry once, then reuse context.parameter_database.
pose, context = pose_stack_from_file(
    "complex.cif", device, prepare_ligands=True, return_context=True,
)

tmol.io.pose_stack_from_pdb() accepts paths, text and lists of lines through the same file path. File and AtomArray inputs preserve matching supplied coordinates for ordinary protein residues. Pose construction builds missing hydrogens but preserves finite authored protein hydrogen coordinates by default. Pass no_optH=False to request OptH packing. Generated ligand types may rebuild hydrogens whose names changed during parameter generation; pass trust_hydrogen_names=True only when those names match the prepared database. Low-level PDB atom-record/DataFrame utilities retain their canonical-only contract.

Structure conversion between external formats and TMol poses.

tmol.io.add_metal_coordination(canonical_ordering: CanonicalOrdering, pose_stack: PoseStack, pose: int, metal: int, donor: int, atom: str, site: int | None = None) → PoseStack[source]#

Bond atom of donor to an open site of metal (by default the one whose virtual faces it best); the donor atom loses its hydrogens.

tmol.io.assemble_input(*parts: AtomArray) → AtomArray[source]#

Join separately-read parts with their bonds, annotations and templates; a chain ID claimed by an earlier part is moved to a free one.

Parameters:

*parts – Structures to join, in the order they should appear.

Returns:

The joined structure.

Raises:

ValueError – If no parts are given, or chain IDs cannot be made unique.

Examples

>>> complex = assemble_input(protein, ligand)
>>> sorted(set(complex.chain_id))
['A', 'L']
tmol.io.atom_array_from_file(cif_path, *, model: int = 1, assembly_id: str | None = None)#

Read a PDB, CIF, compressed CIF, binary CIF, MOL2 or SDF into an AtomArray.

Preserve supplied names, coordinates, hydrogens and covalent connections. Declared unresolved heavy atoms have NaN coordinates. Chemistry the file leaves out is supplemented from the component dictionary.

Missing chemical metadata is allowed here. Pose construction uses existing parameters for known residues and requires additional chemistry only when preparing an unknown residue. Pose construction controls hydrogen optimization and whether to trust regenerated ligand hydrogen names.

tmol.io.atom_array_from_mol2(mol2_path: str | Path, *, res_name: str | None = None, chain_id: str = 'L', res_id: int = 1) → AtomArray[source]#

Read a Tripos MOL2 (or MDL SDF/MOL) ligand as an AtomArray with the file’s bond orders, formal charges and aromatic flags, registered as its own component template.

Parameters:
  • mol2_path – Path to the MOL2, SDF or MOL file.

  • res_name – Component code to give the ligand. Taken from the file’s substructure record when None.

  • chain_id – Chain to place the ligand on.

  • res_id – Residue number within that chain.

Returns:

A single-residue AtomArray with a populated bond list.

Examples

>>> ligand = atom_array_from_mol2("ligand.mol2")
>>> ligand.bonds.get_bond_count() > 0
True
tmol.io.cif_from_atom_array(atom_array: AtomArray, *, path: str | Path | None = None, entry_id: str = 'assembled') → str | Path[source]#

Write a structure as a CIF carrying its own component chemistry: known components from the dictionary, others (a MOL2 ligand) from the structure itself.

Parameters:
  • atom_array – Structure to write.

  • path – Where to write. Returns the CIF text instead when None.

  • entry_id – Data block name.

Returns:

The written path, or the CIF text when path is None.

Examples

>>> text = cif_from_atom_array(assemble_input(protein, ligand))
>>> "_chem_comp_bond" in text
True
exception tmol.io.Atom37MappingError[source]#

Bases: ValueError

An AtomArray cannot be routed unambiguously into an Atom37 tensor.

class tmol.io.CanonicalForm(chain_id: Tensor[slice(None, None, None), slice(None, None, None)], res_types: Tensor[slice(None, None, None), slice(None, None, None)], coords: Tensor[slice(None, None, None), slice(None, None, None), slice(None, None, None), 3], res_labels: NDArray[slice(None, None, None), slice(None, None, None)], residue_insertion_codes: NDArray[slice(None, None, None), slice(None, None, None)], chain_labels: NDArray[slice(None, None, None), slice(None, None, None)], atom_occupancy: NDArray[slice(None, None, None), slice(None, None, None), slice(None, None, None)] | None, atom_b_factor: NDArray[slice(None, None, None), slice(None, None, None), slice(None, None, None)] | None, disulfides: Tensor[slice(None, None, None), 3] | None, res_not_connected: Tensor[slice(None, None, None), slice(None, None, None), 2] | None, cyclic_bonds: Tensor[slice(None, None, None), 3] | None = None, covalent_bonds: Tensor[slice(None, None, None), 5] | None = None, metal_sites: Tensor[slice(None, None, None), 3] | None = None, metal_coordination: Tensor[slice(None, None, None), 5] | None = None, metal_origins: NDArray[slice(None, None, None), slice(None, None, None)] | None = None, residue_annotations: ndarray | None = None, protonation_variants: Tensor[slice(None, None, None), slice(None, None, None)] | None = None)[source]#

Bases: object

This class holds the data that describe a (stack of) structure(s) in a poised, ready-to-use state.

This datastructure holds the information necessary to determine the chemical identities of the residues in the structure(s), which may be under-determined from tmol’s perspective by the source of the structure (e.g. OpenFold does not explicitly model termini). The atoms that are present are represented with non-NaN coordinates in the coords array; the order in which those atoms appear is given by a particular CanonicalOrdering object.

The datastructure also holds convenience information such as author-provided residue labels (ints), chain labels (strings) & insertion codes (strings) as well as the occupancy and B-factor of each atom. These are not strictly necessary but are often useful when processing structures.

as_dict()[source]#

Constructor keyword arguments sharing the original tensors and arrays.

class tmol.io.CanonicalOrdering(max_n_canonical_atoms: int, restype_io_equiv_classes: Tuple[str, ...], restypes_ordered_atom_names: Mapping[str, Tuple[str, ...]], restypes_atom_elements: Mapping[str, Mapping[str, str]], restypes_atom_index_mapping: Mapping[str, Mapping[str, int]], restypes_mainchain_atoms: Mapping[str, Tuple[str, ...] | None], restypes_required_mainchain_atoms: Mapping[str, Tuple[str, ...] | None], restypes_default_termini_mapping: Mapping[str, Tuple[str | None, str | None]], down_termini_patches: Tuple[str, ...], up_termini_patches: Tuple[str, ...], termini_patch_added_atoms: Mapping[Tuple[str, str], Tuple[str, ...]], termini_only_atoms: Mapping[Tuple[str, str], Tuple[str, ...]], termini_only_atoms_by_class: Mapping[str, Tuple[str, ...]], cys_inds: CysSpecialCaseIndices, his_inds: HisSpecialCaseIndices, polymer_conn_inds: PolymerConnectionIndices, name3_aliases: Mapping[str, str] = NOTHING, restypes_virtual_atoms: Mapping[str, FrozenSet[str]] = NOTHING)[source]#

Bases: object

The canonical ordering class describes the integer ordering of residue types and for atoms within those residue types for the collection of available residue types defined by a PatchedChemicalDatabase.

The canonical ordering class’s purpose is to enable creation of a “canonical form” dictionary that describes a molecular system in the way that tmol expects in order to construct a PoseStack.

There is no “canonical form” dictionary is simply a dictionary holding the at-least-three-but-as-many-as-eight arguments to tmol.io._pose_stack_construction.pose_stack_from_canonical_form after the first two. That is, it must contain “chain_id”, “res_types” and “coords” entries.

When constructing a PoseStack, there are multiple residue types for each “equivalence class” (think 3-letter code); e.g. for “CYS” there’s the standard middle-of-a-polypeptide-chain CYS, the standard middle-of-a-polypeptide-chain disulfide-forming CYS, and then for those two, four variants for the N-, C-, and both-N- and-C terminal forms; eight total options for a single “CYS” three-letter code. tmol collects all of the various forms of a single equivalence class and creates a list of all atom names across all the residue types for it. You can then provide tmol the set of atoms that are present at a given position by giving a non-NaN coordinate for that entry in an [n-poses x max-n-res x max-ats-per-res x 3] tensor of coordinates. Atoms with NaN coordinates are taken as possibly present in the residue type; tmol will decide the best fit for which residue type to use at each position. If an atom is provided to tmol and it is not present for a given residue type, then that residue type will be disqualified from consideration. Thus an important part of telling tmol which atoms are present is mapping from an atom name to an index for that atom. The CanonicalOrdering object is where that mapping is encoded. It also handles the mapping from alternate-atom-name to canonical-form-atom index; e.g. in PDBv2, glycine’s two hydrogens were named “HA1” and “HA2”, but in PDBv3, they are named “1HA” and “2HA.” So that we can parse PDB files written in PDBv2 and PDBv3, we have an idea of an “alias” for an atom; see the restypes_atom_index_mapping data member.

Four data members are especially useful:

  • max_n_canonical_atoms

  • restype_io_equiv_classes

  • restypes_ordered_atom_names

  • restypes_atom_index_mapping

The remaining data members are primarily for internal tmol functionality.

max_n_canonical_atoms: the largest number of distinct atom names among all

variants of a single residue type (equivalence class) across all residue types

restype_io_equiv_classes:

essentially the list of 3-letter codes for the residue types that are readable; use the index function (e.g. co.restype_io_equiv_classes.index(“TRP”)) to obtain the integer meant to represent each restype

restypes_ordered_atom_names:

the ordered list of the names of each atom for every allowed residue type; does not include the alternate names for atoms. Atoms should be given to tmol in this order; e.g. by putting the coordinate of the ith atom in the ith entry of the coordinate tensor (e.g. coords[p, r, i] for pose p, residue r)

restypes_atom_index_mapping:

mapping for each name3 from atom name and atom name alias to the index of that atom for every allowed residue type in the restypes_ordered_atom_names list; this is probably more useful than the restypes_ordered_atom_names list, especially if you are using the PDBv2 naming convention (as Rosetta3 does) instead of the PDBv3 convention.

resolve_name3(name3: str) → str[source]#

The residue name to read this input name as.

class tmol.io.PoseBuildContext(canonical_ordering: CanonicalOrdering, packed_block_types: PackedBlockTypes, parameter_database: ParameterDatabase, restype_set: ResidueTypeSet, fragment_definitions: tuple[LigandFragmentDefinition, ...]=(), ligand_names: dict[str, str]=<factory>, cut_covalent_partners: frozenset[str] = frozenset({}))[source]#

Bases: object

Immutable, structure-independent construction context.

Holds only the pieces that depend on the parameter database / ligand set (not on any particular input), so it can be built once and reused across many inputs that share the same ligand(s).

class tmol.io.PreparedAtom37PoseBuilder(context: PoseBuildContext, canonical_template: CanonicalForm, mapped_token_id: Tensor, mapped_slot: Tensor, mapped_residue: Tensor, mapped_atom: Tensor, max_token_id: int, required_mainchain_entries: tuple[tuple[int, int, str], ...], fragment_mapping: FragmentedLigandPoseMapping | None = None, topology_cache_safe: bool = True, pose_topologies: dict[int, _PreparedAtom37PoseTopology] = NOTHING)[source]#

Bases: object

Bind immutable Biotite topology for repeated Atom37 pose construction.

Calls accept float32 coordinates shaped [n_poses, n_tokens, 37, 3] on the context’s device. The first call for a batch size prepares its fixed pose topology; up to four recently used batch sizes are cached. Inputs with coordinate-dependent atom presence or ambiguous histidine hydrogens use the normal uncached construction path.

The builder owns a mutable topology cache and is not safe for concurrent calls. Use one builder per calling thread when pose construction overlaps.

tmol.io.atom_records_from_coords(pbt: PackedBlockTypes, chain_ind_for_block: Tensor[slice(None, None, None), slice(None, None, None)], block_types64: Tensor[slice(None, None, None), slice(None, None, None)], pose_like_coords: Tensor[slice(None, None, None), slice(None, None, None), 3], block_coord_offset: Tensor[slice(None, None, None), slice(None, None, None)], residue_labels: NDArray[slice(None, None, None), slice(None, None, None)] | None, residue_insertion_codes: NDArray[slice(None, None, None), slice(None, None, None)] | None, chain_labels: NDArray[slice(None, None, None), slice(None, None, None)] | None, atom_occupancy: NDArray[slice(None, None, None), slice(None, None, None)], atom_b_factor: NDArray[slice(None, None, None), slice(None, None, None)]) → NDArray[source]#

Create a numpy array holding the atom records needed to write a PDB file from the coordinates and block types of a stack of structures, laid out in pose-stack form.

tmol.io.atom_records_from_pose_stack(pose_stack: PoseStack, merge_fragments: bool = True) → NDArray[source]#

Create a numpy array holding the atom records needed to write a PDB file from a PoseStack.

Fragmented ligands use their original residue identity by default. Pass merge_fragments=False to retain the separate fragment residue numbers.

tmol.io.atomworks_from_pose_stack(pose_stack: PoseStack) → tuple[source]#

Convert a PoseStack back to atomworks UNIFIED_ATOM37_ENCODING tensors.

Parameters:

pose_stack (PoseStack) – The PoseStack to convert. Must contain only standard amino acids.

Returns:

  • coords (Tensor, shape [n_poses, max_n_res, 37, 3]) – Atom coordinates in the atomworks atom37 layout. Absent atoms are 0.

  • residue_type (Tensor[int64], shape [n_poses, max_n_res]) – Atomworks token indices (1..20 for real residues, 0 for padding).

  • chain_iid (Tensor[int64], shape [n_poses, max_n_res]) – Chain identifiers.

tmol.io.biotite_from_canonical_form(cf: CanonicalForm, co: CanonicalOrdering | None = None, include_virtual_atoms: bool = False) → AtomArray | AtomArrayStack[source]#

Export coordinates and chemical/author labels using a shared atom layout.

Multi-model arrays require identical residue and atom annotations. Their atom layout is the union of resolved atoms; absent coordinates remain NaN. Missing author labels default to internal chain IDs and one-based residue positions. Virtual atoms are left out unless include_virtual_atoms. This host-array export detaches coordinates from autograd. A canonical form carries no residue types, so the result has no bonds; biotite_from_pose_stack adds them.

tmol.io.biotite_from_pose_stack(pose_stack: PoseStack, co: CanonicalOrdering | None = None, merge_fragments: bool = True, include_virtual_atoms: bool = False) → AtomArray | AtomArrayStack[source]#

Convert PoseStack back to Biotite structure.

Parameters:
  • pose_stack – Pose stack to convert.

  • co – Canonical ordering used for conversion. Provide the ordering that was used when ligands or custom residue types are present.

  • merge_fragments – Restore fragmented ligands to their original residue identity. Set to False to keep fragment residues separate.

  • include_virtual_atoms – Also write virtual atoms, such as a metal’s site virtuals. A metal split out of a component is then left as its own residue, which its virtuals belong to; otherwise it is put back in the component it came from.

Returns:

Biotite AtomArray for single-pose or AtomArrayStack for multi-pose, with every bond the poses’ residue types and connections declare: bond orders within residues, and chemical bonds between them, metal coordination typed COORDINATION.

tmol.io.build_context_from_biotite(biotite_structure: AtomArray | AtomArrayStack, torch_device: device, param_db: ParameterDatabase | None = None, prepare_ligands: bool = False, ligand_ph: float = 7.4, strict_atom_types: bool = False, strict_ligands: bool = True, ligand_params_files: list[str] | None = None, chem_comp_types: dict | None = None, ligand_seed: int | None = None) → PoseBuildContext[source]#

Build the structure-independent construction context.

The returned context holds only database/ligand-derived pieces (canonical ordering, residue-type set, packed block types, parameter database); it does not depend on the input structure’s coordinates and can be reused across structures sharing the same ligand(s). biotite_structure is used only to detect and prepare ligands (when prepare_ligands=True).

Parameters:
  • biotite_structure – Input AtomArray or AtomArrayStack. Used only for ligand detection/preparation when prepare_ligands=True.

  • torch_device – Target torch device.

  • param_db – Optional parameter database. When provided, canonical ordering, residue types, and packed block types are built from this database. If prepare_ligands=True, it is extended with ligand data. If None, defaults are used.

  • prepare_ligands – If True, detect and prepare non-standard residues (via tmol.ligand, which uses RDKit for atom typing and residue-type construction).

  • ligand_ph – pH AtomWorks protonates residues lacking hydrogens at before ligands are prepared (default 7.4, only used when prepare_ligands=True).

  • strict_atom_types – If True, unknown ligand atom types raise errors instead of using a fallback element heuristic.

  • strict_ligands – If True (default), raise when a detected ligand cannot be prepared and registered (instead of silently dropping it during pose construction). Pass False to fall back to warn-and-skip. Only used when prepare_ligands=True.

  • ligand_params_files – Optional list of tmol YAML params file paths. Residues defined in these files skip the RDKit/OB pipeline.

  • chem_comp_types – {comp_id: type} from the input file’s _chem_comp table (see tmol.ligand.chem_comp_types_from_cif), which says whether a residue belongs to a polymer where the file does not number it along a sequence. Only used when prepare_ligands=True.

  • ligand_seed – Fixed RNG seed for the conformer each prepared residue is built from, making preparation reproducible. Only used when prepare_ligands=True.

Returns:

PoseBuildContext containing canonical ordering, packed block types, parameter database, and residue type set.

tmol.io.canonical_form_from_atomworks(coords: Tensor, residue_type: Tensor, chain_iid: Tensor) → CanonicalForm[source]#

Build a CanonicalForm from atomworks UNIFIED_ATOM37_ENCODING tensors.

Parameters:
  • coords (Tensor, shape [batch, n_res, 37, 3]) – Atom coordinates in the atomworks atom37 layout.

  • residue_type (Tensor[int64], shape [batch, n_res]) – Atomworks token indices. Must be in 1..20 (standard protein only).

  • chain_iid (Tensor[int64], shape [batch, n_res]) – Chain identifiers.

Return type:

CanonicalForm

tmol.io.canonical_form_from_biotite(biotite_structure: AtomArray | AtomArrayStack, torch_device: device, co: CanonicalOrdering | None = None, missing_density_distance_threshold: float = 2.4, atom37_coords: Tensor | None = None, _filter_missing_mainchain: bool = True, _cut_covalent_partners: frozenset[str] = frozenset({})) → CanonicalForm[source]#

Convert a Biotite AtomArray or AtomArrayStack to a CanonicalForm.

This function bridges between Biotite’s data structures and tmol’s internal representation by converting atom and residue information from string-based identifiers to tmol’s canonical integer-based indexing system.

Parameters:
  • biotite_structure – A Biotite AtomArray (single structure) or AtomArrayStack (multiple structures) containing the molecular data. Must contain atom coordinates, residue names, atom names, chain IDs, and optionally B-factors and occupancy values.

  • torch_device – PyTorch device (e.g., torch.device(‘cuda’) or torch.device(‘cpu’)) where the resulting tensors should be allocated.

  • co – A CanonicalForm in case you want to use a non-default database (and thus may need a different mapping)

  • missing_density_distance_threshold – Maximum distance in angstroms for treating a polymer gap as missing density rather than a chain break.

  • atom37_coords – Optional autograd-tracked coordinate tensor of shape [n_poses, n_tokens, 37, 3]. When provided, coordinates are sourced from this tensor (routed by the token_id and atom37_slot annotations on biotite_structure); all-NaN mapped triplets are treated as missing while unmapped finite reference atoms remain context. The geometry-based missing-density check is skipped so the topology stays fixed and gradients flow. See pose_stack_from_atom37_and_topology().

Returns:

A data structure containing:
  • chain_id: Tensor mapping residues to chain indices

  • res_types: Tensor mapping residues to tmol residue type indices

  • coords: 4D tensor of atomic coordinates (poses x residues x atoms x 3)

  • res_labels: Original residue sequence numbers from the structure

  • residue_insertion_codes: PDB insertion codes for residues

  • chain_labels: Original chain identifiers from the structure

  • atom_occupancy: Optional tensor of atom occupancy values

  • atom_b_factor: Optional tensor of atom B-factor values

  • disulfides: Explicit cysteine sulfur bonds, including unresolved SG

  • res_not_connected: Tensor describing whether two consecutive residues should be treated as chemically bonded.

Return type:

CanonicalForm

tmol.io.canonical_form_from_pdb(canonical_ordering: CanonicalOrdering, pdb_lines_or_fname: str | List, device: device, *, residue_start: int | None = None, residue_end: int | None = None, res_not_connected: Tensor[slice(None, None, None), slice(None, None, None), 2] | None = None) → CanonicalForm[source]#

Create a canonical form from either the contents of a PDB file as one long string or a list of individual lines from the file or by providing the name/path of a PDB file

pdb_lines_or_fname must either be a list of the lines in a PDB file or a string representing a file

tmol.io.canonical_form_from_pose_stack(canonical_ordering: CanonicalOrdering, pose_stack: PoseStack, chain_id=None)[source]#

Convert a pose stack to canonical residue and atom tensors.

Parameters:
  • canonical_ordering – Residue and atom ordering for the output tensors.

  • pose_stack – Poses to deconstruct.

  • chain_id – Optional integer chain identifiers shaped [pose, residue].

Returns:

Canonical form containing coordinates, residue metadata, and connectivity.

tmol.io.canonical_ordering_for_atomworks() → CanonicalOrdering#

Construct the CanonicalOrdering for the protein subset used by the atomworks UNIFIED_ATOM37_ENCODING.

tmol.io.canonical_ordering_for_biotite() → CanonicalOrdering#

Construct the CanonicalOrdering object to use for Biotite. This wont be used as a typical CanonicalOrdering object, since we aren’t mapping from int-to-int, and instead are going from string-to-int.

tmol.io.create_pose_stack_from_sequences(seqs, packed_block_types: PackedBlockTypes | None = None, device: device | None = None, param_db=None, termini: bool = True, context: PoseBuildContext | None = None, return_context: bool = False)[source]#

Construct a PoseStack with zero coordinates from sequence strings.

See tmol.pose._sequence for the grammar. Returns (PoseStack, PoseBuildContext) when return_context is set; the context carries the database extended with any ligands the sequence names.

tmol.io.default_canonical_ordering() → CanonicalOrdering#

Create a CanonicalOrdering object from the default set of residue types

tmol.io.default_packed_block_types(device: device) → PackedBlockTypes[source]#

Create a PackedBlockTypes object from the default set of residue types

tmol.io.extended_pose_stack_from_sequences(seqs, device: device | None = None, param_db=None, termini: bool = True, context=None, return_context: bool = False)[source]#

Build a PoseStack from sequences with ideal geometry and extended backbone torsions.

See tmol.pose._sequence for the grammar. Returns (PoseStack, PoseBuildContext) when return_context is set.

tmol.io.fetch_pdb(pdbid)[source]#

Download a PDB-format structure from the RCSB Protein Data Bank.

tmol.io.packed_block_types_for_atomworks(device: device) → PackedBlockTypes#

Construct the PackedBlockTypes for the protein subset used by the atomworks UNIFIED_ATOM37_ENCODING.

tmol.io.packed_block_types_for_biotite(device: device) → PackedBlockTypes#

Construct the PackedBlockTypes (PBT) object that will used for Biotite. We’ll use the defaults since anything might show up in a Biotite AtomArray. Some things may show up in the AtomArrays that are not handled by this PBT, but that is work for the future.

tmol.io.atom37_slot_map_for_ordering(canonical_ordering: CanonicalOrdering, atom_names_by_slot: Mapping[str, Sequence[str]], device: device) → Tensor[source]#

Map each residue type’s Atom37 slots to canonical atom indices.

The mapping is a property of the residue-type table and the slot layout, not of any structure, so it is built once per (ordering, layout) pair.

Atoms that a terminus patch adds are deliberately left unmapped. They exist only on terminal residue variants, and offering one on an interior residue would disqualify that residue’s block types; tmol applies termini itself from the chain boundaries. Which atoms those are is read from the chemical database rather than named here.

Parameters:
  • canonical_ordering – Residue and atom ordering the res_types index into.

  • atom_names_by_slot – Atom name occupying each slot, per residue name3. An empty name marks an unused slot.

  • device – Device for the returned tensor.

Returns:

[n_residue_types, n_slots] canonical atom index per slot, -1 where the slot is unused or the atom is not part of that residue type.

Examples

>>> slot_map = atom37_slot_map_for_ordering(co, layout, torch.device("cpu"))
>>> slot_map.shape[1]
37
tmol.io.canonical_form_from_atom37(atom37_coords: Tensor, res_types: Tensor, chain_id: Tensor, canonical_ordering: CanonicalOrdering, *, slot_map: Tensor | None = None, disulfides: Tensor | None = None, cyclic_bonds: Tensor | None = None, covalent_bonds: Tensor | None = None, res_not_connected: Tensor | None = None) → CanonicalForm[source]#

Scatter Atom37 coordinates into a canonical form.

Parameters:
  • atom37_coords – [n_poses, n_tokens, n_slots, 3], autograd-tracked.

  • res_types – [n_poses, n_tokens] index into canonical_ordering; -1 marks padding.

  • chain_id – [n_poses, n_tokens] chain index per token.

  • canonical_ordering – Ordering the res_types index into.

  • slot_map – [n_residue_types, n_slots] from atom37_slot_map_for_ordering(). Defaults to the standard Atom37 layout, which covers the canonical amino acids; supply one built from an extended layout for other chemistry.

  • disulfides – [n, 3] explicit (pose, res1, res2) rows.

  • cyclic_bonds – [n, 3] explicit head-to-tail closures.

  • covalent_bonds – [n, 5] explicit (pose, res1, atom1, res2, atom2) links in canonical ordering.

  • res_not_connected – [n_poses, n_tokens, 2] suppressed polymer connections.

Returns:

A canonical form whose coordinates carry gradients back to atom37_coords.

tmol.io.pose_stack_from_atom37(atom37_coords: Tensor, res_types: Tensor, chain_id: Tensor, context: PoseBuildContext, *, slot_map: Tensor | None = None, disulfides: Tensor | None = None, cyclic_bonds: Tensor | None = None, covalent_bonds: Tensor | None = None, res_not_connected: Tensor | None = None, no_optH: bool = True, **kwargs: object) → PoseStack[source]#

Build a differentiable PoseStack from Atom37 tensors alone.

One path for canonical and noncanonical poses. Residue identity comes from res_types indexed into the context’s canonical ordering, and any covalent chemistry the polymer connections do not already describe comes from covalent_bonds; no AtomArray is consulted.

Parameters:
  • atom37_coords – [n_poses, n_tokens, n_slots, 3], autograd-tracked.

  • res_types – [n_poses, n_tokens] index into context.canonical_ordering; -1 marks padding.

  • chain_id – [n_poses, n_tokens] chain index per token.

  • context – Chemistry resolved once, from any of the build_context_from_* helpers.

  • slot_map – Atom37 slot layout; see atom37_slot_map_for_ordering().

  • disulfides – Explicit (pose, res1, res2) rows.

  • cyclic_bonds – Explicit head-to-tail closures.

  • covalent_bonds – Explicit (pose, res1, atom1, res2, atom2) links.

  • res_not_connected – Suppressed polymer connections.

  • no_optH – Preserve finite input hydrogens rather than optimizing them.

Returns:

A pose whose coordinates carry gradients back to atom37_coords.

Examples

>>> context = build_context_from_biotite(array, device)
>>> pose = pose_stack_from_atom37(coords, res_types, chain_id, context)
>>> pose.coords.requires_grad
True
tmol.io.pose_stack_from_canonical_aa_atom37(coords: Tensor, residue_type: Tensor, chain_iid: Tensor, **kwargs) → PoseStack[source]#

Build a PoseStack from atomworks UNIFIED_ATOM37_ENCODING tensors.

This is the tensor-only Atom37 entry point. Residue identity comes from residue_type alone, so nothing but tensors is required – but for the same reason it is limited to the canonical amino acids with the canonical n- and c-termini patches.

For anything else – noncanonical residues, ligands, nucleic acids, or covalent links – residue identity cannot be read off a 1..20 token, and pose_stack_from_atom37_and_topology() is the entry point; it takes the chemistry from a supplied topology instead.

Both build a CanonicalForm and finish in the same constructor, so the choice here is only about how residue identity is supplied, not about how the pose is built.

Parameters:
  • coords (Tensor, shape [batch, n_res, 37, 3]) – Atom coordinates in the atomworks atom37 layout.

  • residue_type (Tensor[int64], shape [batch, n_res]) – Atomworks token indices. Must be in 1..20 (standard protein only).

  • chain_iid (Tensor[int64], shape [batch, n_res]) – Chain identifiers (integer IDs, not string labels).

  • **kwargs – Additional arguments passed to pose_stack_from_canonical_form.

Return type:

PoseStack

Raises:

ValueError – If any residue_type value is outside 1..20 (protein-only).

tmol.io.pose_stack_from_file(cif_path, device, *, model: int = 1, assembly_id: str | None = None, residue_start: int | None = None, residue_end: int | None = None, **kwargs)#

Construct a PoseStack from a PDB, CIF or binary/compressed CIF structure.

Reads the structure with atom_array_from_cif(), so the residues carry every atom their chemistry declares and the unresolved ones arrive at NaN for pose construction to rebuild. Further keyword arguments are passed to tmol.io.pose_stack_from_biotite().

tmol.io.pose_stack_from_canonical_form_and_context(cf: CanonicalForm, context: PoseBuildContext, *, no_optH: bool, atom37_coords: Tensor | None, fragment_mapping=None, return_context: bool = False, packer_seed: int | None = None, **kwargs: object) → PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#

Build a pose from a canonical form and a reusable build context.

This is the single construction path. Every entry point – canonical amino-acid Atom37 tensors, an Atom37 batch bound to a fixed topology, a parsed CIF/PDB structure – reduces its input to a CanonicalForm plus a PoseBuildContext and finishes here. The entry points differ only in how they derive those two objects, not in how a pose is built from them.

The context carries the canonical ordering and packed block types, so noncanonical residues, ligands, and covalent links are handled by the same code as standard amino acids: nothing here inspects an AtomArray. Callers that already hold a canonical form and a context – notably repeated guidance or search steps over one fixed topology – should call this directly rather than re-deriving topology per batch.

Parameters:
  • cf – Canonical-form tensors describing residue identity and coordinates.

  • context – Structure-independent chemistry resolved once.

  • no_optH – Preserve finite input hydrogens instead of optimizing them.

  • atom37_coords – When supplied, the autograd-tracked source of cf’s coordinates; retained so gradients survive hydrogen rebuilding.

  • fragment_mapping – Mapping produced when fragmented ligands were expanded.

  • return_context – Also return the context used.

  • packer_seed – Seed of the packer, as for pose_stack_from_biotite.

For ordinary structure inputs, coincident bonded heavy atoms raise before packing; hydrogens coincident with their parent are rebuilt without changing protonation. Differentiable atom37_coords retain their supplied geometry.

Returns:

The constructed pose, optionally with atom mappings or the context.

tmol.io.prepare_atom37_pose_builder(biotite_structure: AtomArray | AtomArrayStack, context: PoseBuildContext) → PreparedAtom37PoseBuilder[source]#

Prepare a callable for repeatedly binding Atom37 coordinates to topology.

This is the campaign-oriented counterpart to pose_stack_from_atom37_and_topology(): immutable residue identity, connectivity, fragmentation, and Atom37 routing are resolved once. Calling the returned builder with a coordinate tensor constructs a differentiable pose while retaining TMol’s usual missing-atom behavior. The returned builder preserves finite input hydrogens by default; pass opt_h=True to optimize them after construction.

tmol.io.remove_metal_coordination(canonical_ordering: CanonicalOrdering, pose_stack: PoseStack, pose: int, metal: int, site: int) → PoseStack[source]#

Open one filled site of a metal; the donor keeps its hydrogens.

tmol.io.pose_stack_from_atom37_and_topology(atom37_coords: Tensor, biotite_structure: AtomArray | AtomArrayStack, context: PoseBuildContext, no_optH: bool = True, **kwargs) → PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#

Build a differentiable PoseStack from atom37 coordinates and a topology.

This is the topology-supplied Atom37 entry point. Unlike pose_stack_from_canonical_aa_atom37(), which reads residue identity from a 1..20 token tensor and so covers only canonical amino acids, this supports any chemistry shared by AtomWorks and the supplied TMol context (including ordinary ligands and nucleic acids): the chemical topology is taken from biotite_structure while mapped coordinates come from the autograd-tracked atom37_coords tensor.

The topology is only read to derive residue identity and the Atom37 slot map. Once derived, construction is pure tensor work in the shared pose_stack_from_canonical_form_and_context(). When the same topology is scored repeatedly, use prepare_atom37_pose_builder() to pay that derivation once; the returned builder never touches the AtomArray again. Unmapped finite reference atoms remain context. This is the entry point for differentiable scoring/guidance over atomized inputs, where the same fixed topology is scored repeatedly as coordinates move.

Build context once with build_context_from_biotite() (with prepare_ligands=True when ligands are present). For repeated diffusion or search steps, bind the topology once with prepare_atom37_pose_builder() and call the returned builder with each coordinate batch. Mapped reference coordinates are ignored, so one reference topology can be reused as those coordinates change. It must carry two integer annotations that map each atom into the atom37 tensor: token_id (the token axis) and atom37_slot (the 0..36 slot).

Topology is derived from chemical identity alone – missing_density breaks and automatic disulfide detection (both coordinate-dependent) are disabled – so the block types, termini, and atom count stay fixed as coordinates change. This adapter does not classify or allowlist residue types or elements: newly supported PTMs, ions, and metals work through the same API once they are represented by the supplied context and canonical ordering.

Parameters:
  • atom37_coords (Tensor, shape [n_poses, n_tokens, 37, 3]) – Autograd-tracked coordinates in the atomworks atom37 layout.

  • biotite_structure (biotite AtomArray) – Reference topology carrying token_id and atom37_slot annotations.

  • context (PoseBuildContext) – Structure-independent context from build_context_from_biotite().

  • no_optH (bool) – Preserve finite input hydrogens and leave newly built hydrogens at ideal positions when True (default). Pass False to run TMol’s hydrogen optimization pipeline.

  • **kwargs – Additional arguments forwarded to pose_stack_from_biotite.

Returns:

Whose coords carry gradients back to atom37_coords.

Return type:

PoseStack

tmol.io.pose_stack_from_biotite(biotite_structure: AtomArray | AtomArrayStack, torch_device: device, param_db: ParameterDatabase | None = None, missing_density_distance_threshold: float = 2.4, no_optH: bool = True, prepare_ligands: bool = False, ligand_ph: float = 7.4, strict_atom_types: bool = False, strict_ligands: bool = True, ligand_params_files: list[str] | None = None, chem_comp_types: dict | None = None, ligand_seed: int | None = None, packer_seed: int | None = None, return_context: bool = False, context: PoseBuildContext | None = None, atom37_coords: Tensor | None = None, **kwargs: object) → PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#

Build a PoseStack from the output generated by Biotite.

To score many structures that share the same ligand(s) efficiently, build the (expensive, structure-independent) context once and reuse it:

context = build_context_from_biotite(struct0, dev, prepare_ligands=True)
for struct in structures:
    pose_stack = pose_stack_from_biotite(struct, dev, context=context)

Reusing a context skips rebuilding the parameter database, canonical ordering, residue-type set, and packed block types; only the per-structure canonical form is recomputed (see the context arg).

Missing non-polymer atoms use prepared conformer geometry and resolved coordinate anchors. If only an attachment’s endpoints are resolved, its first declared torsion sample and the next resolved partner atom can orient the missing component. These are starting conformers for scoring/packing, not recovered experimental coordinates. Supplied heavy-atom coordinates stay unchanged; insufficient or degenerate references still raise.

Parameters:
  • biotite_structure – A Biotite AtomArray or AtomArrayStack.

  • torch_device – Target PyTorch device.

  • param_db – Optional ParameterDatabase. When provided, conversion and pose construction use this database. If prepare_ligands=True, it is extended with ligand data. Mutually exclusive with context.

  • missing_density_distance_threshold – Distance threshold in Angstroms. Adjacent polymer residues whose connection atoms exceed this distance are disconnected only where topology was inferred. A supplied bond across a gap raises; set to 0 to retain it. Default is 2.4.

  • no_optH – Residues the input gives no hydrogens take AtomWorks’ protonation state: database residues then get hydrogens built by tmol, other residues those AtomWorks places. When True (default), preserve finite input hydrogen coordinates, except coincident hydrogen-parent pairs, which are rebuilt with a warning, and build only missing hydrogens and heavy-atom sidechains. When False, residues with complete heavy atoms are packed with OptHSampler to optimize hydrogen positions and NHQ flips, while residues with missing heavy atoms are rebuilt with DunbrackChiSampler. Generated ligand types may still rebuild hydrogens whose names changed during parameter generation; pass trust_hydrogen_names=True only when those names are known to match the prepared database.

  • prepare_ligands – If True, detect and prepare non-standard residues (see build_context_from_biotite for details).

  • ligand_ph – pH AtomWorks protonates residues lacking hydrogens at (default 7.4).

  • strict_atom_types – If True, unknown ligand atom types raise errors instead of using a fallback element heuristic.

  • strict_ligands – If True (default), raise when a detected ligand cannot be prepared and registered, instead of silently dropping it. Pass False to warn-and-skip. Only used when prepare_ligands=True.

  • ligand_params_files – Optional list of tmol YAML params file paths.

  • chem_comp_types – {comp_id: type} from the input file’s _chem_comp table, which says whether a residue belongs to a polymer where the file does not number it along a sequence. Only used when prepare_ligands=True.

  • ligand_seed – Fixed RNG seed for the conformer each prepared residue is built from, making preparation reproducible. Only used when prepare_ligands=True.

  • packer_seed – Seed of the packer that places hydrogens and builds missing side chains; unseeded, it continues torch’s global random state.

  • return_context – If True, return (pose_stack, PoseBuildContext).

  • context – Reusable context from build_context_from_biotite. It must be on torch_device and is mutually exclusive with param_db and prepare_ligands=True.

  • atom37_coords – Optional coordinates shaped [pose, token, 37, xyz]. When supplied, mapped coordinates are read from this tensor using the input structure’s integer token_id and atom37_slot annotations. Wholly finite mapped triplets are authoritative; all-NaN mapped triplets are missing and do not fall back to the Biotite coordinates. Unmapped finite Biotite atoms remain context, and absent leaf atoms are completed normally. Partial-NaN or infinite mapped triplets raise Atom37MappingError. The resulting pose coordinates remain connected to this tensor for autograd. Geometry-based missing-density and additional-disulfide detection are disabled so topology is fixed.

  • **kwargs – Additional arguments passed to pose_stack_from_canonical_form.

Returns:

PoseStack when no optional values requested and return_context is False. (PoseStack, PoseBuildContext) when return_context is True. (PoseStack, dict) when optional return values were requested via kwargs. Fragmented poses expose their block mapping as pose_stack.split_block_mapping.

tmol.io.pose_stack_from_cif(cif_path, device, *, model: int = 1, assembly_id: str | None = None, residue_start: int | None = None, residue_end: int | None = None, **kwargs)[source]#

Construct a PoseStack from a PDB, CIF or binary/compressed CIF structure.

Reads the structure with atom_array_from_cif(), so the residues carry every atom their chemistry declares and the unresolved ones arrive at NaN for pose construction to rebuild. Further keyword arguments are passed to tmol.io.pose_stack_from_biotite().

tmol.io.atom_array_from_cif(cif_path, *, model: int = 1, assembly_id: str | None = None)[source]#

Read a PDB, CIF, compressed CIF, binary CIF, MOL2 or SDF into an AtomArray.

Preserve supplied names, coordinates, hydrogens and covalent connections. Declared unresolved heavy atoms have NaN coordinates. Chemistry the file leaves out is supplemented from the component dictionary.

Missing chemical metadata is allowed here. Pose construction uses existing parameters for known residues and requires additional chemistry only when preparing an unknown residue. Pose construction controls hydrogen optimization and whether to trust regenerated ligand hydrogen names.

tmol.io.component_chemistry_from_cif(cif_path) → dict[source]#

The complete chemistry a CIF declares for each component it names.

A structure records only the atoms it resolved, but it declares the whole component in chem_comp_atom and chem_comp_bond – names, elements and bond orders, independent of what the density showed. That declaration is what residue-type generation needs, since a residue built from resolved atoms alone is a different molecule: an unresolved sidechain tip becomes a real methyl once the pipeline protonates it.

Returns {comp_id: AtomArray} with no coordinates – the pipeline generates its own conformer – and only for components the file bonds: connectivity cannot be guessed, so a multi-atom component with no authored bonds is left out for the caller to resolve another way.

tmol.io.pose_stack_from_pdb(pdb_lines_or_fname: str | list | Path, device: device, *, residue_start: int | None = None, residue_end: int | None = None, res_not_connected: Tensor[slice(None, None, None), slice(None, None, None), 2] | None = None, **kwargs) → PoseStack | tuple[PoseStack, dict] | tuple[PoseStack, PoseBuildContext][source]#

Read PDB through AtomWorks and the shared annotated-array constructor.

Accept a path, PDB text or a list of lines. Residue slicing uses a half-open index range. Additional keywords follow pose_stack_from_file(), including prepare_ligands and return_context. Supplied hydrogens are retained and hydrogen optimization is disabled unless explicitly requested.

tmol.io.pose_stack_to_pdb_string(pose_stack: Any) → str[source]#

Convert a PoseStack into PDB text suitable for molecular viewers.

Return one interactive viewer for several labeled AtomArray selections.

Selection values may be boolean atom masks. Query strings are also accepted when the supplied AtomArray provides a callable aa.mask(query) method. Selection results are resolved in Python and exact PDB atom serials are baked into the HTML. This avoids viewer-side query-language differences and remains exact when atom names or residue identifiers are duplicated. Clicking a label restyles the same model and animates the camera to the selected atoms, following the AtomWorks selection-gallery interaction.

Parameters:
  • atom_array – Structure displayed by the shared viewer.

  • selections – Display labels mapped to boolean atom masks or query strings.

  • notes – Optional explanatory text keyed by selection label.

  • width – Viewer width in pixels.

  • height – Viewer height in pixels.

  • highlight_color – Color used for the active selection.

tmol.io.switchable_view(structures: Mapping[str, Any], *, notes: Mapping[str, str] | None = None, width: int = 720, height: int = 420)[source]#

Return HTML that switches one 3Dmol viewer among labeled structures.

Parameters:
  • structures – Ordered mapping of display labels to structures accepted by view().

  • notes – Optional mapping of structure labels to short explanatory text.

  • width – Viewer width in pixels.

  • height – Viewer height in pixels.

tmol.io.to_atom_lines(atom_records)[source]#

Convert atom records into ATOM lines.

tmol.io.to_pdb(atom_records)[source]#

Atom record DataFrame as pdb text.

tmol.io.to_pdb_lines(atom_records)[source]#

Yields atom record DataFrame as pdb lines.

tmol.io.view(model: Any, *, width: int = 720, height: int = 420, style: Literal['cartoon', 'stick'] = 'cartoon', background_color: str = 'white', cartoon_color: str = 'spectrum', show_sidechains: bool = True, show_heteroatoms: bool = True, show_hover: bool = True, zoom_to: dict | None = None, highlighted: object | None = None, highlight_color: str = '#e83e8c')[source]#

Create a draggable py3Dmol viewer for a molecular structure.

model may be a PoseStack, a Biotite AtomArray or AtomArrayStack, PDB text, or a PDB path. highlighted is an optional boolean mask over the atoms in the first model; highlighted atoms are shown as thicker sticks and spheres. The return value remains a real py3Dmol.view object for compatibility with existing notebooks.

tmol.io.write_pose_stack_pdb(pose_stack: PoseStack, fname_out: str, merge_fragments: bool = True, **kwargs)[source]#

Write a PDB-formatted file to disk given an input PoseStack. Optionally, additional arguments may be passed to the inner function “atom_records_from_pose_stack.” Fragmented ligands use their original residue identity by default; pass merge_fragments=False to keep fragment residues separate.