Development#
For build requirements and platform support, see installation.
Local setup#
git clone https://github.com/uw-ipd/tmol.git
cd tmol
python -m pip install "scikit-build-core>=0.10" "cmake>=3.24,<4" "pybind11>=2.12" ninja packaging
python -m pip install --no-build-isolation -e ".[dev]"
Install the intended PyTorch build first. Disabling build isolation keeps the compiler and runtime on the same PyTorch installation.
TMol locates the Torch and pybind11 CMake packages from this Python
interpreter; CMAKE_PREFIX_PATH and pybind11_DIR do not need to be set.
Requirements:
Python 3.11 or newer.
PyTorch 2.8 or newer.
A C++20-capable compiler (TMol uses C++17 with PyTorch 2.8–2.12).
CMake 3.24 or newer.
CUDA toolkit with
nvccfor CUDA builds.
Without CUDA, use a CPU-only build:
pip install --no-build-isolation -e . -Ccmake.define.TMOL_ENABLE_CUDA=OFF
Build extensions#
TMol builds extensions with CMake through scikit-build-core.
# Production extensions
pip install --no-build-isolation -e .
# Include test-only C++/CUDA extensions
pip install --no-build-isolation -e ".[dev]" \
-Ccmake.define.TMOL_BUILD_TESTS=ON
# Select GPU architectures
pip install --no-build-isolation -e . -Ccmake.define.CMAKE_CUDA_ARCHITECTURES="80;90"
# Control parallelism
MAX_JOBS=4 pip install --no-build-isolation -e . -Ccmake.define.TMOL_NVCC_THREADS=2
Build settings:
Variable |
Default |
Meaning |
|---|---|---|
|
|
|
|
|
Build test-only extensions. |
|
|
Threads per |
|
|
Turn off for CPU-only builds. |
|
auto |
Maximum parallel build jobs. |
Extension loading#
TMol can load kernels two ways:
AOT: pre-built shared libraries bundled in the installed wheel.
JIT: source files compiled on first use through
torch.utils.cpp_extension.
Environment variables:
Variable |
Effect |
|---|---|
|
Force JIT mode. |
|
Try AOT first, then JIT if AOT is unavailable. |
Use JIT mode while editing kernels. Use AOT mode for normal installed packages.
On Linux, the JIT loader detects PyTorch’s OpenMP CPU backend and propagates
the corresponding compiler and linker flags, so at::parallel_for has the
same thread behavior as the CMake-built extension. OMP_NUM_THREADS and
torch.set_num_threads() therefore apply consistently in both modes.
Tests#
# All tests
pytest tmol/tests/ -v
# Specific file
pytest tmol/tests/score/test_score_function.py -v
# Skip CUDA-parametrized cases
pytest tmol/tests/ -v -k "not cuda"
# Coverage
pytest tmol/tests/ --cov=./tmol --junitxml=results.xml
# Benchmarks
pytest --benchmark-enable --benchmark-only --benchmark-max-time=.1
Ligand charge generation is intentionally strict. Partial charges come from the SMILES to OpenBabel MMFF94 mol2 step and are applied by atom index. There is no RDKit/Gasteiger fallback and no charge-mode switch.
Containers#
Docker:
docker build -t tmol-dev -f containers/docker/tmol-dev.Dockerfile .
docker run --gpus all -it -v "$(pwd):/tmol_host" -w /tmol_host tmol-dev bash
pip install --no-build-isolation -e .
Apptainer:
apptainer build tmol-dev.sif containers/apptainer/tmol-dev.def
apptainer run --nv --bind "$(pwd):/tmol_host" tmol-dev.sif
CI#
GitHub Actions runs linting, the Linux x86-64, Linux aarch64, and Apple Silicon CPU suites, and CPU documentation builds on hosted runners. CUDA tests, benchmarks, and the CUDA-only tutorial smoke tests use the self-hosted DIGS runner. PR docs builds upload rendered HTML artifacts; same-repository PRs also publish previews under the Pages site.
Releasing#
Versioned wheel and sdist publication happens from v* tags. The tag version
must match [project].version in pyproject.toml.
scripts/release_matrix.py defines the release wheels. For releases after
0.1.59, the matrix contains 34 CUDA wheels and 12 CPU wheels. CPU wheels use
PyTorch 2.14 with Python 3.11–3.14 on Linux x86-64, Linux aarch64, and Apple
Silicon. CPU support is listed separately from the broader CUDA matrix.
Before publication, CI installs every wheel through a staging index in a fresh environment, loads the native extension, and checks scoring and gradients. CPU wheels are also tested under their public PyPI versions. A separate lane builds and installs the indexed source distribution. Fast PR tests cover pip’s resolver, source metadata, PyTorch constraints, missing variants, and hashes.
After these checks pass, the workflow publishes the ABI-qualified wheels to GitHub, adds versioned wheel pages to GitHub Pages, and uploads standard CPU wheels plus the source distribution to PyPI. CPU repackaging preserves native code, retains the exact PyTorch minor constraint, and rebuilds wheel RECORD hashes. PyPI metadata must use package-index dependencies; direct Git and URL dependencies are rejected.
Before using a versioned wheel URL, check the GitHub Releases page. The version in a checkout is not proof that a release has been published.
Code style#
TMol uses Black for Python formatting, Flake8 for linting, and clang-format for C++/CUDA formatting.
black --check .
black .
flake8
Run pre-commit before opening a PR:
pre-commit run --all-files