Benchmarking#

Use the developer harness to measure implementation regressions. For application throughput, batch size, and memory use, see GPU batching. Both require a development installation.

Run benchmarks#

Use dev/bin/benchmark with pytest selectors:

dev/bin/benchmark tmol/tests/score -k cuda-full-lk_ball

The wrapper enables pytest benchmarks, prints a summary, and writes JSON results under dev/benchmark/.

Compare revisions#

dev/bin/compare_benchmark compares benchmark results across revisions. Put pytest arguments first, then revisions after --.

dev/bin/compare_benchmark tmol/tests/score -k cuda-full-lk_ball -- origin/master

The meta-revision TREE means the current working tree:

dev/bin/compare_benchmark tmol/tests/score -k cuda-full-lk_ball -- TREE HEAD

Ancillary benchmark plots live near the tests as plot_*.py scripts.

Profiling#

dev/bin/profile_benchmark runs a short pytest benchmark under Nsight Systems by default:

dev/bin/profile_benchmark --output profile/ljlk \
  tmol/tests/score -k cuda-full-ljlk

Use Nsight Compute when kernel-level counters are needed; arguments after -- are forwarded to the profiler:

dev/bin/profile_benchmark --tool ncu --output profile/ljlk-kernels \
  tmol/tests/score -k cuda-forward-ljlk-100 -- \
  --kernel-name regex:ljlk --launch-count 20

The output prefix defaults to dev/profile/<host>/<UTC timestamp>. Keep pytest selectors narrow: profiling every parametrized benchmark produces a very large trace and makes hardware-counter collection unnecessarily slow. If the wrapper is not launched from the development environment, pass its interpreter with --python /path/to/venv/bin/python or set TMOL_PROFILE_PYTHON.