Repository navigation
Releases: lightvector/KataGo
Release list
CUDA Speedup for Turing (RTX 20xx, etc)
This is a quick minor optimization release for the CUDA backend that helps certain older GPUs (RTX 20xx, and T4, mainly). For best performance, use this release if you have an NVIDIA GPU. For all other backends (AMD GPUs, etc) see the prior release v1.18.1.
Download the latest neural nets to use with this engine release at https://katagotraining.org/.
Also, for 9x9 boards or for boards larger than 19x19, see https://katagotraining.org/extra_networks/ for networks specially trained for those sizes!
KataGo is continuing to improve at https://katagotraining.org/ and if you'd like to donate your spare GPU cycles and support it, it could use your help there!
Getting Started / Choosing a Backend
If you're a new user, this section has tips for getting started and basic usage! For choosing a backend:
NVIDIA GPU
- You'll also have to install CUDA and one of CUDNN or TensorRT from nvidia depending on your choice.
- CUDA+CUDNN
- Faster startup times than TensorRT, good performance, and v1.18.x makes it faster for transformer models on most GPUs.
- Use CUDNN >= 9.8.0. The CUDNN 8.9.7 builds will be a LOT slower when running transformers - consider upgrading if you're still on old CUDNN. We're offering CUDNN 8.9.7 builds only for continuity with prior releases.
- TensorRT
- Longer startup time. May be faster on older convolutional nets still, but tends to be slower than CUDA+CUDNN on transformers on recent versions and strong GPUs.
- TensorRT versions older than 10 are not supported.
- See the prior release v1.18.1 for prebuilt exes.
Other GPUs
See the prior release v1.18.1
New changes in v1.18.2
- Added support for tensor-core attention on Turing GPUs on CUDA sm_75 (RTX 20xx GPUs, Tesla T4). On the GPUs of this generation with tensor cores, this can make transformer models more than 50% faster. Thanks to @nlzy for this improvement!
- Note: GTX 16xx GPUs are also sm_75 but lack tensor cores and aren't expected to benefit. It in some testing it also seemed to be the case that on such GPUs, FP16 is generally slower than FP32, so if you have a GTX 16xx GPU, try
useFP16 = falsein your config and you might get a major speedup.
- Note: GTX 16xx GPUs are also sm_75 but lack tensor cores and aren't expected to benefit. It in some testing it also seemed to be the case that on such GPUs, FP16 is generally slower than FP32, so if you have a GTX 16xx GPU, try
Better Benchmark/Genconfig, Minor bugfixes
For CUDA backend specifically, see v1.18.2, for a newer release! For all other backends, this v1.18.1 is the latest release.
This is a bugfix and minor improvement release on top of v1.18.0, including making benchmark and genconfig help test and recommend additional performance optimizations that might not be obvious for casual users, beyond the major ones that v1.18.0 added.
You can download the latest neural nets for KataGo, including the new strong transformer neural nets from KataGo's main training run.
Also, for 9x9 boards or for boards larger than 19x19, see https://katagotraining.org/extra_networks/ for networks specially trained for those sizes!
KataGo is continuing to improve at https://katagotraining.org/ and if you'd like to donate your spare GPU cycles and support it, it could use your help there!
Getting Started / Choosing a Backend
If you're a new user, this section has tips for getting started and basic usage! For choosing a backend:
NVIDIA GPU
Use CUDA+CUDNN or TensorRT. Both are decent, either may be better performance. For transformers, CUDA is typically a bit faster on good GPUs (new in 1.18.x - recent optimizations have let it overtake TensorRT).
- You'll also have to install CUDA and one of CUDNN or TensorRT from nvidia depending on your choice.
- CUDA+CUDNN
- Faster startup times than TensorRT, good performance, and v1.18.x makes it much faster for transformer models on many GPUs.
- Use CUDNN >= 9.8.0. The CUDNN 8.9.7 builds will be a LOT slower when running transformers - consider upgrading if you're still on old CUDNN. We're offering CUDNN 8.9.7 builds only for continuity with prior releases.
- TensorRT
- Longer startup time. May be faster on older convolutional nets still, but tends to be slower than CUDA+CUDNN on transformers on recent versions and strong GPUs.
- TensorRT versions older than 10 are not supported.
AMD GPU
New in v1.18.x - try the new ROCm backend, which should be much faster than OpenCL.
- The Linux build requires installing ROCm on your own, builds are offered for ROCm 7.2.4 and ROCm 7.14.0.
- It's worth testing and installing ROCm 7.14.0 if you can. ROCm 7.2.4 may be a LOT slower on some systems.
- The Windows build does NOT require installing ROCm (you need a reasonably recent AMD (Adrenalin) driver installed), and bundles everything it needs from ROCm 7.13. Download the package for your GPU family:
gfx103X- Radeon RX 6000 series (RDNA2)gfx110X- Radeon RX 7000 series (RDNA3), including RDNA3 APUs such as the Radeon 780Mgfx1151- Ryzen AI Max ("Strix Halo") APUsgfx120X- Radeon RX 9000 series (RDNA4)
- The windows builds for ROCm are a bit experimental and were cross-compiled from a machine that doesn't actually have an AMD GPU for testing it - please let me know if there are any issues with them!
- The DLLs provided and the
rocblasandhipblasltfolders must stay next to katago.exe.- The exe is actually the same across these packages and is compiled for more GPUs than these, only the bundled AMD libraries differ. If your GPU is not in the list, for example an RX 5000 series (RDNA1) GPU or a Ryzen AI 300 series APU (gfx1150/gfx1152), it may still work: download the ROCm 7.13 Windows tarball for your GPU family from https://repo.amd.com/rocm/tarball-multi-arch/ (a large download), then copy the DLLs from its
binfolder over KataGo's and replace KataGo'srocblasandhipblasltfolders with the ones from there. Copy them into KataGo's own directory rather than relying on PATH, since AMD driver installation can itself leave conflicting copies of some of these DLLs on the system. - The first run with a given neural net and board size can take 45 seconds or more before anything seems to happen, while MIOpen tunes its convolution kernels for your GPU - KataGo is not hung. The results are cached under
%USERPROFILE%\.miopen\, so later startups are fast.
- The exe is actually the same across these packages and is compiled for more GPUs than these, only the bundled AMD libraries differ. If your GPU is not in the list, for example an RX 5000 series (RDNA1) GPU or a Ryzen AI 300 series APU (gfx1150/gfx1152), it may still work: download the ROCm 7.13 Windows tarball for your GPU family from https://repo.amd.com/rocm/tarball-multi-arch/ (a large download), then copy the DLLs from its
Intel GPU / NPU
New in v1.18.x - try the new ONNX Runtime backend with the OpenVINO execution provider (the onnx-openvino package).
- You must set
onnxProvider = openvinoin your config. - Both the Windows and Linux builds are self-contained (ONNX Runtime and the OpenVINO runtime are bundled, no OpenVINO install needed), so beyond downloading them you only need the Intel drivers for your hardware:
- Intel GPU: on Windows, the normal Intel graphics driver. On Linux, the GPU compute driver (e.g. the
intel-opencl-icdpackage on Ubuntu). - Intel NPU: on Windows, the Intel NPU driver, which also comes via Windows Update (Windows 11 only). On Linux, the Intel NPU driver. Also set
onnxOpenVINODeviceType = NPUin your config, since the default device is GPU.
- Intel GPU: on Windows, the normal Intel graphics driver. On Linux, the GPU compute driver (e.g. the
- OpenVINO recompiles the model at every startup. To cache the compiled model between runs, set
onnxOpenVINOCacheDirto a directory in your config. - Windows also has a separate
onnx-directmlpackage, which runs on any DirectX 12 GPU (Intel, AMD, or NVIDIA) if you setonnxProvider = directml. It is self-contained too, and mostly worth trying if the native backends and OpenCL do not work for you. - See Compiling.md for the status of the other execution providers it supports.
- ONNX is not currently enabled for contribute on this release, due to some concern about the large surface area of possibly flaky execution providers and newness of backend, but it should work fine for all other KataGo usage.
Other / Old GPUs
If you have some other GPU, or older GPUs on which the above doesn't work, or don't want to install extra stuff, try OpenCL. OpenCL will often work in cases others don't support, but be much slower, particularly for transformer models.
MacOS
For MacOS, use the Metal backend, which you can generally get by installing KataGo from homebrew, which usually updates not too long after KataGo's own release.
CPU
If you need a pure-CPU version of KataGo, use Eigen AVX2. It will be quite slow compared to GPU. If somehow you're on an ancient CPU as well and Eigen AVX2 doesn't work, you can try Eigen, which will be even slower.
Other notes
-
+bs50- these are just for fun, and don't support distributed training but DO support board sizes up to 50x50. They may also be slightly slower and will use much more memory, even when only playing on 19x19, so use them only when you really want to try large boards. -
Linux executables were compiled on a 22.04 Ubuntu machine using AppImage. You will still need to install e.g. correct versions of Cuda/TensorRT or have drivers for OpenCL, etc. on your own. Compiling from source is also not so hard on Linux, see the "TLDR" instructions for Linux here.
Changes in v1.18.1
- The
benchmarkcommand now also tests a couple of other settings and recommends the config changes for them if they are faster:- Testing a max batch size of half the number of search threads, which splits batches more evenly and be better on many backends or GPUs.
- Testing 2 NN server threads per GPU (CUDA and ROCm only, the main backends where this is known to often be a speedup). As mentioned in the v1.18.0 release notes, this is now often the best-performing setup on these backends, and now the benchmark tries it for you and if it's good will indicate how to configure it.
- New command line flags
-no-server-thread-testand-no-half-batch-size-testskip these extra tests. - The
genconfigcommand runs the same extra tests and automatically writes any faster settings into the config it generates for you.
- Fixed a bug on the CUDA backend where every NN server thread would initialize and allocate some memory on GPU 0, even when KataGo was configured to use only other GPUs.
- If an NN server thread dies with an unrecoverable error, such as running out of GPU memory, KataGo now logs and prints the error before exiting, rather than possibly aborting with no message at all on some platforms.
- Fixed a possible invalid memory access on the CUDA and ROCm backends when shutting down after a GPU error partway through a neural net evaluation.
- Added slightly better error checking and tests for threaded queries to GPUs.
New backends, major CUDA optimizations, rules fix support
This is not the latest release - see a more recent release at v1.18.1
This release adds major optimizations to the CUDA backend, and two new backends - ROCm for AMD GPUs, and ONNX for various other GPUs/accelerators. See below!
Getting Started / Choosing a Backend
If you're a new user, this section has tips for getting started and basic usage! For choosing a backend:
NVIDIA GPU
Use CUDA+CUDNN or TensorRT. Both are decent, either may be better performance.
- You'll also have to install CUDA and one of CUDNN or TensorRT from nvidia depending on your choice.
- TensorRT
- For transformers (which will soon be the best models), CUDA 13 + TensorRT 10.16 is good but older TensorRT versions will often be outperformed by CUDA+CUDNN.
- TensorRT versions older than 10 are not supported.
- CUDA+CUDNN
- Faster startup times than TensorRT, still pretty good performance, and this release makes it much faster for transformer models on many GPUs (see below).
- Use CUDNN >= 9.8.0. The CUDNN 8.9.7 builds will be a LOT slower when running transformer models - consider upgrading if you're still on old CUDNN. We're offering CUDNN 8.9.7 builds only for continuity with prior releases.
AMD GPU
New in this release - try the new ROCm backend, which should be much faster than OpenCL.
- The Linux build requires installing ROCm on your own, builds are offered for ROCm 7.2.4 and ROCm 7.14.0.
- The Windows build does NOT require installing ROCm (you need a reasonably recent AMD (Adrenalin) driver installed), and bundles everything it needs from ROCm 7.13. Download the package for your GPU family:
gfx103X- Radeon RX 6000 series (RDNA2)gfx110X- Radeon RX 7000 series (RDNA3), including RDNA3 APUs such as the Radeon 780Mgfx1151- Ryzen AI Max ("Strix Halo") APUsgfx120X- Radeon RX 9000 series (RDNA4)
- The windows builds for ROCm are a bit experimental and were cross-compiled from a machine that doesn't actually have an AMD GPU for testing it - please me know if there are any issues with them!
- The DLLs provided and the
rocblasandhipblasltfolders must stay next to katago.exe.- The exe is actually the same across these packages and is compiled for more GPUs than these, only the bundled AMD libraries differ. If your GPU is not in the list, for example an RX 5000 series (RDNA1) GPU or a Ryzen AI 300 series APU (gfx1150/gfx1152), it may still work: download the ROCm 7.13 Windows tarball for your GPU family from https://repo.amd.com/rocm/tarball-multi-arch/ (a large download), then copy the DLLs from its
binfolder over KataGo's and replace KataGo'srocblasandhipblasltfolders with the ones from there. Copy them into KataGo's own directory rather than relying on PATH, since AMD driver installation can itself leave conflicting copies of some of these DLLs on the system. - The first run with a given neural net and board size can take 45 seconds or more before anything seems to happen, while MIOpen tunes its convolution kernels for your GPU - KataGo is not hung. The results are cached under
%USERPROFILE%\.miopen\, so later startups are fast.
- The exe is actually the same across these packages and is compiled for more GPUs than these, only the bundled AMD libraries differ. If your GPU is not in the list, for example an RX 5000 series (RDNA1) GPU or a Ryzen AI 300 series APU (gfx1150/gfx1152), it may still work: download the ROCm 7.13 Windows tarball for your GPU family from https://repo.amd.com/rocm/tarball-multi-arch/ (a large download), then copy the DLLs from its
Intel GPU / NPU
New in this release - try the new ONNX Runtime backend with the OpenVINO execution provider (the onnx-openvino package).
- You must set
onnxProvider = openvinoin your config. - Both the Windows and Linux builds are self-contained (ONNX Runtime and the OpenVINO runtime are bundled, no OpenVINO install needed), so beyond downloading them you only need the Intel drivers for your hardware:
- Intel GPU: on Windows, the normal Intel graphics driver. On Linux, the GPU compute driver (e.g. the
intel-opencl-icdpackage on Ubuntu). - Intel NPU: on Windows, the Intel NPU driver, which also comes via Windows Update (Windows 11 only). On Linux, the Intel NPU driver. Also set
onnxOpenVINODeviceType = NPUin your config, since the default device is GPU.
- Intel GPU: on Windows, the normal Intel graphics driver. On Linux, the GPU compute driver (e.g. the
- OpenVINO recompiles the model at every startup. To cache the compiled model between runs, set
onnxOpenVINOCacheDirto a directory in your config. - Windows also has a separate
onnx-directmlpackage, which runs on any DirectX 12 GPU (Intel, AMD, or NVIDIA) if you setonnxProvider = directml. It is self-contained too, and mostly worth trying if the native backends and OpenCL do not work for you. - See Compiling.md for the status of the other execution providers it supports.
- ONNX is not currently enabled for contribute on this release, due to some concern about the large surface area of possibly flaky execution providers and newness of backend, but it should work fine for all other KataGo usage.
Other / Old GPUs
If you have some other GPU, or older GPUs on which the above doesn't work, or don't want to install extra stuff, try OpenCL. OpenCL will often work in cases others don't support, but be much slower, particularly for transformer models.
MacOS
For MacOS, use the Metal backend, which you can generally get by installing KataGo from homebrew, which usually updates not too long after KataGo's own release.
CPU
If you need a pure-CPU version of KataGo, use Eigen AVX2. It will be quite slow compared to GPU. If somehow you're on an ancient CPU as well and Eigen AVX2 doesn't work, you can try Eigen, which will be even slower.
Other notes
-
+bs50- these are just for fun, and don't support distributed training but DO support board sizes up to 50x50. They may also be slightly slower and will use much more memory, even when only playing on 19x19, so use them only when you really want to try large boards. -
Linux executables were compiled on a 22.04 Ubuntu machine using AppImage. You will still need to install e.g. correct versions of Cuda/TensorRT or have drivers for OpenCL, etc. on your own. Compiling from source is also not so hard on Linux, see the "TLDR" instructions for Linux here.
Major Changes this Release
New backends
(thanks to @Looong01 @seniorfish)
This release adds a ROCm backend for AMD GPUs using HIP and MIOpen, which should be a much faster option than OpenCL for modern AMD GPUs, especially for transformers, and a new ONNX runtime backend (ONNX Runtime) which supports many kinds of hardware, useful for hardware with no other KataGo backend - in particular Intel GPUs and NPUs via the OpenVINO provider, and DirectML on Windows.
Major CUDA and ROCm performance optimizations
(thanks to wyz @doomoooo) and various others for helping explore the major optimizations possible!)
The CUDA backend received a large round of optimization work (batched/fused matrix multiplies, custom flash attention kernel, fused feed-forward layers, reduced CPU-GPU synchronization, and more), with most of the optimizations shared with the new ROCm backend as well:
- Transformer models can be anywhere from 1.2x to 2x faster in overall search speed depending on the GPU, with the largest gains on the newest GPUs (e.g. NVIDIA Blackwell / RTX 50xx).
- Convolutional models are slightly faster, with the biggest gains in configs using multiple neural net server threads.
- GPU memory usage for transformer models on CUDA is greatly reduced.
- Setting
numNNServerThreadsPerModel = 2(or two threads per-GPU, if multiple GPUs) is now often the best-performing config on both CUDA and ROCm - you can try this if you're optimizing for performance.
Support for upcoming rules fix
Added support for fixing an issue with territory scoring (Japanese-like rules). In some thousand-year-ko situations, the cleanup phase could contain pass fights over the ko mouths in seki, making the score and winrate evaluation wrong. KataGo's formal rules are now updated to version 3, where empty points adjacent to a region in atari no longer count as territory - see the updated rules documentation.
The fix will not immediately take effect, but once contributors to https://katagotraining.org/ upgrade and neural nets are trained sufficiently on the new version, then the issue should be fixed. Older nets will still evaluate with the old behavior.
Other user-facing feature additions/changes:
- Added new command
katago benchmarknnthat benchmarks raw neural net evaluation throughput without any search, useful for comparing backends and GPU settings in isolation. - Added better GPU error checks and training data quality safeguards to
contributecommand. - Removed the
cudaDisableWarmupconfig option, warmup now always runs.
User-facing bugfixes:
- Fixed a crash in the
evalsgfcommand when printing the lead estimate together with the search graph, and fixed tree-averaged ownership missing from its JSON output.
Dev-facing details and changes:
- Added new command
katago dumponnx(TensorRT and ONNX backends) that writes out the ONNX graph KataGo builds for a model, and those two backends can also directly load a.onnxfile in place of a.bin.gzmodel, including ONNX files produced by other tooling. See ONNX_Model_Files.md for the format. - The CUDA and ROCm backends gained several config knobs that disable individual new optimization paths for debugging (e.g.
cudaUseMmaAttention,cudaUseFusedFFN). See the source code for the relevant debugging knobs and flags.
Python scripts changes:
- The Python board implementation is updated to match the territory scoring rules fix, and model confi...
TensorRT bugfixes
This is not the latest release - see a more recent release at v1.18.0
This is a quick bugfix release on top of v1.17.1 to fix some bugs in TensorRT. Only the TensorRT executables are provided in this release since this is the only backend that was affected by the bugs fixed. See v1.17.1 for all other backends and for the new strong transformer nets, and v1.17.0 for the much longer list of changes new in v1.17.x!
Bugs fixed in v1.17.2
- TensorRT 10.16 could nondeterministically segfault or hang on startup when using more than one GPU, when multiple GPU threads built their engines at the same time. Engine builds are now serialized across GPUs. (#1225, thanks @zsqdx!)
- TensorRT could fail to build the network with very large
maxBatchSizeand/or a large number of search threads, because KataGo was capping the TensorRT workspace at 1 GiB and attention tensors alone could exceed it. KataGo now leaves the workspace at TensorRT's automatic device-dependent sizing. (#1229, thanks @zsqdx!) - The humanSL model (
b18c384nbt-humanv0.bin.gzfrom here), failed on TensorRT in v1.17.0 and v1.17.1, erroring with "OnnxModelBuilder: SGF metadata encoder not yet supported". It should work now. - The built TensorRT executables on linux now have protobuf provided as part of the appimage, to resolve the issue where some systems have different versions (#1226)
This release bumps TensorRT plan and timing cache version, both because of the metadata encoder support above and to avoid reusing any caches that might have been polluted by prior bugs. So the first run of this version will be slower to start up while it re-tunes, as usual.
OpenCL, Windows, TRT Bugfixes
This is not the latest release - see a more recent release at v1.18.0
However, you can still download some of the new transformer models from this release, see downloads below
New Transformer Models (same as v1.17.0)
Included in this release for download are three very strong transformer models of various sizes that newly work with these 1.17.x releases:
- b10c384h6nbttflrs.bin.gz - smallest, stronger per visit than the strongest b18 main run models but usually as fast or faster than them
- b10c512h8nbt3tflrs-fson-silu-rsnh.bin.gz - medium, stronger per visit than the strongest b28 main run models but usually as fast or faster than them
- b11c768h12nbt3tflrs-fson-silu.bin.gz - largest transformer so far, stronger than the b40 zhizi main run models but usually as fast or faster than them
(see here for pytorch .ckpt versions of these transformers)
See https://katagotraining.org/networks/ for the main run models so far which still all work too. The main run as of the time of this release has not switched to transformers since this is the first version supporting them, but will switch soon once enough contributors upgrade to v1.17.0. See also https://katagotraining.org/extra_networks/ for other unusual models that can be used with KataGo.
Getting Started / Choosing a Backend
If you're a new user, this section has tips for getting started and basic usage! For choosing a backend:
-
If you have an NVIDIA GPU, use CUDA+CUDNN or TensorRT. You'll also have to install CUDA and one of CUDNN or TensorRT from nvidia depending on your choice.
- TensorRT - Usually best performance but longer startup times.
- For transformers, i.e. the best models, prefer CUDA 13 + TensorRT 10.16, older versions may be outperformed by CUDA+CUDNN.
- CUDA+CUDNN - Faster startup times than TensorRT, still pretty good performance, occasionally actually fastest.
- For transformers, i.e. the best models, usually slightly slower than TRT 10.16 (but not always), and faster than older TRT versions.
- Use CUDNN >= 9.8.0 if at all possible. The CUDNN 8.9.7 builds will be a LOT slower when running transformer models.
- (CUDA 13, CUDNN 9.24, TensorRT 10.16 builds are new in this release. TensorRT < 10 support has been dropped)
- TensorRT - Usually best performance but longer startup times.
-
If you have a non-NVIDIA GPU, or have an NVIDIA GPU but really don't want to install CUDNN/TensorRT, OpenCL works across GPUs but will be much slower. (Future releases will likely also add faster and better support for other GPUs, stay tuned)
-
For MacOS, use the Metal backend, which you can generally get by installing KataGo from homebrew, which usually updates not too long after KataGo's own release.
-
If you need a pure-CPU version of KataGo, use Eigen AVX2, or if your CPU is also ancient and Eigen AVX2 doesn't work, then try Eigen.
Other notes
-
+bs50- these are just for fun, and don't support distributed training but DO support board sizes up to 50x50. They may also be slightly slower and will use much more memory, even when only playing on 19x19, so use them only when you really want to try large boards. -
Linux executables were compiled on a 22.04 Ubuntu machine using AppImage. You will still need to install e.g. correct versions of Cuda/TensorRT or have drivers for OpenCL, etc. on your own. Compiling from source is also not so hard on Linux, see the "TLDR" instructions for Linux here.
Bugs fixed in v1.17.1
- OpenCL was crashing during tuning or contribute (#1221)
- Windows versions would fail at runtime if you built them using the latest 2026 MSVC and then tried to use them with outdated DLLs that didn't correspond to the version you built them with. (#1216)
- (technically not something you're supposed to do but if you're copying around exes and dlls from various places you can run into this)
- TensorRT if it was encountering certain GPU errors due to other configuration issues, was also slightly too picky about them and would fail even if it could fall back to other tuning that would work.
Transformers - Strong New Nets
This is not the latest release - see a more recent release at v1.17.1 for some bugfixes!
However, you can still download some of the new transformer models from this release, see downloads below
New Transformer Models
Included in this release for download are three very strong transformer models of various sizes that newly work with these 1.17.x releases:
- b10c384h6nbttflrs.bin.gz - smallest, stronger per visit than the strongest b18 main run models but usually as fast or faster than them
- b10c512h8nbt3tflrs-fson-silu-rsnh.bin.gz - medium, stronger per visit than the strongest b28 main run models but usually as fast or faster than them
- b11c768h12nbt3tflrs-fson-silu.bin.gz - largest transformer so far, stronger than the b40 zhizi main run models but usually as fast or faster than them
Major Changes this Release
Transformer support! (yay)
This release adds support for transformer neural nets in the C++ engine on all major backends (thanks @ChinChangYang for Metal/CoreML, and thanks @zsqdx for performance optimizations and @hzyhhzy for various training improvements, along with many others). Transformer models are generally much stronger given the same compute cost, and the main training run at katagotraining.org will soon be switching to a strong transformer and will therefore require v1.17.x to contribute to.
GPU Backend Fixes and Performance Improvements:
- Added CUDA config option
cudaUse1x1Matmulcontrolling whether 1x1 convolutions run as a cuBLAS matrix multiply, which is slightly faster in FP16. New default uses it when FP16 is enabled, which may result in a few percent speedup on some machines. - Although OpenCL is still not as optimized as other backends, made a variety of robustness improvements, and fixed a bug that prevented the use of tensor cores for 1x1 convs, speeding up models noticeably on some GPUs.
- Fixed a major memory leak in Metal backend that could eventually exhaust memory in long-running processes.
- Greatly reduced the memory usage of the on-device CoreML model conversion and ANE (peak memory during load roughly halved, and steady-state reduced also by a large factor).
- Various improvements to error handling, checks, and robustness in the Metal/CoreML backend.
(Metal/CoreML fixes by thanks to @ChinChangYang)
Other Feature additions/changes for KataGo users:
- GTP extensions
kata-raw-nnandkata-raw-human-nnnow accept an optional color argument to specify the player to move, and fixed the parsing of the optional optimism argument ofkata-raw-nn. See GTP_Extensions.md. - Analysis engine
allowMovescan now specify one entry for each player separately, instead of only a single entry total. - Increased the max komi that users can specify from 150 to 400.
- KataGo now logs the number of parameters of the neural net being loaded.
- Dropped support for TensorRT versions < 10.
Bugfixes for KataGo users:
- Fixed bug where the experimental eval cache (added v1.16.4, off by default) could reuse cached search results across changed search parameters in GTP and the analysis engine, causing major errors as very different winrates/scores could get blended together.
- Fixed New Zealand rules to use a default komi of 7 instead of 7.5, matching the actual New Zealand rules.
- Fixed OpenCL tuner bug where performing the initial tuning with FP16 disabled would incorrectly cache the GPU as not supporting FP16, silently degrading performance of later FP16 runs.
C++ selfplay/training-data changes and other internal changes:
- TensorRT backend now works by emitting ONNX in memory and parsing it with TensorRT's ONNX parser. Set
trtDisableOnnx=trueto fall back to the previous hand-constructed network graph (convnets only). - For TensorRT, added a
trtDumpDebugPlanToDirconfig option that dumps the emitted ONNX, the serialized engine, and per-layer engine info for debugging precision or fusion issues. - CUDA backend now requires cudnn-frontend library, for fused attention kernels for transformers.
- Transformer nets are exported using a new model version 17.
- Added experimental option to reanalyze positions during selfplay data generation (
useReanalyze,reanalyzeProp, and related options). After a game finishes, a random subset of the cheap-search positions are re-searched fully and recorded as training data, favoring positions where the cheap search result was surprising. - Added experimental option
useSearchValueSurpriseto compute value surprise for data weighting/selection from the search result of the turn itself relative to the raw net evaluation, rather than from the realized game outcome, so that data weighting does not condition on the game's actual result. - Added experimental options to support future migration work for pass-alive handling in trained nets and self-play to fix a pathological corner case caused by KataGo's rules, including new support for trained models to flag themselves as supporting it.
- Selfplay now logs some additional statistics about game generation.
- Adjusted handling of illegal moves when loading starting positions and sgfs, adding a new option to tolerate them.
- Fixed a minor concurrent access bug in training data writing.
Python scripts feature additions/changes:
- Major training performance optimizations, as much as 1.7x faster on some hardware setups on some model architectures (probably a bit less for convnets). Most optimizations are on by default and individually controllable via
KATAGO_*environment variables. - Added experimental support for the Aurora optimizer, a variant of Muon.
- Added
-lr-scheduleargument totrain.pyto specify an explicit piecewise-constant learning rate scale schedule from the command line. - Major cleanup of
shuffle.pyarguments and defaults. NOTE: deliberate compatibility break - some formerly-required arguments now have sensible defaults,-keep-target-rowsis now required (and accepts 'all'), and the deprecated-add-to-window-sizewas removed in favor of-add-to-data-rows. shuffle.pyhas a-num-wavesoption for very large whole-dataset shuffles that still maintains perfect permutation, a-dry-run-print-resource-costmode to estimate disk/memory costs without running, and better logging.- Added new script
refresh_batchnorm_stats.pythat exactly recomputes batchnorm running statistics of a checkpoint (for both the plain and SWA model) over a sample of data, improving models for checkpoints whose batchnorm stats are stale (particularly those at higher LR or large batch size). - Improved training data loading: added a
-data-prefetch-depthoption, reduced CPU and memory overhead, improved failure handling. benchmark_fresh_model.pyimprovements, including a mode that benchmarks the full training loop step rather than only the model forward/backward.- Model export now computes a bound on the possible attention logit magnitudes of transformer nets and refuses to export nets whose logits could grow large enough to compromise FP16 attention masking in inference backends (overridable), and
train.pynow has an optional penalty to discourage large attention logits during training. - Cleaned up model configs and added configs for some larger transformer models.
- Fixed optimizer name not being set in the training state when loading a checkpoint without optimizer state (thanks @loker404!).
- Fixed a barrier issue in distributed training that could cause validation to time out.
- Fixed
shuffle.pylogged data row ranges to account for-add-to-data-rows. - Minor fixes and updates to various other scripts.
Build/CI fixes:
- Fixed Windows CI to work with newer Visual Studio versions rather than requiring VS 2022 (thanks @ChinChangYang).
- Fixed stale build cache issues in GitHub Actions builds.
Bunch of bugfixes yay, and better Metal for MacOS
This is not the latest release - see a more recent release at v1.17.0
Changes this Release
Improved Metal backend for macOS
- The Metal backend now supports hybrid CPU+GPU+ANE (Apple Neural Engine). Each server thread can be configured to run on GPU (via MPSGraph) or ANE (via CoreML) using the
metalDeviceToUseThread<N>config option. - The CoreML model converter (katagocoreml) is now vendored into the repo, so building the Metal backend does not require an external Homebrew package.
- Various improvements to the internal implementation, error reporting
Other feature additions/changes
- Added option to report MCTS triple-ko (no-result) probability in GTP and analysis via
includeNoResultValue. - Minor improvements for book generation - multithreaded book processing, some additional commands and arguments.
User-facing Bugfixes
- Fixed crash on
clear_cachein the analysis engine. - Fixed endless recursion on Windows in threadsafequeue.h.
- Fixed
cpuctUtilityStdevPriorallowing a value of 0, which could cause a divide-by-zero. - Replaced recursive SGF parsing with iterative approach to avoid stack overflow, other small fixes.
- Fixed priority mutex bug where low-priority path called the high-priority path.
- Fixed bad interaction between certain hacks and avoid moves in book generation.
- Fixed inconsistent genmove params setting in GTP.
- Fixed incorrect error field name when reporting komi errors in the analysis engine.
- Fixed issue where firstReportDuringSearchAfter could fire before any results were available.
Build/compatibility fixes
- Added support for CUDA 13.0.
- Added support for CMake 4.* with the Metal backend.
- Fixed nvinfer library detection on Windows (use nvinfer_10).
- Fixed Eigen cmake configuration to be more general.
- Fixed Metal backend CMAKE_OSX_SYSROOT not being set.
- Fixed Makefile git revision detection.
- Now compiles with CMake build type flags.
Major Python/training script changes preparing for new architectures
- Added support for Muon and AdamW optimizers. Muon trains a lot faster than all prior optimizers in this repo.
- Added experimental support for training transformers, with a variety of basic features. No support on C++ side yet.
- Added support for torch.compile, pytorch side benchmarking script.
- Variety of python training script changes and fixes.
- Fixed some bash scripts to better handle whitespace in file paths, detect python, etc.
- Fixed
play.pycrash due to undefined var and escape sequence.
Internal/code quality/dev-relevant changes
- Upgraded to tclap 1.2.5.
- Various const correctness improvements and refactoring, slight CPU-side performance improvements.
- Tightened assertions and testing outside of hot paths.
- Removed demoplay command.
- Fixed out-of-bounds write in training data score distribution.
- Various other minor fixes and improvements
Experimental Eval Cache, Bugfixes
This is not the latest release - see a more recent release at v1.16.5!
Changes this Release
- Added experimental eval-caching feature, NOT enabled by default yet. Enable it by setting
useEvalCache=truein the gtp.cfg or analysis.cfg. This will make it so that while analyzing with KataGo interactively, if you walk deeper into a variation and have KataGo realize a good move that was a blind spot, then when you walk backward to an earlier point, the search will be far more likely to solve that tactic as well and now able to analyze the earlier position in light of the newly solved tactic. - Subtree value bias now no longer applies to nodes following passing, to avoid conflating evals that don't have a discriminating local pattern.
- Fixed issue in contributing to distributed training where where unnecessary locking might block starting new games while old games were uploaded.
- Fixed some typos that could cause the python testing script
python/play.pyor getting input features in python to crash - Various internal refactors and cleanups
Various bugfixes, mingw support
This is not the latest release - see a more recent release at v1.16.5!
Changes this Release
Added support
- Added support for mingw compile for opencl/eigen in windows (TODO)
- Allowed human policy to be used for book generation
Fixes for contribution to distributed training
- Fixed bug regarding already-finished games that occasionally caused contribution to katagotraining.org to crash. (thanks to luotianyi for reporting and testing!)
- Linux executables for CUDA and TensorRT have been built now with TCMalloc enabled in the hopes that this will might reduce long-term memory fragmentation during contribute.
Other fixes
- Fixed issue that would result in faulty search behavior with some passing hack options
- Fixed incorrect turn number when converting books to startposes for training.
- Fixed debugging output for illegal game history checking
- Fixed some issues with compilation and scripts on Metal and/or MacOS
Metal Backend and TensorRT Compile Bugfix
This is not the latest release - see a more recent release at v1.16.5!
Changes this Release
This is a quick bugfix release for two issues:
- Fix major bug in v1.16.1 neural net weight calculations that completely broke the Metal backend. (Thanks to @dfannius and @ChinChangYang for quick reporting and fix and testing!)
- Fix issue in detecting versions from TensorRT header files with certain recent TensorRT versions when compiling from source.
The changes in this release compared to v1.16.1 mainly only affect users using the Metal backend or who were building TensorRT from source, so users on v1.16.1 on backends besides Metal and who are not building from source don't need to upgrade. (Users on version v1.16.0, though, should still upgrade to fix issues with potential KataGo crashes).