Skip to content

Latest commit

 

History

220 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Self learning Gomoku AI

Patchnotes

v.0.1.2 -> v.0.1.3
Fixed random bugs.

TODO:

  • Rewrite Python codebase: Goal selfplay / selftrain loop
  • Implement RIF ruleset in both Python and C++

Basics

In Gomoku black starts. In code:
Black is 0 and -1.0 is black winning evaluation
White is 1 and 1.0 is white winning evaluation
(Draw is 2 and 0.0 is drawn evaluation)

Algorithm

The general idea is to use a variation of MCTS, which uses a Neural Network with policy and value-head.
The Neural Networks parameters are improved over time via selfplay. The model is retrained on better moves calculated by the MCTS. If the new model wins at least 55% of games against the old one keep it.

This loop is not yet implemented

Multithreading

Multithreading is implemented via a worker pool.
When creating a batcher it calculates the amount of GCP (Gamestate conversion processes), and SIM (Simulation) workers based on provided hyperparameters. These workers are all the workers ever created by that Batcher, we dynamically reduce the amount of workers to start when needed but we never create new ones.
Multithreading is implemented on a Batcher level since every environment is independent of all other, which makes multithreading fairly efficient due to not needing any mutexes or atomics. (Every environment will only ever be worked on by one thread).

Hyperparameters for threading are:

  • PerThreadSimulations: How many threads for simulating (MCTS Tree).
  • PerThreadGamestateConvertions: How many threads for converting nodes to gamestate tensors.

For n environments

  • MaxThreads: How many max threads per type (can result in twice the amount of workers then expected but different types will never be active simultaniously).

Data

Gamestate

Gomoku gamestates for neural network input are encoded as follows:
Tensor of shape [HistoryDepth(HD) + 1, BoardSize, BoardSize].
HD is a even number and at least 2 (HD 2 is no history).
Axis 0 is described as:

[0]: Next players color
[1 - HD/2]: Black stone history
[HD/2+1 - HD]: White stone history

This makes it so we encode the last HD - 2 played moves on the board.
For each step back (index - 1) the last placed stone of given color is removed from that board.

Reading/Writing

Python:
Gamestates can be sliced (render history board states) via the sliceGamestate function in Utilities.py
Datapoints are only ever written by the DatasetCreator.py
C++:
The sliceGamestate function also exists in Utilities.h
Gamestates are created from Nodes via the nodeToGamestate function in Node.cpp
(Reading gamestates should never be necessary except for debugging)

Datapoint

A datapoint contains:

  • History of moves that lead to gamestate
  • The best move at gamestate according to MCTS
  • If current player won

Selfplay

In selfplay the last 500.000 datapoints are stored so we can sample datapoints for retraining.
They are stored in a text file, reading and writing is only recommended via the Storage class.

Naming Conventions

Datasets from human data:
HD(HistoryDepth),[AUG (is augmented)],TS(TrainsplitFraction),Rulesets(All rulesets in the dataset)

Neural Network

Architecture

Filter and Layer count are by user choice

Gamestate
        ↓
   ResNet
        ↓                      ↓
Policy Head    Value Head

ResNet

x = input
residual = x
    Convolution: 3x3 Kernal with padding
    Batchnorm
    ReLU
    Convolution: 3x3 Kernal with padding
    Batchnorm
x + residual
ReLU

Policy Head

Convolution: 1x1 Kernal with 2 Filters
Flatten
Batchnorm
ReLU
Linear: 450 → 225
(Softmax)

Value Head

7 times:
    Convolution: 3x3 without padding
    Batchnorm
    ReLU
Flatten
Linear: Filters → Linear Filters
Batchnorm
ReLU
Linear: Linear Filters → Linear Filters
Batchnorm
ReLU
Linear: Linear Filters → 1
Tanh

Rules

Proper rules not yet implemented.

For now just "freestyle" 5 in a row wins.

Modes

The AlphaGomoku executable can be called with 1 of 3 modes:

  • DUEL: Evaluate 2 models against each other (used in retrain validation).
  • SELFPLAY: Let model play against itself to generate datapoints for retraining.
  • HUMAN: Lets you play against a model with MCTS.

Environment Variables

LOGGING:

  • INFO: Logs non verbose information
  • WARNING: Logs warnings, usually not terminal
  • ERROR: Logs errors, can be terminal
  • FATAL: Logs terminal errors

Parameters

  • help : Print help message.
  • version : Print version.
  • mode : Mode to run the program in (duel, selfplay, human).
  • model : Name of the model.
  • simulations : Number of simulations to run per move.
  • environments : Number of environments to run in parallel.
  • randmoves : Number of random moves to make before starting.
  • seed : Set seed for randmoves.
  • humancolor : Color of the human player (0 = black, 1 = white).
  • stones : Stone skin to use for rendering.
  • board : Board skin to use for rendering.
  • renderenvs : Render the environments.
  • renderanalytics : Render the analytics.
  • renderenvscount : Number of environments to render.
  • datapath : Path to store the data.
  • modelpath : Path to store the models.
  • device : Device to use for inference (cpu, cuda, mps).
  • scalar : Scalar to use for inference (float16, float32).
  • threads : Number of threads to use for batching.
  • batchsize : Batchsize cap for inference.
  • nocache : Previous simulation cache should be deleted before next simulation.
  • policybias : Policy bias to use for MCTS.
  • valuebias : Value bias to use for MCTS.
  • explorationbias : Exploration bias to use for MCTS.
  • modelpath : Where models are stored.
  • datapath : Where datapoints are stored.
  • outputtrees : Trees should be output as graphviz files.
  • outputtreespath : Where graphviz tree outputs are stored.

Italic args can pe specified per model like: --device1 [model1 device] --device2 [model2 device].

Compatibility

Tested with Pytorch 2.1.1 on:
Ubuntu 22.04 - (CUDA.12.1)
MacOS 17.1 - (MPS)

About

! WIP ! AlphaGo based Gomkou algorithm

Resources

Stars

0 stars

Watchers

1 watching

Forks

Contributors

Languages