Game of Life on FPGA
Conway's Game of Life in custom hardware, with a hand-gesture interface for drawing the grid
- 722 clock cycles per 1280×720 generation
- ~200k generations per second in theory, at about 150 MHz
- 60 fps on screen, against about 8 fps for our C++ version
The brief for Imperial’s second-year Electronics Design Project was an educational tool that visualises a mathematical function, with hardware acceleration that beats a CPU-only version. Our team of six (Benjamin De Vos, Anderson Lo, Ajay Samaranayake, Ching Bon Tang, Joshua To and me) chose Conway’s Game of Life on a PYNQ-Z1, which pairs an ARM processor with FPGA fabric. The system shows a 1280 by 720 grid over HDMI, one cell per pixel, and users draw the starting pattern in the air with hand gestures.
Sizing the problem on a CPU
I wrote the Python implementations we used as a baseline. The first tracks only live cells, so its cost scales with the number of live cells rather than the size of the grid. I then parallelised it across 16 threads, first by splitting the grid into bands of rows and then into square blocks. At 1920 by 1200 the block split was more than twice as fast as the row split, because each thread’s neighbour lookups stay within a small region of memory.
The best software result was still far off our target. Our C++ version computed about 8 generations per second at 1280 by 720, so a 60 Hz display needed a speedup of at least eight times, which set the case for custom hardware.
Next-state logic
I wrote the core of the hardware: a module that applies the Game of Life rule to one cell from its eight neighbours, and a Verilog generate loop that instantiates 1280 of them side by side, so a whole row’s next state is computed at once. Around it, Ajay built the row counter and the line buffer that feeds each row with the rows above and below it, and Ching Bon built the logic that swaps the two block RAMs between “current” and “next” each generation, which avoids copying the grid. With one row per clock cycle, a full generation takes 722 cycles. At the fabric’s clock of roughly 150 MHz that is about 200,000 generations per second, although the display itself refreshes at 60 Hz.
Loading the starting grid
I wrote the logic that loads a user’s pattern into memory. The processor passes data to the fabric through 32-bit registers, so each 1280-cell row is split into forty 32-bit words, written into forty registers, and concatenated in hardware into one 1280-bit word that fills a whole row of block RAM in a single cycle. The full grid loads in 720 cycles, after which a flag starts the evolution.
Hand-gesture interface
I built most of the gesture interface, which runs on a laptop. It tracks hand landmarks with MediaPipe through OpenCV: with the drawing gesture, each point the user traces sets a 3 by 3 block of cells, and a thumbs-up sends the finished grid to the board. An open hand pauses the simulation on the current generation, and closing it resumes. The grid is sent to the board row by row over a TCP socket, and I wrote the Jupyter notebook on the board that receives each row, writes it into the registers and steps the generations. In user testing the interface responded in about 30 ms on average.