All projects
October 2023 – December 2023 Coursework

Pipelined RISC-V CPU

A 32-bit CPU in SystemVerilog, built three times: single-cycle, pipelined, then with a cache

  • 4 people on the team
  • 5-stage pipeline with hazard handling
  • 255 cycles saved by reordering one program's loops

Team coursework in the second year of my degree at Imperial: a RISC-V CPU in SystemVerilog implementing most of the RV32I base instruction set, delivered in three versions: single-cycle, pipelined with hazard handling, and pipelined with a data cache. I worked with Dima Askarov (control unit and hazard unit), Sam Barber (ALU and data cache) and Meric Song (verification testbench). The design was simulated with Verilator and run on the Vbuddy development board.

Block diagram of the single-cycle CPU: control unit, PC register, PC adder, ALU, register file, result selector and data memory, with the signals between them.
The single-cycle CPU. Click to open it full size.

Program counter and memory

I designed the program counter logic. The PC register takes the next PC value from a multiplexer that I merged into it: either PC + 4 or a branch or jump target, selected by PCSrc from the control unit. The PC adder computes that target as PC + ImmExt, and for JALR adds the register offset instead and clears the bottom two bits so the address stays word-aligned. Dima and I integrated it with the control unit once that was complete.

I also wrote the first version of the data memory and the result multiplexer that selects what is written back to the register file. Supporting jumps meant widening ResultSrc to three bits, so the write-back value can be the ALU result, loaded data, PC + 4 (the return address), an immediate, or PC + ImmExt. For the instruction memory, I made the program file a runtime argument loaded with $readmemh, so any test program could be run without rebuilding the design.

Pipelining

For the pipelined CPU I implemented the four pipeline registers between the fetch, decode, execute, memory and write-back stages, and connected them in the top-level module. The fetch and decode registers support stall and flush signals driven by Dima’s hazard unit: a flush clears the register, and a stall holds the previous values so the correct PC is not lost when a predicted jump turns out to be wrong. By the execute and memory registers every hazard has been resolved, so those two pass values through on each clock edge.

Simulation waveform of the pipelined CPU showing the PC in each stage, the stall and flush signals, forwarding selects and register values.
The pipelined CPU in simulation. The stall and flush signals go high where an instruction needs a value that is still being loaded from memory; the value is then forwarded to the ALU.

F1 lights program

I wrote the F1 starting-lights test program: the lights come on one at a time, then switch off together after a pseudorandom delay. I prototyped it in C++ and then implemented it in RISC-V assembly. The delay uses a 4-bit linear feedback shift register (1 + X³ + X⁴), placed in a subroutine so the program also exercises JAL and JALR.

On the pipelined CPU I reordered the program’s loops around the branch prediction. The hazard unit assumes backward jumps are taken and stalls for one cycle when that is wrong, so arranging the loops to avoid mispredictions saves one cycle per iteration: 255 cycles over a full count to 255.

Build and test tooling

I wrote the build script that verilates the full design and runs it with a given program, fixed the Makefile that assembles .s files into hex, and wrote the top-level testbench that drives the Vbuddy, outputting register a0 to the LED bar and showing the program name on the screen. The same testbench was used to test both the single-cycle and pipelined CPUs.

The CPU running a reference program that plots the distribution of a Gaussian data set on the Vbuddy.