Pipelined RISC-V CPU
A 32-bit CPU in SystemVerilog, built three times: single-cycle, pipelined, then with a cache
- 4 people on the team
- 5-stage pipeline with hazard handling
- 255 cycles saved by reordering one program's loops
Team coursework in the second year of my degree at Imperial: a RISC-V CPU in SystemVerilog implementing most of the RV32I base instruction set, delivered in three versions: single-cycle, pipelined with hazard handling, and pipelined with a data cache. I worked with Dima Askarov (control unit and hazard unit), Sam Barber (ALU and data cache) and Meric Song (verification testbench). The design was simulated with Verilator and run on the Vbuddy development board.
Program counter and memory
I designed the program counter logic. The PC register takes the next PC value from a
multiplexer that I merged into it: either PC + 4 or a branch or jump target, selected by
PCSrc from the control unit. The PC adder computes that target as PC + ImmExt, and for
JALR adds the register offset instead and clears the bottom two bits so the address stays
word-aligned. Dima and I integrated it with the control unit once that was complete.
I also wrote the first version of the data memory and the result multiplexer that selects
what is written back to the register file. Supporting jumps meant widening ResultSrc to
three bits, so the write-back value can be the ALU result, loaded data, PC + 4 (the return
address), an immediate, or PC + ImmExt. For the instruction memory, I made the program
file a runtime argument loaded with $readmemh, so any test program could be run without
rebuilding the design.
Pipelining
For the pipelined CPU I implemented the four pipeline registers between the fetch, decode, execute, memory and write-back stages, and connected them in the top-level module. The fetch and decode registers support stall and flush signals driven by Dima’s hazard unit: a flush clears the register, and a stall holds the previous values so the correct PC is not lost when a predicted jump turns out to be wrong. By the execute and memory registers every hazard has been resolved, so those two pass values through on each clock edge.
F1 lights program
I wrote the F1 starting-lights test program: the lights come on one at a time, then switch
off together after a pseudorandom delay. I prototyped it in C++ and then implemented it in
RISC-V assembly. The delay uses a 4-bit linear feedback shift register (1 + X³ + X⁴),
placed in a subroutine so the program also exercises JAL and JALR.
On the pipelined CPU I reordered the program’s loops around the branch prediction. The hazard unit assumes backward jumps are taken and stalls for one cycle when that is wrong, so arranging the loops to avoid mispredictions saves one cycle per iteration: 255 cycles over a full count to 255.
Build and test tooling
I wrote the build script that verilates the full design and runs it with a given program,
fixed the Makefile that assembles .s files into hex, and wrote the top-level testbench
that drives the Vbuddy, outputting register a0 to the LED bar and showing the program name
on the screen. The same testbench was used to test both the single-cycle and pipelined CPUs.