Back to Projects

Oct 2025 - Dec 2025

RISC-V RV32I Processor

This is an RV32I processor in SystemVerilog. I did it in a team. We started from a single-cycle core and ended up with a 5-stage pipeline and a real cache. The machine does the full set we were asked for, everything except ECALL, EBREAK, CSR, and FENCE. That's 40-odd instructions. It goes IF, ID, EX, MEM, WB, with hazard detection, forwarding, and a flush when a branch guess is wrong. On top of that we put a 4 KiB 2-way set-associative write-back cache. Fifteen assembly tests, all passing.

My job was execution and memory: the ALU, the register file, the Execute stage, the forwarding unit, and the data cache controller. The other teammates handled the pipeline registers, the control unit, and a lot of the cache testing.

The ALU bit I actually care about is SRA. Arithmetic right shift has to keep the sign. SystemVerilog's >>> does that if you treat the value as signed first. Forget $signed() and you silently get a logical shift. I didn't want to stall on every RAW hazard, so I put forwarding in Execute instead. Three-way muxes: take the ALU result from MEM if that's the newest write, otherwise writeback, otherwise the register file. You have to check RegWrite and ignore x0, or a disabled instruction forwards garbage.

The cache is write-back, write-allocate, 2-way, 128 sets. The controller is a five-state FSM. LRU is a single flip-bit per set. The bug that taught me the most was refill data going into the wrong set. The pipeline was stalled, the PC wasn't moving, but the address from Execute was still twitching because forwarding was still doing its thing. The cache was writing the incoming line to whatever index happened to be on the bus. The fix was a shadow address: latch it when you leave IDLE, and don't look at the live bus again until you're done.

We also tried a BTB as the stretch goal. It broke tests, so we pulled it. Shipping the core that works was the point.

Reflections

If I did it again I'd spend the extra time on a 2-bit predictor or a return address stack, and a write buffer so dirty evicts don't stall the core. Dual-issue is the fun one, but that's a different project.

Correct combinational logic is not the same as a working pipeline. Even stalled, signals still move, which is how refill data ended up in the wrong set until we latched a shadow address.

Seeing all 15 tests pass with forwarding and the cache actually behaving was the bit that made it feel done.

If I had more time, I would consider implementing the CPU onto an FPGA!

SystemVerilogRISC-VFPGA