Back to Projects

Feb 2026 - Mar 2026

Real-Time FPGA-Accelerated Crypto Trading System

This was my second-year Information Processing project at Imperial, done with six other people. We came up with an idea to take live crypto data, predict where the price was going over a very short window, and actually trade on those predictions. The catch was that the model had to run on FPGA, not just in Python.

We ended up with two PYNQ-Z1 boards that never talk to each other directly. One is the trading node: it sits on a Binance WebSocket, computes microstructure features every 200ms, and runs a linear regression on the programmable logic to predict the price differential over the next two seconds. Those predictions turn into paper BUY/SELL orders. The other board is a voice controller, so you can switch assets and features by talking to it. Shared state lives in a Flask API on AWS. If the cloud drops, both boards keep running locally and resync when it comes back.

What stuck with me is what we actually had to improve to make it fast, and how often the first bottleneck we pointed at was the wrong one.

The model is incremental linear regression: rank-one updates to AtA and Atb, which maps pretty cleanly onto pipelined FPGA logic. We also kept two software paths, one naive and one optimised, and ran all three on every batch. That sounds like overkill until you watch it catch an overflow or a normalisation bug before it hits a live trade. On the kernel itself, hardware accumulation for 4000 samples took 3.49ms against 30.4ms for the better software version, and peak multiplies were about 68x faster. We spent a while assuming compute was the thing to squeeze. It wasn't. DMA was. Getting samples on and off the fabric mattered more than shaving another cycle off the multiply.

We talked to people in Imperial's Financial Signal Processing lab about whether a fancier model was worth it. The honest answer was that a lot of practical HFT edge is latency, not model complexity. A neural net was a non-starter on this board anyway. There wasn't enough BRAM to hold the matmuls, which is a very specific way of being told to keep it simple.

Paper trading grew a $1000 book to $2126 at 20x leverage, with 57.6% directional accuracy. I'm putting that here because it is what the dashboard shows, not because I think we built a money printer. It is a lab system on a PYNQ-Z1. The number I care about more is that the hardware path was actually faster, and that we could explain why.

Highlights

  • Running naive software, optimised software, and hardware on the same batch. Agreement between them caught overflow bugs we would have otherwise blamed on the market.
  • The kernel was fast. Moving data onto the FPGA was not. DMA, not the multiply, was the thing we kept coming back to.
  • 77% of the DSP slices on a small PYNQ-Z1, 16-bit fixed-point into a 64-bit accumulator. The board was genuinely full.

Reflections

The thing that stuck with me is how often the obvious bottleneck is the wrong one. We went in thinking we needed a cleverer model, or a tighter inner loop, and the work that actually moved the needle was plumbing, fixed-point, and checking the hardware against two software implementations.

If we had more time, the useful next steps would be moving the pseudo-inverse onto the fabric with Cholesky or QR, and putting a second trading node on a correlated asset so the multi-node setup actually gets used.

SystemVerilogFPGAPythonNumPyAWS EC2Flask