May 2026 - Jun 2026
Reinforcement Learning Smart Grid Controller
This was my 2nd year EIE summer project, and honestly one of my favourite projects I've worked on recently.
The hardware was a real lab smart grid that we worked with. Everything sat on a shared 10 V DC bus: a PV emulator for solar, a supercapacitor for storage, a PSU standing in for the grid, and an LED load that was meant to look like a house. Picos sat on the bus so we could actually read voltages and currents instead of guessing. Azure sent a new tick every five seconds with prices, demand, and a few deferrable tasks, sixty ticks to a simulated day, and the whole thing had to keep up live.
We were given this grid and a specific set of requirements to fulfil on it:
- Keep instant demand served. Lights on at all times.
- Finish the deferrable tasks inside their windows.
- Buy and sell energy so the day is as cheap as possible, ideally profitable.
- Do it live, on the hardware, inside that five second tick.
My job was the controller. Everyone else could wire the bus perfectly and we'd still lose if the policy was stupid, so it mattered a lot. Every other team used linear programming. We were the only ones who actually trained an AI model for it.
Getting the model onto the bench was its own problem. There were a few controllers plus the inference machine, and they all needed to talk without eating the five second window. We put everything on a common IPv4 address so the Picos, the controllers, and the model were just hitting the same place instead of a mess of point-to-point links. Inference would read the bus, decide, and publish the capacitor command and the load before the tick was over. When that path was clean, the loop felt obvious. When it wasn't, nothing else mattered.
I didn't know RL before this. I sat down, learned PPO properly, built a Gymnasium environment in PyTorch, and just started iterating. Collect a day of ticks, train, look at what the agent actually did, realise I'd rewarded the wrong thing, change the env, do it again. Four prototypes later, my model beat every other team.
- Prototype 1: I learned the agent will do exactly what you pay it for. I made unmet demand expensive, so it imported everything and lost money every day. Next time, don't let reliability crowd out the actual goal.
- Prototype 2: I got better at putting the physics in the environment instead of hoping the network would discover it. Demand got served automatically and we finally turned a profit, but it was still cheating with hardcoded arbitrage. Next time, make it learn the money itself.
- Prototype 3: I got better at writing a reward that is the thing I actually care about. Strip the heuristics, two actions, profit in cents. That's when it started looking like it understood prices. Next time, stop letting the simulator be nicer than the lab.
- Prototype 4: I got better at not fooling myself. Match the hardware, kill the lookahead I wouldn't have on the bench, and make missing a deferrable task accurately hurt the reward. That's the policy we put on the bus.
On the day, my model outperformed every other team.
Highlights
- Watching the model actually turn a profit. Not only according to my calculations, it was making money on the grid.
- Getting the hardware and the software talking, running live inference, and it all working perfectly after weeks of effort and many iterations.
- The demo. We were all panicking trying to get everything up, high stress, and then it ran perfectly. That high is hard to describe.
Reflections
If I did it again I'd spend more time actually tinkering with the model, and I wouldn't make the same reward mistakes twice.
Also, integrating with hardware is way harder and way more error prone than I thought. We had to change the IPs five times across five devices just to get a reliable connection between the model and the Picos.
Oh, and I couldn't have done any of this without my friends.
