
Autonomous drone racing for Anduril and DCL's AI Grand Prix. Flew a 17-gate course using only a camera and IMU at 30 Hz, finishing 25th out of more than 3,000 entries with a 35.37 second lap.
This summer, one of my teammates and I competed in Anduril and DCL's AI Grand Prix, an autonomous drone racing competition. We ended up finishing 25th out of more than 3,000 entries with a final time of 35.37 seconds.
This was probably one of the most fun technical projects I have worked on. We came into it with basically no background in autonomous drone racing, very limited compute, and seven weeks to figure everything out.
There were two virtual qualifiers, VQ1 and VQ2. VQ1 gave us absolute position, which made the problem much easier. VQ2 removed that. Now the drone had to fly a 17 gate course using only a camera and IMU while running everything live at 30 Hz.
That changed basically the entire problem.
The first big question was localization. If the drone does not know where it is, having a good controller does not really matter. We needed to detect the gates, estimate their geometry accurately enough to recover pose, fuse that with the IMU, and keep the estimate stable while the drone was moving at race speed.
I went through close to 30 versions of the vision stack.
At different points I was training YOLO models, custom corner detection networks, crop trackers, and different versions of the localization pipeline. One of the biggest improvements came from looking specifically at where our models sucked. I spent around six hours manually clicking gate corners on images where the detector was failing, especially blurry frames, close approaches, partial occlusions, and weird viewing angles. We then combined those high quality manual labels with much larger amounts of automatically generated simulator data.
By the end we had around 39,000 labeled frames. The final corner detector could localize gate corners to subpixel accuracy, and our localization became good enough that some of the other top teams were asking us how we were doing it.
Then there was the control problem.
There were a million possible ways to approach it. We could train everything with PPO. We could build a simplified simulator and learn a residual policy there. We could imitate human flights. We could hand tune a controller. We could try to model the dynamics ourselves. We tried a lot of them.
I pulled out my old FPV controller and personally flew the course hundreds of times. Some of that data was useful for training, but honestly a lot of the value was just learning what a fast line actually looked like. I could see where the drone needed to accelerate, where it needed to start rotating early, which gates were actually hard, and which problems were coming from perception instead of control.
Pure imitation learning did not end up being enough. The drone basically inherited how cautious the human pilot was. Pure end to end RL also had transfer problems.
What eventually worked was a hybrid system. We had a reference controller that could already finish the course, optimized the racing line offline, and then trained PPO residual policies to make smaller corrections in the sections where learned control actually helped.
One of the biggest constraints through all of this was compute.
A real simulator flight took around 40 seconds. That is completely terrible if you want to do millions of RL steps. We also did not have a lab full of GPUs. Most of the competition I was working with one shared RTX 5090 and my laptop 5070.
So I trained a dynamics world model from our flight data.
Instead of making the model predict the entire simulator from scratch, we modeled the physics we understood and trained a small neural network ensemble to predict the remaining error. Eventually we could run around 300,000 simulated steps per second on one GPU. That made it possible to test huge numbers of controller changes and PPO policies offline before wasting another real simulator flight.
The world model was definitely not perfect. Over longer horizons the prediction error got large enough that an optimizer could find ways to exploit it. We learned pretty quickly that it was useful for short horizon search and ranking candidates, not as a replacement for the actual simulator.
Compute management itself became part of the competition. I basically tried to keep every GPU I could get access to at 100 percent utilization. At one point I had six laptops connected and doing different jobs. One would be training a vision model, another running post-training on YOLO, another fitting dynamics, another processing flight logs. If a friend had a GPU sitting around, I wanted it.
We collected over 200 GB of data and ran a few thousand live flights. Most of them failed.
That was another thing I learned from this project. Making one good flight happen is very different from building something that repeatedly works. We would fix one gate, suddenly start reaching three gates farther, and immediately discover an entirely new failure mode. Sometimes a change looked amazing and then failed six times in a row. Sometimes the simulator itself was lagging and we spent hours thinking our policy had gotten worse.
Eventually we got pretty obsessive about evaluation. We kept a frozen champion, tested changes against it, replayed failures, tracked exactly where runs died, and stopped trusting a change just because one flight looked good.
The final jump happened really fast. Our first complete VQ2 lap was 40.01 seconds. Fixing the gate map brought that to 39.79. Better multigate localization and a better racing line brought it to 37.62. Adding routed residual policies got us to 36.83. Then a bunch of speed tuning and work on the later part of the course brought the final time down to 35.37 seconds.
A lot of this project was just iteration. Train something, fly it, crash, figure out why, collect the failure, train again, and repeat. I worked around the clock for a lot of the competition and tried basically every idea I thought had a chance of making the drone faster.
We started with almost no domain knowledge and ended up learning a ridiculous amount about reinforcement learning, PPO, behavior cloning, world models, system identification, visual localization, EKFs, real-time inference, control, model fine-tuning, and just how difficult it is to make a drone fly fast when the only thing telling it where it is is a blurry camera.
It was extremely scrappy, which was probably my favorite part. We did not have the compute or number of people that some of the labs we were competing against had. We just had to keep finding ways around that.
Seven weeks later, we had a drone flying a 17 gate course autonomously in 35.37 seconds.

















