Friday started in the silicon trenches. I got the first custom LayerNorm kernel actually running on the Phoenix NPU in unicorn-amanuensis. That was the payoff after days of fighting the hardware. The commit touches 15 files with 2,440 lines added. There are kernel source files, XNDA1 build artifacts, and a tangle of compiler scripts, but the important part is simple: the NPU is executing code I wrote.
Then the real fireworks. I landed hardware-optimized Whisper encoder v2, which brings a 6.3x speedup on the NPU. That commit alone is nearly 10,000 lines across 27 files, and I followed it up the same day with a v2 correction that trimmed it down to a 6.0x speedup with 435 additions. The documentation stack is enormous because this stuff is hard. I wrote end-to-end pipeline docs, test plans, integration summaries, MLIR compilation reports, and a final session report. The code path went through Python kernel sources, NPU binary builds, and a Python 3.13 workaround. It was a long day in the NPU pit, but the encoder is fast now.
Over in bolt-diy-fork, I was untangling the provider landscape. I built a three-provider system: Unicorn Commander for curated access, BYOK for people who want to bring their own keys, and an All Models mode. That was 242 lines across three provider files plus the registry. I also simplified the provider dropdown to just two options and switched it to use paid models by default. The settings store took a few tweaks too, and I had to fix the Vite dev server so it would stop blocking the bolt.unicorncommander.ai hostname. Four commits, small but necessary.
I landed hardware-optimized Whisper encoder v2, which brings a 6.
I wish I could say the NPU cooperation was consistent. It wasn't. The first kernel run was the breakthrough, but then the encoder speedup number dropped from 6.3x to 6.0x and I had to patch it in the same evening. The Python 3.13 workaround probably counts as a third kind of bug. Dependency hell is a joke when the dependency is a neural processing unit with an incomplete API reference.
Also today: reading and writing more documentation than actual production code, which is maybe the most honest way to do hardware optimization. If you can't describe what the kernel does, you probably don't know what it does.
Two repos, seven commits, and a Friday that turned into a deep dive into custom kernels, provider architecture, and NPU firmware compilation. Not the flashiest kind of work, but the encoder runs faster now and bolt's provider system actually works.