Dispatch 054
Week ↗

Gemma 3 ships on the edge

It was Thursday, July 10, 2025. I made two commits to the Unicorn Execution Engine this week. Most days I am wrestling with race conditions or memory leaks. Today I was wrestling with documentation and a mod…

Commits
2
Systems
1
Read
3min
2025-07-10 · SIGNAL2 commitsuc-Unicorn-Execu2
Commit signal

Rendered from this day’s 2 commits — no stock art.

It was Thursday, July 10, 2025. I made two commits to the Unicorn Execution Engine this week. Most days I am wrestling with race conditions or memory leaks. Today I was wrestling with documentation and a model that refused to behave until I stopped treating it like a toy.

The headline is simple: I got the production Gemma 3 27B server running on the NPU and iGPU framework. It is not just a demo anymore. It is a real server.

I started by updating the README. The project status was lying. It said one thing, the code did another. I fixed the README to match reality. That is easy. The hard part came next. I had to wire the 27B parameter model into the hardware acceleration stack.

The headline is simple: I got the production Gemma 3 27B server running on the NPU and iGPU framework.

The files changed were eight across two commits. Eight files for one goal. I rewrote the quantization engine docs because the old ones were useless. I updated the project handoff summary so the next poor soul knows what to expect. Then I got to the code.

I built a new loader for the quantized Gemma 27B model. It talks to both the NPU and the integrated GPU. The attention kernel needed a real rewrite. The old one was fast but wrong. The new one is slower but correct. I added a Vulkan FFN compute engine to handle the feed-forward layers. It is a bit of a beast.

The server itself got a fresh start. I wrote a new entry point that actually works. It is not a hack. It is a proper server. I spent hours staring at the logs. The model would crash if I did not feed it data in the right order. The NPU wants things in chunks. The iGPU wants things in streams. I had to make them agree.

There was dependency hell. Of course there was. I had to pin a few libraries because the latest versions broke the Vulkan bindings. I do not like pinning. It feels like a trap. But the server needs to run. So I pinned them.

The README got another update. I had to explain how to run the new server. It is not one command anymore. It is two. And you need to set an environment variable. I wrote that down. I also noted the FastAPI issues. They are still there. They are not fixed. But the server runs.

I did not touch the test suite today. The tests were passing before. They are still passing. I trust them. I trust the logs more. The logs showed the model decoding tokens. They showed the NPU working. They showed the iGPU waiting its turn. That is enough for now.

The quantization engine is ready. The attention kernel is real. The server is live. It is not perfect. It is not fast enough yet. But it works. I can run a 27B model on this hardware. That is what I wanted.

I stopped working when the logs stopped scrolling. I had two commits. Two files changed in the README, eight files changed in the code, across two commits. The line counts are high. Two thousand eight hundred forty-six lines added. Forty-four removed. That is a lot of code for one day. It feels heavy. It should.

Also today: I updated the handoff summary so you know what I broke and what I fixed.

The Gemma 3 server is no longer a dream. It is a process on my machine. It is slow. It is loud. It is mine.