Friday and I spent it chasing the NPU. The thing was supposed to be Phase 3 done already. It wasn't. The bias weights were wrong, the ONNX graph had stale nodes, and the hybrid pipeline kept stalling on the projection layer like it didn't know what BERT was. I went in and fixed the bias, rewrote the graph mutator to stop leaving garbage behind, and pushed the whole hybrid TTS chain to Phase 3.
The commits tell the story: 2,025 lines added, 7 removed, spread across 9 files in unicorn-orator. The core changes are in xdna2/kokoro_hybrid_npu_phase3.py and xdna2/modify_onnx_graph.py (bert_projection_weight.npy and bert_projection_bias.npy). I also updated the status doc and the README to reflect that this is actually production-ready now, not \"we think it might work on a good day\".
The hardest part wasn't the code. It was convincing the ONNX runtime to stop complaining about a projection bias that should have been fine. Turns out the graph mutator was leaving a dangling node that only showed up under real NPU load, not in my local CPU tests. Standard pattern. I spent an hour on that one bug and then five minutes on everything else.
I also updated the status doc and the README to reflect that this is actually production-ready now, not \"we think it might work on a good day\".
The Phase 3 pipeline now runs end to end on the NPU with the corrected weights and the cleaned-up graph. That's the arc: NPU acceleration for Kudo's TTS, finally solid enough to ship.
Also today: updating the docs to match what the code actually does.
That's it for the day. One repo, one big push, and the NPU path is finally real.