Dispatch 186
Week ↗

Switching the voice engine to faster-whisper

Saturday morning started with the usual quiet hum of maintenance. I had two small chores lined up in uc-cloud, just keeping the Unicorn-Brigade submodule pointed at the right commits. It is the kind of work…

Commits
3
Systems
2
Read
2min
Product capture

Unicorn Commander — models.

Saturday morning started with the usual quiet hum of maintenance. I had two small chores lined up in uc-cloud, just keeping the Unicorn-Brigade submodule pointed at the right commits. It is the kind of work that feels like dusting the shelves, necessary but not exactly thrilling. I updated the submodule once to catch health-check.sh, and then again to pull in the current main branch. These are the invisible gears that keep the rest of the machine from grinding itself apart, and I got them done before the coffee even cooled.

The real story of the day, however, happened in unicorn-amanuensis. I have been wrestling with the speech-to-text backend for a while now, and the old setup was showing its age. The latency was creeping up, and the resource usage was getting heavy. I decided it was finally time to make the switch to faster-whisper using the CTranslate2 backend for production use. This was not a minor tweak. It was a foundational change that touched three key files: the NPU optimization layer, the OpenAI wrapper for the NPU, and the server-side WhisperX implementation.

The diff was substantial. I deleted 183 lines and added 87. It felt like shedding a heavy coat. The old code was bloated with workarounds for hardware acceleration that no longer made sense. The new backend is leaner and significantly faster. I spent the bulk of the day tracing through the optimization logic, making sure the NPU integration still held up under pressure. It was a bit of a puzzle, untangling the dependencies between the ONNX models and the WhisperX server. But once the pieces clicked into place, the system felt lighter.

I have been wrestling with the speech-to-text backend for a while now, and the old setup was showing its age.

There was a moment of dry humor when I realized how much dead code I was carrying. Those 183 lines deleted were mostly legacy support for hardware configurations we do not use anymore. It was a good reminder that software, like any good wardrobe, needs regular purging. The new setup is not just faster; it is cleaner. The CTranslate2 backend handles the heavy lifting with less fuss, and the NPU optimizations are now focused on what actually matters.

I also took a moment to verify that the health checks in Unicorn-Brigade were still passing. They were. A small victory, but one that matters when you are pushing changes to production. The submodule updates were just the formalities, ensuring that the latest fixes were available if needed. But the heart of the day was the backend switch.

It is easy to get bogged down in the details of line counts and file changes, but the real metric is how the system feels when you run it. The new whisperx implementation is snappier. The latency is gone. The resource usage is down. It is a quiet kind of win, the kind that does not make headlines but makes the whole stack happier.

I wrapped up the commits with the usual chore updates, but my mind was already on the next optimization. The foundation is solid now, and the speed is there. It is a good day to be building. The voice engine is ready for whatever comes next.

Also in the frame

Real product captures — click any to enlarge.

uc cloud
uc cloud
shared memory