Monday. I spent the whole day on Unicorn Amanuensis, and it came down to one question: can we run Whisper transcription on Intel integrated graphics without a discrete GPU, and make it fast enough to actually ship?
The answer turned out to be yes, with caveats. I integrated whisper.cpp for Intel iGPU via SYCL, which is Intel's OpenCL-like API for their GPUs. The codebase absorbed the entire whisper-cpp-igpu submodule plus a bunch of custom server wrappers. It's a big addition. One commit touched 1130 files, and another added 10,420 lines of Dockerfiles, converters, and server code.
The performance numbers are where it gets interesting. The commit messages claim 80x realtime on the SYCL implementation, and a separate production server hit 7 to 11x realtime. I'm going to trust the commit messages on this, but 80x on an iGPU sounds almost too good. Either way, even the conservative 7x figure changes the cost equation significantly.
I'm going to trust the commit messages on this, but 80x on an iGPU sounds almost too good.
I also built out the full production deployment story. Dockerfiles for every variant now exist lightweight, IGPU, INT8, OpenVINO, comprehensive. Docker Compose stacks for UC1 Pro and production. A device manager that handles multi-GPU systems intelligently. Server variants with INT8 quantization and diarization support. Deployment documentation is complete.
The INT8 servers are worth calling out. Quantized models run at 8-bit precision instead of 16-bit, which cuts memory usage roughly in half with minimal quality loss. Combined with the SYCL backend, that makes the iGPU path viable on hardware that doesn't have a lot of VRAM.
The other commit that matters: device manager for multi-GPU. When you have both an iGPU and a discrete GPU, you want to route transcription intelligently rather than just throwing everything at the biggest GPU. The new device_manager.py does exactly that.
Also today: updated CLAUDE.md, README, the production roadmap, the IGPU master checklist, and a bunch of build artifacts that will never change.
What it added up to: a transcription stack that can run on a $200 NUC instead of a $2000 GPU, with production deployment docs, Docker Compose files, and enough server variants to cover every use case from bare-bones to full diarization. The 80x claim is bold, but even if the real number is half that, it's a meaningful shift.