Monday started with a specific technical question that I needed to answer before the week could move forward: can we run Whisper transcription on Intel integrated graphics without a discrete GPU, and make it fast enough to actually ship? I spent the entire day in Unicorn Amanuensis chasing that answer. The short answer is yes, with some serious caveats.
The solution involved integrating whisper.cpp for Intel iGPU via SYCL, which is Intel's OpenCL-like API for their GPUs. This wasn't a small patch. The codebase absorbed the entire whisper-cpp-igpu submodule plus a bunch of custom server wrappers. One commit touched 1,130 files, and another added 10,420 lines of Dockerfiles, converters, and server code. It was a massive injection of new code, but it was necessary to build the foundation.
The performance numbers are where things get interesting. The commit messages claim 80x realtime on the SYCL implementation, and a separate production server hit 7 to 11x realtime. I’m going to trust the commit messages on this, even though 80x on an iGPU sounds almost too good to be true. Either way, even the conservative 7x figure changes the cost equation significantly. It means we can run this stack on a $200 NUC instead of a $2000 GPU.
I’m going to trust the commit messages on this, even though 80x on an iGPU sounds almost too good to be true.
I also built out the full production deployment story on Monday. Dockerfiles for every variant now exist, including lightweight, IGPU, INT8, OpenVINO, and comprehensive versions. We have Docker Compose stacks for UC1 Pro and production, along with a device manager that handles multi-GPU systems intelligently. The INT8 servers are worth calling out specifically. Quantized models run at 8-bit precision instead of 16-bit, which cuts memory usage roughly in half with minimal quality loss. Combined with the SYCL backend, that makes the iGPU path viable on hardware that doesn’t have a lot of VRAM. The new device_manager.py ensures that when you have both an iGPU and a discrete GPU, the system routes transcription intelligently rather than just throwing everything at the biggest GPU.
By Tuesday, the focus shifted to making that iGPU path actually work for users. I got the Intel integrated GPU SYCL transcription working with real-time streaming on the UC1 Pro deploy. That is the hardware path people actually want to run on. It was a bigger lift than I expected. Across seven files, I made 1,435 additions and 368 deletions. The main server file got updated, the Dockerfile for the ultimate iGPU path got rewritten, and the web interface got adjusted on the frontend side too. The docker compose file and deployment docs were updated to match, and I bumped requirements.txt with whatever the new dependencies were. The whisperx Dockerfile was the key piece, and the static frontend got touched alongside it.
I also took some time to clean up the visual presentation. I added theme screenshots to the docs and updated the unicorn-amanuensis README to show off the dark, light, and magic themes. Four files, thirteen additions and three deletions. It is just visual cleanup, but it makes the project look polished instead of rough around the edges. The iGPU path being live means people without dedicated GPUs can actually run transcription at real-time speeds. No NVIDIA required. Intel's SYCL support was the missing piece, and now it is there.
Wednesday was a different kind of day. I spent the morning building the infrastructure layer for tool servers in center-deep-pro. This is basically giving the system the ability to run specialized research and search tasks as dedicated services rather than inline functions. I added 17 files covering the Docker setup, the OpenWebUI integration hooks, and the actual tool implementations for academic research, deep search, and report generation. It was a big chunk of work, almost 3,000 lines added across configuration files, server scripts, and the Python tool modules themselves. The Gemma 3 setup was part of this too, tying the tool server into the broader model infrastructure.
The uc-cloud repo took a different kind of hit that afternoon. I pushed a massive update there, 74,000 lines across 283 files, which was mostly the Authentik SSO integration work. I am talking full setup guides, OIDC configuration docs, backend developer starter guides, and all the deployment documentation for tying Center Deep into Authentik for single sign-on. It was a big lift, but it was necessary plumbing. I also pushed a separate documentation commit for UC-1 Pro v1.0 that cleaned up the README and features architecture docs to keep things consistent.
Then there was unicorncommander.com, which got its entire website pushed live in one commit. That was 1,137 files and about 146,000 lines, including all the built assets, images, the Colonel background artwork, and the full site structure. This was the marketing site that ties the product narrative together, and having it in the repo means the whole thing ships with the code.
Wednesday was one of those days where three completely different kinds of work happened at once. Infrastructure for a new capability, a massive documentation and SSO integration push, and a full website deployment. Each one alone would be a solid day. Doing all three in 5 commits is just how it goes when you are the one who has to make it happen. The tool servers are the most interesting part technically, because they open up the system for real research workflows instead of just chat. But honestly, having the website in the repo and the SSO wired up are what make the whole thing feel less like a prototype and more like a product.
This week was about closing the gap between experimental code and deployable reality. We moved from asking if something could work on integrated graphics to shipping the full stack to do it. We added the infrastructure for serious research tools and finally got the marketing site into version control. It was a lot of lines of code, a lot of Docker files, and a lot of documentation. But the result is a system that is actually usable on affordable hardware, with the plumbing to support real users. That is a good week.
Real product captures — click any to enlarge.