The headline today was simple. I got Qwen3-Embedding-0.6B running on the Unicorn NPU. It is not a massive refactor. It is not a rewrite of the core engine. It is a focused spike to see if the hardware can handle the specific demands of embedding generation without choking on memory or latency. The project landed in the Unicorn Execution Engine repository with a single commit that added nearly four thousand lines of code. That number looks big, but most of it is scaffolding. The real work is in the kernels.
I started with the assumption that the standard CPU baseline was too slow for production workloads. The benchmarks in the new directory confirm that. I wrote a baseline script to measure the CPU performance first. That gave me a hard number to beat. If the NPU cannot beat that number by a significant margin, the whole exercise is just academic. The results are in the JSON file. The CPU took a certain amount of time. The NPU path needs to be faster.
The core of the work is in the kernels directory. I implemented an embedding lookup kernel in MLIR. This is the part where we fetch the vectors. It sounds trivial until you try to make it fast. Then you realize that memory access patterns matter more than you think. I also wrote a matrix multiplication kernel. Embeddings are not just lookups. They often involve linear transformations. The NPU handles these operations differently than the CPU. The MLIR code has to talk directly to the hardware instructions.
I started with the assumption that the standard CPU baseline was too slow for production workloads.
I wrapped these kernels in a Python executor. The npu_kernel_executor.py file is the bridge. It takes the high-level requests from the application and translates them into the low-level commands the NPU understands. The matrix_mul_npu.py file handles the heavy lifting for the matrix operations. I kept the interface simple. If the interface is complex, I will use it wrong. Simple interfaces are easier to test.
Testing was the other half of the day. I wrote test_npu_basic.py to verify that the kernels actually produce correct results. Correctness is non-negotiable. Speed is important, but wrong answers are useless. The test suite checks the output against the CPU baseline. If the numbers match within a reasonable tolerance, the kernel is good. If they do not match, I have to dig into the MLIR to see where the logic diverged.
The project structure includes a checklist and a README. These are not fluff. They are necessary for the next person who picks this up. The checklist ensures that no step is skipped. The README explains why we are doing this and how to run the benchmarks. Documentation is part of the code. If it is not documented, it does not exist.
I also added a .gitignore file. We do not want to commit generated results or temporary files. That keeps the repository clean. Clean repositories make it easier to find the actual code.
This is a small step. It is not the final product. But it proves the concept. The NPU can run Qwen3 embeddings. The benchmarks will show if it is fast enough. For now, the code is there. The tests pass. The baseline is set. Tomorrow I look at the numbers. If the speedup is significant, we proceed. If it is not, we rethink the approach. But for today, the NPU is awake and working.
Also today: I updated the project checklist to reflect the new kernel implementations and ensured the benchmark results are captured for comparison.
The day ended with a working prototype and a clear path forward.