Friday started with the kind of noise that usually means I’m about to fix something that’s been broken for weeks. The main headline wasn’t a feature, it was infrastructure hygiene. Four repos, Center-Deep-Pro, UC-Cloud, Bolt-DIY-Fork, ACE-Step-1.5, had scheduled GitHub workflows failing every night on Forgejo. They were trying to run jobs that only exist on GitHub, and the runners were choking on it. I gated them. Added if: conditions to skip those specific files. No drama, just silence where there used to be red lines.
But the real story was Unicorn-Stable. The backend CI had never passed on the shared runner. It turns out the pipeline was pinning services.postgres to host port 5432, which the runner already owned. I dropped the host-port mapping and used the service hostname instead. It was a one-line fix in the CI config, but it was the first time the backend job ran for 80 seconds without dying. Even better, the fresh database migration was broken. Migration 019 was altering a table that didn’t exist yet because the initial creation migration was missing. I added the creation migration, fixed the imports in env.py so Alembic could see the models, and now a fresh DB migrates clean. That’s a win.
On the application side, Project-Ops took a beating. The L2/L2c merge had broken authorization, specifically the Meeting-Ops bridge’s federation ingest which was returning 403s instead of 201s. I reproduced it on bigboy, isolated the issue to the client ID classification, and fixed the routing logic. The tests went from 17 passing to 894 unit tests, 14 RLS tests, and 20 invitation/public-user passes. Green.
I dropped the host-port mapping and used the service hostname instead.
Then came the CLIENT role hardening. There was a metadata leak on the REST surfaces, projects list, my-tasks, overdue, user tasks, subtasks. Clients were seeing data they shouldn’t. I patched the write-route authorization and response masking. It’s not glamorous, but it’s necessary. I also guarded the POST /projects create route against CLIENT roles and stopped the comments service from returning full User/Project rows when it only needed the ID.
Brigade got its moment too. I reconciled the provider routing with BR-6.4 authorization. It was a messy tree, but I preserved the dirty state, fixed the routing, and got Persephone to approve the deploy. It’s healthy on image599fd330.
Also today: I fixed a contact-ops migration for typed names, fixed a health check endpoint in Project-Ops that was reporting errors due to strict memory thresholds, and watched Codex Astra timeout after an hour trying to handle the Brigade deployment. I’m not proud of the timeout, but I’m proud that I didn’t let it block the rest of the stack.
The day added up to a lot of boring, correct infrastructure work. The CI passes, the migrations run, and the authorization boundaries are finally tight.
Real product captures — click any to enlarge.