I build the infrastructure that keeps AI agents honest: it checks what they change, limits what they can touch, and writes down what they did.
Most of my code is written by coding agents that I plan, direct and verify. So the problem I care about is how you trust software that a model wrote. I'm looking for an AI engineering role on agent infrastructure, evals or developer tools.
Six self-hosted tools for running AI agents safely. Each works on its own, and all of them log to Ledger.
I built a labelled set of 71 diffs across Go, Python, JavaScript and TypeScript: 44 with a planted bug and 27 clean. This is Gate with no LLM, using only its static and security checks:
| bug class | cases | caught | |
|---|---|---|---|
| sql injection | 4 | 4 | |
| hardcoded secret | 3 | 3 | |
| path traversal | 2 | 2 | |
| logic, off-by-one | 10 | 0 | |
| authz, concurrency | 8 | 0 | |
| test tampering, deleted tests | 8 | 0 | |
| api misuse, dependency confusion | 9 | 0 | |
| clean diffs wrongly flagged | 27 | 0 |
Precision 1.0, recall 0.21. The deterministic checks never raise a false alarm, but they miss every bug that needs you to understand the intent. Catching those is the LLM review stage's job, and I haven't benchmarked it yet. I published both numbers because the low one tells you more.
Turns a one-paragraph game idea into a playable browser game. You can play three of them below. An LLM splits the idea into about nine units, agents build them in parallel, every unit has to pass verification, and a repair agent fixes any game that still fails the play test.
Three finished builds that passed Prosper's headless play test. These are the merged games exactly as Prosper produced them. Click the game once so it gets your keyboard.
| 8 Aug to 11 Sep | 26 to 29 Sep | |
|---|---|---|
| games started | 223 | 24 |
| units approved | 58% | 87% |
| games that reached the browser play test | 60 | 18 |
| playable on the first test | 4 (22%) | |
| playable at their last test | 35 (58%) | 13 (72%) |
From Prosper's own build database: 246 games and 2,285 units in total. The second period adds the genre starter games, a stricter play test that fails any runtime error, and Jules. Because the play test is stricter, only 22% pass it first time. Regeneration and repair lift that to 72%. The first period didn't record first-test results separately.
When a game fails the play test, Prosper copies it to a repair folder and hands it to Jules, an OpenCode agent running on free models. Jules gets the spec, the exact failure evidence, and a check command that runs the real merge and play test. Two sessions race on separate API keys, and the first to pass wins.
Prosper doesn't take Jules's word for anything. After every round it re-runs the play test itself, restores any pipeline-generated file Jules touched, and copies a fix back only if it passes. Every round is logged with its diff, and when the same kind of repair keeps showing up, I fix that bug in the pipeline instead.
In its first two days, Jules worked on 7 games that failed the play test and made 5 of them playable, over 33 rounds and about 22 hours of agent time. One fix took a single round: a ReferenceError from a variable used before it was initialized. Another took three rounds and four files, after a sprite call that didn't exist crashed the game loop. The two it couldn't fix both came from add-ons calling methods the game didn't have, so I fixed that at the source: each add-on now declares what it needs from the game, and installation fails immediately if the game doesn't provide it.
Single steps ran for over 10 minutes because step time grows with reasoning tokens. I switched the default to low reasoning effort.
29 of the 33 rounds ended on a "no progress" timeout. A 7-minute stall limit killed sessions in the middle of a step. I raised it to 13 minutes, and a hung request now retries in the same session after 4 minutes instead of ending the round.
The repair folder sits inside the repo, so Jules wandered into pipeline code. It's now blocked from reading or searching outside the game.
A new repair session now starts from the previous session's edits, so a restart keeps the progress already made.
Nine hand-built starter games across eight genres, like Star Lancer (space shooter), Gem Cascade (match-3) and Bastion Road (tower defense), come with asset kits and 16 add-ons like bosses, extra waves and relics that scale a game up. Seven genres have a play test that actually plays the game. The shooter test, for example, fails a run where one enemy shot destroys the ship.
To check coverage, I classify 101 realistic game requests. 97 land in a genre (96%) with 0 misclassified, and 81 (80%) have a starter game today. The four it can't place (3D shooter, MMORPG, rhythm game, fighting game) are out of scope for now. Management and story games have no starter yet.
In August I audited the pipeline against four real builds. The verification looked strict but was mostly letting everything through:
The browser play check set ok = False on failure, but the function never returned it. No build could fail.
The "is anything moving" check compared two array objects instead of their contents, so it always saw motion, even in a frozen game.
The Phaser boot check skipped itself because the platform string was phaser+js, not phaser.
Empty units all hashed to the SHA-256 of an empty string, so one unit's failed verdict was reused for every later empty unit.
All four are fixed. Every JavaScript and TypeScript file is now parsed and bundled with esbuild before approval, and the play test fails a game on blank rendering, placeholder art, game-code errors, or a run that ends without taking any input.
A research crawler for agents, so they don't need a search API. It crawls outward from seed URLs, streams each page to the agent as it's extracted, and stores everything in vector memory the agent can search later over MCP.
Literal queries worked. "How does octopus camouflage work" put the octopus pages first. Paraphrased ones didn't:
"what makes something intelligent or smart" missed the octopus article, which says octopuses are "considered the most intelligent of all invertebrates." It returned thesaurus pages instead, with the best score at 0.16.
There were two causes. Each page was embedded as a single vector, so one sentence in a 10,774-word article got drowned out. And search always returned the top results, however weak they were. The next day I split pages into overlapping chunks, each with its own vector and a link back to its page, and added a minimum score so a weak match returns nothing instead of noise.
I build with coding agents, mostly Claude Code, and I treat their output the way I'd treat a new hire's first pull request.
Before an agent touches code I write the design: the units, the contracts between them, and what "done" means. My Prosper decomposer redesign ran to 1,150 lines before any code changed.
Each plan gets its own branch, like plan-a/core or plan-c/worm, and is merged only after its tests pass, so I can review, revert or bisect each one.
Tests on every package, and a labelled set wherever there's a quality question. I report the result even when it's bad, like Gate's 0.21 recall.
Before I call something done, a fresh agent with no context follows the README on a clean machine. Everything it trips over becomes a fix on a nu/fixes branch.
Gate, Warrant and Ledger started as tools I needed to trust my own agents. Gate recognises which agent wrote a commit from its trailers and tracks how often each one's changes are rejected.
Go, Python and TypeScript. Postgres, Redis, Qdrant and Docker. Claude Code and MCP to build with; Anthropic, Gemini, Groq and DeepSeek APIs behind provider fallback in production code. Playwright and headless Chromium when the question is "does it actually run".