Prosper Nwo

I build the infrastructure that keeps AI agents honest: it checks what they change, limits what they can touch, and writes down what they did.

Most of my code is written by coding agents that I plan, direct and verify. So the problem I care about is how you trust software that a model wrote. I'm looking for an AI engineering role on agent infrastructure, evals or developer tools.

Work
Sep 2026 v0.1.0 Go · TypeScript · Postgres

Prolu

Six self-hosted tools for running AI agents safely. Each works on its own, and all of them log to Ledger.

Gate
Decides whether a code diff from any agent can be merged: approve, review or reject. Runs as a CLI, a GitHub Action, or an MCP server the agent calls on its own work. 472 tests
Ledger
Append-only decision log. Records are SHA-256 hash-chained in Postgres, chain roots are Ed25519-signed, and the auditor export maps to EU AI Act Articles 12 and 14. 244 tests
Warrant
When one agent delegates to another, it issues a short-lived token scoped to that job. Every tool call passes through a proxy the model can't reach. 157 tests
Harbour
Runtime for long-running agents. Goals survive crashes, you can attach to one from any machine, and two-phase commit means a retried action never runs twice. 158 tests
Bench
Turns a repo's own bug fixes into benchmark tasks, then ranks coding agents on them. It answers "which agent should we use?" with your code instead of a public leaderboard. 107 tests
Proof
For crawlers: records what robots.txt, content signals and the page licence said at fetch time, in signed manifests, and reports when those rules change later. 191 tests

How I know Gate works

I built a labelled set of 71 diffs across Go, Python, JavaScript and TypeScript: 44 with a planted bug and 27 clean. This is Gate with no LLM, using only its static and security checks:

bug classcasescaught
sql injection44
hardcoded secret33
path traversal22
logic, off-by-one100
authz, concurrency80
test tampering, deleted tests80
api misuse, dependency confusion90
clean diffs wrongly flagged270

Precision 1.0, recall 0.21. The deterministic checks never raise a false alarm, but they miss every bug that needs you to understand the intent. Catching those is the LLM review stage's job, and I haven't benchmarked it yet. I published both numbers because the low one tells you more.

Aug–Sep 2026 Python · FastAPI · React · Phaser 62nd of 484, AMD Developer Hackathon ACT II

Prosper

Turns a one-paragraph game idea into a playable browser game. You can play three of them below. An LLM splits the idea into about nine units, agents build them in parallel, every unit has to pass verification, and a repair agent fixes any game that still fails the play test.

Void Runner, a game Prosper built

open in a new tab

Three finished builds that passed Prosper's headless play test. These are the merged games exactly as Prosper produced them. Click the game once so it gets your keyboard.

Prosper conductor screen: a snake game brief on the left, and eight build units with their dependencies and status on the right
Conductor, 23 Aug 2026: one brief split into eight units, building in parallel

By the numbers

8 Aug to 11 Sep26 to 29 Sep
games started22324
units approved58%87%
games that reached the browser play test6018
playable on the first test 4 (22%)
playable at their last test35 (58%)13 (72%)

From Prosper's own build database: 246 games and 2,285 units in total. The second period adds the genre starter games, a stricter play test that fails any runtime error, and Jules. Because the play test is stricter, only 22% pass it first time. Regeneration and repair lift that to 72%. The first period didn't record first-test results separately.

Jules, the repair agent

When a game fails the play test, Prosper copies it to a repair folder and hands it to Jules, an OpenCode agent running on free models. Jules gets the spec, the exact failure evidence, and a check command that runs the real merge and play test. Two sessions race on separate API keys, and the first to pass wins.

Prosper doesn't take Jules's word for anything. After every round it re-runs the play test itself, restores any pipeline-generated file Jules touched, and copies a fix back only if it passes. Every round is logged with its diff, and when the same kind of repair keeps showing up, I fix that bug in the pipeline instead.

In its first two days, Jules worked on 7 games that failed the play test and made 5 of them playable, over 33 rounds and about 22 hours of agent time. One fix took a single round: a ReferenceError from a variable used before it was initialized. Another took three rounds and four files, after a sprite call that didn't exist crashed the game loop. The two it couldn't fix both came from add-ons calling methods the game didn't have, so I fixed that at the source: each add-on now declares what it needs from the game, and installation fails immediately if the game doesn't provide it.

  • speed

    Single steps ran for over 10 minutes because step time grows with reasoning tokens. I switched the default to low reasoning effort.

  • stalls

    29 of the 33 rounds ended on a "no progress" timeout. A 7-minute stall limit killed sessions in the middle of a step. I raised it to 13 minutes, and a hung request now retries in the same session after 4 minutes instead of ending the round.

  • scope

    The repair folder sits inside the repo, so Jules wandered into pipeline code. It's now blocked from reading or searching outside the game.

  • restarts

    A new repair session now starts from the previous session's edits, so a restart keeps the progress already made.

Genre coverage

Nine hand-built starter games across eight genres, like Star Lancer (space shooter), Gem Cascade (match-3) and Bastion Road (tower defense), come with asset kits and 16 add-ons like bosses, extra waves and relics that scale a game up. Seven genres have a play test that actually plays the game. The shooter test, for example, fails a run where one enemy shot destroys the ship.

To check coverage, I classify 101 realistic game requests. 97 land in a genre (96%) with 0 misclassified, and 81 (80%) have a starter game today. The four it can't place (3D shooter, MMORPG, rhythm game, fighting game) are out of scope for now. Management and story games have no starter yet.

What broke first

In August I audited the pipeline against four real builds. The verification looked strict but was mostly letting everything through:

  • gate

    The browser play check set ok = False on failure, but the function never returned it. No build could fail.

  • motion

    The "is anything moving" check compared two array objects instead of their contents, so it always saw motion, even in a frozen game.

  • boot

    The Phaser boot check skipped itself because the platform string was phaser+js, not phaser.

  • cache

    Empty units all hashed to the SHA-256 of an empty string, so one unit's failed verdict was reused for every later empty unit.

All four are fixed. Every JavaScript and TypeScript file is now parsed and bundled with esbuild before approval, and the play test fails a game on blank rendering, placeholder art, game-code errors, or a run that ends without taking any input.

Aug 2026 TypeScript · BullMQ · Playwright · Qdrant

Discover

A research crawler for agents, so they don't need a search API. It crawls outward from seed URLs, streams each page to the agent as it's extracted, and stores everything in vector memory the agent can search later over MCP.

What broke

Literal queries worked. "How does octopus camouflage work" put the octopus pages first. Paraphrased ones didn't:

"what makes something intelligent or smart" missed the octopus article, which says octopuses are "considered the most intelligent of all invertebrates." It returned thesaurus pages instead, with the best score at 0.16.

There were two causes. Each page was embedded as a single vector, so one sentence in a 10,774-word article got drowned out. And search always returned the top results, however weak they were. The next day I split pages into overlapping chunks, each with its own vector and a link back to its page, and added a minimum score so a weak match returns nothing instead of noise.

How I work

I build with coding agents, mostly Claude Code, and I treat their output the way I'd treat a new hire's first pull request.

  1. Write the plan first.

    Before an agent touches code I write the design: the units, the contracts between them, and what "done" means. My Prosper decomposer redesign ran to 1,150 lines before any code changed.

  2. One plan per branch.

    Each plan gets its own branch, like plan-a/core or plan-c/worm, and is merged only after its tests pass, so I can review, revert or bisect each one.

  3. Measure it, then publish the number.

    Tests on every package, and a labelled set wherever there's a quality question. I report the result even when it's bad, like Gate's 0.21 recall.

  4. Newcomer pass.

    Before I call something done, a fresh agent with no context follows the README on a clean machine. Everything it trips over becomes a fix on a nu/fixes branch.

  5. Build the guardrails I'm missing.

    Gate, Warrant and Ledger started as tools I needed to trust my own agents. Gate recognises which agent wrote a commit from its trailers and tracks how often each one's changes are rejected.

Also built
  • CelarisBrowser space MMO: React and Three.js client, 40 Node services, Postgres and MongoDB. One command deploys it to a server.2026
  • NagBotHomework reminders that escalate across Discord, WhatsApp and email until a vision model confirms a photo of the finished work.2026
  • AERISMy first multi-service system: a personal assistant split into nine Python and Node services with shared vector memory.2026
Tools

Go, Python and TypeScript. Postgres, Redis, Qdrant and Docker. Claude Code and MCP to build with; Anthropic, Gemini, Groq and DeepSeek APIs behind provider fallback in production code. Playwright and headless Chromium when the question is "does it actually run".