CriticalKPI
Builds

Applied AI Engineering

In the last three months I've designed, directed, and verified about two dozen applications, tools, and sites with AI coding agents, and built the system that keeps them honest. Here's what exists, what it does, and how I know it works.

What this is, and what it isn't

I design the systems, write the specs, direct the AI agents, and verify what they produce. The agents write most of the code. That's a different job from "vibe coding," where you accept whatever runs. The difference is everything around the code: specs written first, work isolated so parallel agents can't collide, review gates that have to run a check instead of reading a diff, automated tests, an incident log, and a human approving anything that touches the outside world.

Claude Code does most of the building. Other models play supporting roles: OpenAI's Codex for second-opinion reviews, local models for offline work. The tools rotate. The process doesn't.

By the numbers

Counted from AgencyIQ's own usage tracking and git history as of October 2, 2026, and rounded down. A raw count means little on its own, so each one comes with what it tells you.

MeasureCountWhat it tells you
AI-assisted work sessions tracked6,800+ since July 30Sustained daily use, not a weekend experiment. Every session is logged, so these figures come from records, not memory.
Commits across 14 repositories6,600+Every change is versioned and reversible, which is what makes unattended agent work safe to undo.
Usage events logged800,000+Every AI request is recorded. That record is how AgencyIQ tracks spend against plan limits.
Applications, tools, and sites builtAbout 24Breadth: the same method working across apps, automation tools, and production websites.
Custom skills27Procedures I've written down for the agents, so the same job gets done the same way every time.
Agent definitions23Specialist roles, such as accessibility QA or security audit, each with defined limits on what it may touch.
Hook and guard scriptsAbout 30Automated tripwires that block risky actions, like editing the shared checkout, before they happen.
Incidents logged and triaged400+Problems get written down and fixed at the cause. A long log means I track failures instead of hiding them.
Backend test code vs. application codeAbout 1.5 to 1 (by lines)Tests are how I verify what the agents write without reading every line. More test code than app code means changes get checked, not assumed.

Running many agents without collisions

Several AI sessions working on one codebase will eventually edit the same file, claim the same database migration number, or merge on top of each other. Early on, that's exactly what kept happening. When about a dozen sessions ran at once, a pile of them silently froze.

The fix wasn't a rule. It was structure. Every task gets its own git worktree, so most collisions can't happen at all. A work-claims registry covers what can't be isolated, like the migration sequence. Hooks block edits to the shared checkout unless a session has declared it belongs there. The Conductor caps concurrency at about five, a detector reaps frozen sessions, and merges land one at a time behind a lock. I instrument all of it: more than 1,400 collision events logged so far, which is how I tune it.

Before anything merges, it passes a task-level review and a whole-branch review. Both were hardened after the system caught itself rubber-stamping bad code. Each one has to actually run something, such as a rebuild, the tests, or an attempt to break the change, and CI has to be green on current main.

Hollering Toad: audits that get verified

SEO and GEO audits usually end as a document someone forgets. Hollering Toad (named with a nod to Screaming Frog) crawls a site, scores it, and records every recommendation as a tracked requirement. Fixes go to the coding agent, an independent pass verifies them against the spec, and the site is re-audited and re-scored. The audit process behind it has run on three of my own sites, including this one.

Home intelligence: data into findings

The analytics instinct from my day job, pointed at my own house. A private platform pulls solar production, utility bills, water bills, spending, and investment holdings into one place, with scheduled ingestion, an idempotent importer that de-dupes real-world exports, and a spending engine that learns merchant categories from history.

The point is what it found. An anomaly detector compares each period's solar output to the same calendar month in other years, and it flagged a nine-week outage on its own, one I'd missed, which explained the two highest bills on record. It raised no false flags anywhere else in the history. A second analysis showed the rising bills were mostly more usage, not higher rates: net usage climbed about 30% year over year while the per-kWh rate fell. That's the difference between a dashboard and an insight.

The platform holds real family and financial data, so it stays private. I describe the method here and rebuild it on synthetic data when I need to show it.

Everything I've built

Status is stated plainly. "Running" means it's in regular use. "Proof of concept" means it works on fabricated data and hasn't been in production.

BuildWhat it doesStatus
AgencyIQThe operating platform behind my businesses: CRM, email, social, SEO and GEO audits, project and resource management, reporting.Running
The ConductorQueues approved milestones by priority and available capacity, dispatches isolated agent sessions, and routes results to review or blocked. I accept the finished work.Running
Review and incident systemHardened review gates, an incident log, health sweeps, and an independent playbook for verifying AI-implemented work.Running
Overnight runner and usage governorRuns headless work overnight, tracks AI usage against plan limits, and throttles concurrency when headroom runs low.Running
Agent roster and autonomy tiersDefined agents (accessibility QA, anomaly detection, security audit, test generation, and more) under a written charter. External publishing, deploys, spending, and deletion always need a human.Built, expanding
Hollering ToadSEO and GEO audit with a requirements ledger and a verify-and-rescore loop.Built, in use
SocialFlowMulti-brand social automation with swappable plugins for content, publishing, models, and analytics. Instagram, Buffer, and GA4 attribution.Built
CRM and LinkedIn outreachContact import, inbox and de-duplication, and an outreach timeline. Real contacts, so private.Running, private
CommerceIQAn ecommerce analytics and personalization demo on simulated shopper sessions.Proof of concept
PL Claims AIA claims-management copilot with a ranked daily worklist that explains every ranking in plain language. Every case in it is fabricated.Proof of concept
SellerCopilotA single-brand marketing app for small sellers: brand voice, approval queue, scheduler, multi-channel publishing.Built
Catalog enrichment and scoringAI that fills gaps in product data and scores catalog quality, built on my experience in ecommerce product discovery and analytics.In development
AI Video StudioTurns raw footage, or Q&A pairs from an authority site, into cut, captioned or narrated, branded short videos with no manual editing.Built
KATiqA private, local-first family intelligence platform: search across archived files, finance, energy, and household records. Described, never published.Running, private
Home intelligenceThe energy, spending, and portfolio analysis described above.Running, private

There's also a shelf of smaller tools: a frozen-session reaper, a Search Console sitemap submitter, and a photo restoration script. Around ten sites sit under all of it, several deploying automatically through GitHub Actions.

The code

Most of this is private because it runs my businesses and touches real data. I'm preparing clean public versions of the self-contained pieces, built on synthetic data, and I'll list them here when they're up. Until then, I'm glad to walk through any of it live.

Want to see how it works?

I'll walk you through the Conductor, the review gates, or the analytics, whichever matters most to what you're building.