Day One With an AI Agent Team: 6 Things That Broke (and the Fixes)

 


On September 26, 2026, we set up a small "company" whose staff are AI agents. Claude does research, writing, and review. OpenAI's Codex and Google's Antigravity make images. The person in charge doesn't write code.

Most setup guides show the happy path. This post is the other half: the six things that actually went wrong on day one, why they happened, and what we changed. Every item comes from our internal failure log, so none of this is hypothetical.

If you're a non-developer wiring AI tools together on a Windows PC, at least one of these will probably save you an hour.

1. The image tool had no quota left before we started

What happened. Our first image request to Codex failed instantly with "You've hit your usage limit," and a time when it would reset.

Why. Codex usage on a ChatGPT plan is shared across the places you use it. OpenAI says the same limits apply in the app, the CLI, the IDE extension, and the cloud (OpenAI Help Center, checked 2026-09-26). We hadn't run anything through the CLI yet, so the quota for that window had been used up elsewhere on the same account. Image generation also uses limits 3 to 5 times faster on average than a normal turn, according to OpenAI's docs (checked 2026-09-26).

Fix. Before a batch of image work, check the usage page first. When we measured it after the reset, one image took the 5-hour meter from 100% to 99% on our Plus account. That's one measurement on one account, not a rule.

2. A tool installed fine but "didn't exist" for anyone else

What happened. We installed Google's Antigravity CLI from inside a desktop app. It worked there. When the founder opened their own terminal, Windows said the program didn't exist at that path.

Why. The desktop app we were working in is a packaged (MSIX) Windows app. For packaged apps, Windows copies writes to the user's AppData folder into a private, per-app location (Microsoft Learn, checked 2026-09-26). The install really happened, just somewhere only that app could see.

Fix. Install tools that live in AppData from your own terminal, outside the app. Keep your project files outside AppData too.

3. The agent quietly skipped a step

What happened. We asked Antigravity, run without a human watching, to make an image and save it into a folder. It reported success, but no image appeared.

Why. In headless mode, actions that normally need your approval can't prompt anyone, so they're auto-denied and the run keeps going (Antigravity docs, checked 2026-09-26). The agent had tried to run a command to move the file, and that command was denied.

Fix. We didn't switch on the "skip all permissions" flag. Instead we told the agent to use only its image tool and return the file path, and we moved the file ourselves. Narrow permissions, one extra step.

4. The image model invented our text

What happened. We asked for a blog thumbnail showing "a checklist with three checked items." The image looked great, but the model wrote its own three items, which weren't the principles the post was about.

A second try put stray quotation marks around the headline, because we had wrapped the text in quotes in the prompt.

Fix. Spell out every piece of text that should appear, and add: "Render all text exactly as given, without any quotation marks or extra words." Don't put quote marks around the text you want drawn.

5. "Keep everything the same" didn't keep it the same

What happened. We gave the model the approved image and asked it to change only the checklist labels. The labels came back right, but the white background had turned into gray paper texture and the magnifying glass had turned gold.

Fix. Restate what must stay, specifically: background color and texture, panel positions, object colors, and style ("flat vector, no shading"). With that spelled out, the second attempt matched the original.

6. We published the wrong page, and only checking caught it

What happened. We publish by copying prepared HTML and pasting it into Blogger. The Privacy Policy page went live containing the Contact page text. The clipboard still held the previous page.

Fix. Two changes. The copy button now says which file it copied. And after anything is published, we fetch the live page and check for a few key phrases that must be there. That check is what caught this one.

The pattern behind all six

None of these were exotic bugs. Each one was a gap between what we assumed and what was actually happening:

  • Limits are shared. Check before you start.
  • Tools run in contexts. Where you install something decides who can see it.
  • Unattended agents fail quietly. Give them narrow jobs and verify the output.
  • Image models fill in blanks. Leave none.
  • "Published" isn't "correct." Look at the live result.

We'll keep logging what breaks as we go. Next up: how we set up the dashboard the founder uses to give orders and approve anything that goes public, without opening a terminal.

*How this post was made: drafted by an AI agent, fact-checked by a separate AI reviewer against the sources linked above, and approved by our human founder. All six incidents happened on our setup on September 26, 2026.*

Comments