← All posts Mobile

Six Apps, One Person, and an Agent That Ships Them

What it actually takes to let an AI agent run App Store and Play Store releases: the rules that had to be written down, the one thing it is never allowed to do, and why the tedious parts are the ones worth handing over.

There are six apps under my name on the App Store and Google Play — MoveWell, Restoria, Bout2, GroupFood, Commuticate and Breaking News Maker. There is no release engineer, no build team and no one to hand a checklist to. There is me, and an agent that runs the pipeline.

The interesting part is not that this works. It is what had to be written down before it did, because almost none of it is the sort of thing that appears in documentation. It is the accumulated residue of things going wrong.

Releases are exacting, not hard

Shipping a mobile build is not intellectually difficult. It is a long sequence of precise steps where a single wrong detail costs a day, and where nothing tells you which detail was wrong.

That combination — well specified, tedious, unforgiving, high consequence — is exactly where handing work to an agent pays. Not because it is clever, but because it does not get bored on step nine of eleven, and because writing the rules down forces you to actually have rules.

What follows are the constraints that turned out to matter. They are the content here; the automation is just what enforces them.

A version and build pair is used once, ever

The build step guarantees the (version, build) pair is fresh before it does anything else. Never re-shipping a pair sounds obvious and is trivially easy to violate at 6pm when you are re-archiving after a small fix.

App Store Connect will reject a duplicate build number, which is the good case. The bad case is Android, where you can quietly produce a bundle that collides with something already in a track and spend twenty minutes reading an unhelpful Play Console error.

iOS first, Android only after TestFlight

The two platforms are built in separate phases on purpose, and the ordering is the point: iOS archives now, Android waits until TestFlight has confirmed the build is good.

If a build is broken, TestFlight tells you before Android was ever built. Building both in parallel feels faster and means a bad build costs two bundles, two uploads and two sets of store metadata instead of one. The Play build then mirrors whatever version and build the iOS archive shipped with, so the two can never drift.

The build is not done until the file is in Finder

This one is genuinely surprising the first time, and it is why the rule is written in capital letters in my setup.

An archive built with xcodebuild from the command line does not appear in Xcode Organizer. Archives built through the Xcode UI are registered; CLI archives are not. So the terminal says ** ARCHIVE SUCCEEDED **, you open Organizer to distribute it, and the list is empty.

The fix is to reveal the .xcarchive in Finder so it can be double-clicked, which registers it. Which makes the operational rule:

The build is not "done" when the command exits zero. It is done when the artifact is highlighted in Finder.

That is a small thing, and it is precisely the kind of small thing that makes automation useless if it is missed — a pipeline that reports success and leaves you with nothing to click has not saved you anything.

Do not pass authentication keys to the archive step

A signing failure with a misleading message, worth knowing because the error sends you in entirely the wrong direction.

Passing -authenticationKey arguments to xcodebuild forces cloud signing. Cloud signing then fails, claiming the certificate is not in the provisioning profile — so you go and inspect your certificate and your profile, both of which are fine.

The certificate was never the problem. Do not pass those arguments to the archive.

Store copy has hard limits that look soft

Release notes are where an agent writing text will cheerfully produce something that gets rejected or, worse, published looking broken.

  • Play's "What's new" is 500 characters, and line breaks count. Not "about 500". Verify with wc -c before printing, and cut if it is over.
  • Plain text only. Store listings render these blocks as-is. Markdown does not get parsed, it gets displayed — asterisks and underscores sitting there in your listing.
  • Use the real bullet character, not a hyphen and not an asterisk.
  • No emoji in App Store copy. Play renders them reliably and they suit the terser style there. If the feature genuinely is about reactions, spell them out rather than listing glyphs.
  • No commit hashes, file names or internal vocabulary. "Fixed the deep-link handler" is not a sentence for a user. Say what they can now do.

There is also a hard stop: if the iOS marketing version and the Android versionName disagree, the pipeline refuses to continue. Mismatched versions across stores are the kind of mess that is invisible for a month and then impossible to reason about, so it is better to fail loudly at the point of release.

The one thing the agent never does

Here is the boundary, and it is the part I would argue about with anyone automating this.

Every archive uploads itself to TestFlight without asking. Export, validate, upload — no prompt, no confirmation. That work is repeatable and low consequence. A bad TestFlight build costs a build number.

Submitting for App Store review waits for me. Always.

The difference is not risk exactly, it is reversibility. Submission puts you in a queue where the cost of being wrong is not "try again", it is "go to the back". I have watched an app sit fifteen days at Waiting for Review, and resubmitting would have reset that clock. Anything whose failure mode is measured in weeks keeps a human in the loop.

Android is deliberately asymmetric in the same way, and the asymmetry is the design rather than an inconsistency I have not gotten around to fixing.

Approved is not downloadable

The last step is the one most likely to be skipped, because by then the interesting part is over.

When a release goes live, the API's recommended-version defaults and the admin dashboard's version-drift indicator both need bumping. Those two have drifted apart before, which is why it is a written procedure rather than something I remember.

And it has a mandatory precondition: do not treat "approved" as the trigger. Approved and actually downloadable are different states. Bump the recommended version too early and your API starts nudging users toward a build they cannot install yet, which is a strange and annoying bug to receive a report about.

Why this adds up to scale

None of the above is impressive on its own. Together they are the reason six apps is manageable for one person.

The work that makes shipping expensive at small scale is not the code. It is that each release carries thirty small exact obligations, most of which you only learn by violating them, and all of which have to be right every single time. A team absorbs that with process and people. Alone, you absorb it with attention — and attention is the thing that runs out.

Writing the rules down so something else can enforce them converts a recurring tax into a one-time cost. The agent is not making judgment calls. It is holding thirty details steady so that the judgment calls — is this build good, should this ship — are the only things I have to spend attention on.

What it does not solve

Being straight about the limits, because the genre this post belongs to is usually oversold.

It does not decide whether a release is ready. It does not catch a regression TestFlight would have caught. It does not write store copy that is good — it writes store copy that is valid, and those are different, and I still edit it. It cannot talk to App Review. And every rule above exists because something went wrong first, which means the list is not finished; it is just the part I have learned so far.

What it does is make the boring half of shipping reliable, which turns out to be most of the reason shipping is hard.