There's a story going around from a frontend dev who was migrating his website tools into a mini-program. Image compression, a nine-grid splitter, a QR generator, a mortgage calculator, the usual utility drawer. Most of it was straight code migration since both sides were TypeScript, and with AI help even that stopped being the slow part.
Then he hit the ID-photo tool. The hard bit was separating a person cleanly from the original background so it could be swapped to white, blue, or red, while keeping hair edges and face outlines from looking fake. His first pass was, in his words, functional but unusable. Edges were stiff, hair and ears especially. His plan was to go find an open-source algorithm or just call a third-party API.
Instead the model went off and worked the problem itself. It kept rewriting the algorithm, testing on images, adjusting the edge handling, and spent about an hour and a half before landing on a new approach. It even offered a second route: load a lightweight segmentation model locally. He didn't fully adopt that one because of mini-program constraints, model size, and performance, but the plain image-processing version it wrote ended up better than he expected.
The part everyone talks about, and the part worth keeping
The headline everyone latches onto is that the AI crossed into a domain he wasn't expert in and "wrote the algorithm." Fair enough, that's the dramatic bit. But watch what actually happened in that hour and a half: tweak the algorithm, re-run it against some test images, eyeball whether the hair looks less fake, tweak again. That's the real work, and it was completely ad hoc. No record of which version was which, no fixed set of images to judge against, no way to say "version 4 is better than version 2 at ears" except by memory.
That's the thing worth keeping. Not the specific segmentation code, which is disposable, but the loop around it. If you're doing any kind of image processing where quality is subjective, you're going to run that loop dozens of times, and doing it by feel is how you end up shipping a regression you can't explain.
So the reusable move is to turn the loop into infrastructure. Three pieces, all inside things a small builder can actually stand up.
Piece one: the algorithm as a hosted Python endpoint
The background-removal logic wants to be a real service, not a snippet you paste into a chat and lose. You can run Python online and expose the processing function as an API endpoint that takes an input image plus a background color and returns the composited result. Every version of the algorithm becomes a callable you can hit the same way, which matters because consistency of invocation is what makes comparison honest.
One thing to be clear about: keeping the processing in a Python backend, rather than loading a model on the client, is a different trade-off than the local-model route the original dev considered. Server-side means your test harness and the production path share the exact same code, which is what you want when you're comparing versions. The mini-program side is outside what I'd map here, since that's a separate runtime and framework; the endpoint is the reusable core, and how a given client consumes it is a boundary I won't pretend to cross.
Worth flagging up front: an image-processing endpoint sitting open on the internet is something you should put access control on before you point real traffic at it. It's easy to stand up an unauthenticated endpoint and forget it's public.
Piece two: a SQLite ledger that pins versions to the same images
This is the piece that fixes the "eyeball it" problem. Keep a table where each row is one algorithm version run against one test image. Columns for version label, a short note on what changed, the source image, the output, and whatever quick judgment you want to record about the hard regions, hair and ears in this case. Use an in-platform SQLite database for it, and when you need to inspect or correct rows by hand, the SQLite editor lets you do that without writing a throwaway script.
The discipline here is boring but it's the whole point: the test images stay fixed. Same set every time. When the model hands you version 5, you run it against the identical inputs version 2 saw, and now "is this actually better" is a query, not a memory. You can also spot the classic trap where a change fixes hair but quietly wrecks something that was fine before.
Piece three: an HTML before/after board to judge by eye
Segmentation quality is subjective, so at some point a human has to look. Give yourself somewhere to look properly. You can run HTML online to publish a simple board that reads from the ledger and lays out source, version A, and version B side by side, zoomed in on the edge regions. That original dev spent his time flipping between outputs to judge hair edges; a fixed comparison view is exactly what turns that from a vibe into a decision you can defend.
If you want the board to pull live comparison data, it can call the same hosted endpoint through the platform's tool calls, so the page you're eyeballing and the algorithm you're testing never drift apart.
Why this generalizes past ID photos
The pattern shows up anywhere output quality is fuzzy and iteration is the job. Look at the OCR browser extension someone shipped that runs local recognition on hover, or the tools being migrated where compression had a mature open-source answer but segmentation didn't. The moment a model starts iterating on something judged by eye rather than by a passing test, you need a way to record what you tried and compare fairly. Otherwise the model's hour of tuning produces a result nobody can reproduce or verify.
I'd treat any quality or performance claim about the resulting algorithm as unverified until it's run against the pinned image set and the numbers are in the ledger. The original account is one developer's impression that the output beat his expectations, which is fine as motivation but isn't evidence. The harness is what turns impression into something you can actually stand behind.
And if the tool earns its keep, the endpoint, the ledger, and the board are already hosted in one place, which means publishing or monetizing it is a step you've set yourself up for rather than a rebuild. The algorithm the model wrote will probably get replaced twice over. The loop you built around it is what stays.