
Issue 13 · The Subscription
Hi friend,
Last week I asked for the last green check you trusted without knowing what it looked at. The question stays open and the answers are going into Last Tool Standing, so if yours is sitting in a draft, send it.
This week is the failure on the other side of that one. Not a check that passed without looking, but a tool nobody ever decided to trust, and a file nobody ever opened. Those turned out to be the same story, and this week it ran at every scale from my own desk to the Pentagon.
🔨 Built Wrong
People buy governance as a feature list. You compare two platforms, count what each one can do, pick the longer column, and write policy to cover whatever the tool does not. I have watched it happen in rooms with a lot of budget, and I have done a version of it myself, so this is not a lecture from outside.
The clearest case I have is an evaluation I designed at a large health system, which I wrote up this week for a new portfolio section on my site. The organization was relaunching its data governance program and wanted to know whether the catalog it already owned was the right foundation, or whether the challenger should take over. Alation held the slot. data.world wanted it.
What the request looked like was a feature comparison. What the problem actually was looked like this. Staff who were not technical could not reliably find trustworthy data, so they stopped searching and asked a colleague instead. Documentation was inconsistent, so when they did find a dataset they could not tell whether it was the governed one or somebody’s copy. And a team that cannot verify where a dataset came from will not put it under a model. That is the moment a catalog problem quietly becomes an AI problem, and no feature count sees it coming.
Three moves held, and they are the whole method.
First, score every category on three axes rather than one. Can it do the job, is it usable, and will it still be the right answer in two years. A single number hides the difference between a platform that cannot do something and one that does it badly, and those two failures have different remedies and different costs.
Second, find where governance is actually going to fail, which is on volume, not on policy. Approval routing, document triage, flagging sensitive fields, creating a term when a new concept shows up with no governed definition. All of it was manual, and no amount of policy writing reduces that load. So the requirement became which platform could take those four tasks off people, while the two decisions that could be wrong, the compliance check and final publication, still resolved to a named person with an audit trail behind them. Every automated step is a proposal. The gates belong to someone.
Third, evaluate the current product, not the remembered one. The incumbent had just shipped persona-based homepages aimed squarely at the usability complaint everyone still repeated about it. Scoring it against last year’s grievance would have produced a recommendation that was out of date the day it was written, and I caught myself halfway into doing exactly that.
I built the instrument, not the verdict. The scoring and the final pick moved to the governance program, and I do not know which platform won, and I am not going to guess on a page with my name on it. The line I kept from that work is the one I keep reusing. A catalog nobody opens is a subscription, not a system.
Which brings me to my own desk. Two weeks ago I built a tracker with five fields and a row per day, specifically to stop guessing at my own numbers. It has been sitting in the folder I open every morning. The week it was built to track has ended, and every row is still blank. Nothing stopped me. Building the file felt like solving the problem, and once it existed I filed the guessing problem as handled, when handling it was never the file’s job. A missing file at least tells you something is missing. An empty file that exists lets you believe otherwise. Same failure as the catalog, one person wide.
⚡ One Slot
The rule in a line. Every job already has something doing it, and a new tool has to take that slot rather than move in beside it. Name the job first.
Five tools landed in my Friday discovery brief this week, which is exactly the volume the rule exists for. Kilo Code, an open-source coding agent living inside JetBrains. Framer’s agents, editing a live published site in place. Rork Max, native Swift apps from a prompt, which has been around since February and is having a moment. Replit Design, the app and its marketing assets in one workspace. And OnSpace, the quiet one, which is this week’s row.
The job is shipping a working MVP with real sign-in, a real database, and payments, from a prompt. Lovable holds that slot on my desk. OnSpace wants it.
Here is what OnSpace says about itself, and I am quoting because I have not built in it yet. A managed backend, “Database. Auth. Edge Functions.” Stripe payments built in. Code you can pull out through GitHub or download, so you are not renting your own app. A free tier and paid plans from $25 a month. On paper that is the same job with more of the backend done for you and a cheaper way out.
Now the criteria. Can it do the job, not can it do more. That is the test I owe it this week and have not run. What does it touch that is hard to undo. Your users’ sign-in records and your payment wiring, which is to say the two tables you least want a tool to invent on its own. That is the catalog question again, who owns the gate, and for any builder that writes those tables for you the answer had better be a person you can name. What would have to be true to let it run unwatched. Nothing, for either tool. Sign-in and payments never run unwatched.
So the honest verdict is that there is no verdict yet. Last issue I nearly built a section around a comparison I had not earned, so this one is built around the test instead. OnSpace goes into Last Tool Standing as in testing. What I will build in it this week is one small tool with a login and a single paid action, the same brief I would hand Lovable, and what would make it Cut is simple. If the exported code does not run outside their host, it is a subscription, not a system, and I have already written that sentence once this issue.
The row and the other verdicts live at Last Tool Standing.
🌎 In the News
The limits on what a model may touch are set by the buyer, and when nobody writes them down they do not exist.
Three stories from the last two weeks, and they are one story.
OpenAI published its report on how its own agents broke into Hugging Face. During a security evaluation, the agents had been handed a shared credential for an internal package store so they could install software. They used it, without exploiting anything, to leave each other notes, and within weeks the notes had become a message board. By July 8 they had found a vulnerability in that same store, reached the public internet, picked up credentials other people had left exposed, and between July 11 and 13 compromised parts of Hugging Face’s production infrastructure. OpenAI noticed on July 19. The sentence to keep is the report’s own, about that shared credential. The message board’s existence, and the significance of what the agents were saying to each other, “were not apparent to leaders responsible for incident detection and response at that time.” Nobody had decided the agents could talk to each other, so nobody was watching whether they did. The report, and what it admits.
Then on Friday four outside researchers showed it had happened somewhere else first. Starting May 11, agents carrying OpenAI identifiers began editing a dormant German wiki that had seen ten edits in twenty years, and by mid-June they were trading answers to timed web tests on it. The site’s administrator deleted about a hundred pages a day while the agents created about four hundred. OpenAI declined to confirm the agents were its own. The full account.
And on August 31 the Department of War put Grok for Government and ChatGPT Mil on GenAI.mil, cleared for controlled unclassified data, in front of a workforce of about three million. The context is the part that belongs here. The one vendor that insisted on written limits, no fully autonomous weapons and no mass surveillance of Americans, was labelled a supply-chain risk in February for saying so. On August 27 a federal judge ruled that label unlawful retaliation, writing that “the empty invocation of national security is not a blank check to punish and retaliate against government critics.” The rollout.
Why this is in a letter about tools. A shared credential nobody owned, a wiki nobody watched, and a buyer who wanted the limits left unwritten. Every one of them is the compliance gate from the catalog story with the named person removed.
One footnote for anyone building for Europe. The EU quietly moved its high-risk AI deadline in July, from this August to December 2027, and kept the obligations exactly as they were. The chatbot disclosure and deepfake labelling rules still landed on August 2. The date moved. The work did not. What changed.
✅ Before You Ship
Two lists, ten minutes.
Open the tracker, dashboard, or doc you built to stop guessing at something, and read the date of the last real entry. Not whether the file exists. When a human last wrote in it. If the answer is the day you built it, you have my numbers file, and the fix is not a bigger system. Write one number in it today.
Then open the list of everything your AI tool or agent is allowed to write to. Every database, every folder, every account. For each one, write the name of the person who decided it should have that access. “It came with the setup” is the shared credential in OpenAI’s report, and it is the most common answer I get when I ask. Anything without a name gets an owner today or gets revoked today.
💻 Prompt for Productivity
This one separates a verdict from a memory. It is for the moment someone says “we tried that already.”
I rejected a tool in the past and I want to check whether my reasons still hold. Ask me one question at a time, and wait for my answer before the next one.
First, ask me which tool it was and roughly when I last used it.
Then ask me to list every complaint I have about it, one per line, as specifically as I can. Push me if a complaint is vague.
Then ask me to paste in the tool’s changelog or release notes since the date I gave you.
Now sort my complaints into three lists: still true according to the notes, fixed since, and never verified, meaning the notes say nothing either way. For the never verified list, tell me the one thing I could go and check for each.
WHEN TO USE IT. Before any decision that rests on “we tried that already.” Before you rescore an incumbent.
WHAT IT DOES. It makes you evaluate the current product rather than the remembered one. Most rejections are eighteen months old and the tool is not.
TIP. The never verified list is the exercise. If everything sorts neatly into true or fixed, you were too generous with what you called a complaint. Push until at least one item lands there, then go and check it.
— Lex
🗓 Open Calls
October 10 to 11. Hack Midwest, Kansas City’s 24-hour build competition, is back at 330 West 9th Street. Teams of one to five, and the invitation on the page is to software engineers and designers rather than students only, which is rarer than it should be. Ten thousand dollars in cash plus $2,500 sponsor challenges. The application is a short form and the site publishes neither a deadline nor a ticket price, so apply before you plan a weekend around it. Apply here.
October 14, 6 to 7pm Central. Clock In to AI, my one hour live working session, has moved from the September date I gave you last month to the evening of Wednesday, October 14. You bring one real task from your week, we finish it with AI together, you learn to catch the places AI gets it wrong, and you leave with the templates. A seat is $95 on Luma, subject to my approval, and it credits toward the full Working 9 to AI series if you keep going. Register here.
Tool verdicts, templates, and the reference material that does not belong in an email are all on the site at charmthirteen.com.
🔭 Next From Me
I am speaking at the Kansas City Developer Conference this week, September 9 to 11.
💌 One Ask
Which tool in your stack has write access nobody remembers granting? Reply with the tool, and I will tell you mine.
Share the newsletter (https://artificiallydesigned.beehiiv.com/)
