Field notes / AI systems · Case study
I Don't Have an AI Employee. I Have a Chief of Staff.
Everyone selling AI right now promises a tireless employee. I built the opposite: an assistant with full context on my business, real capability, and no authority. It is the best hire I never made.
One afternoon this month, my operation cleared a 13-day email backlog across every deal in the pipeline, produced a six-figure offer on a property with math I could defend line by line, killed 13 stale deals with a logged reason on each, and prepped a negotiation brief for a partnership conversation the next morning.
I made maybe ten decisions that afternoon. I did not type most of what went out, and I did not open a spreadsheet once.
This is not a post about a product you can buy, and it is not a hypothetical. It is a case study of my own operation, written down because I get asked about it constantly and because it is the best answer I have to the question underneath the whole Five Levels series: what does this actually look like when someone who builds AI systems for a living runs his own business on one?
The answer is not an AI employee. An employee implies delegation of judgment, and judgment is the one thing I refuse to delegate. What I run instead is a chief of staff: an assistant with full context on every deal, every inbox, and every number, the freedom to prepare anything, and a hard rule that nothing leaves the building and no dollar moves without a human yes.
Everything below is a consequence of that one split: maximum capability, deliberately scoped authority. Here are the ten patterns that make it work.
1. A chief of staff, not an autopilot
The assistant can see everything: three email accounts, my calendar, a searchable archive of my WhatsApp deal conversations, the full pipeline with every deal's numbers and history, and a browser it can drive for anything behind a login. On the capability axis, I have held nothing back. It reads faster than I do and it never forgets a thread.
The authority axis is a different story. It drafts every outward email; it sends only when I say "send it," to a recipient I have named. It computes offer numbers and tells me what it would do; I decide. It can mark a deal dead in the pipeline, with a reason, only when I tell it to. The pattern in one sentence: it prepares everything and decides nothing.
That is exactly what a great chief of staff does. Full context, total preparation, and a clean handoff of every decision to the principal. The moment you blur that line, you have an employee whose judgment you have to audit. Keep the line sharp and you never audit judgment at all; you only review work product, which is fast.
2. The business lives in files, not in the AI's head
The single most important architectural decision, and the least glamorous: the business does not live inside the assistant. It lives in a folder of plain files under version control. One folder per deal, holding the numbers, a log of every communication, and running notes. A decisions file that records every modeling default I have ever set, with the reasoning. A daily journal. Playbooks for each stage of the pipeline.
The assistant reads and writes that record; it does not replace it. Which means any session, any model, any future tool, or any human I bring on can be rebuilt from the files alone. If my AI vendor disappeared tomorrow I would lose a very good assistant and none of my business.
This is what makes the chief-of-staff pattern safe. The assistant is replaceable. The record is not. If you adopt one pattern from this post, adopt this one, because every other pattern stands on it.
3. Deterministic core, AI at the joints
None of my underwriting math is done by a language model. The engine that turns a rent roll and a purchase price into a verdict is ordinary tested code, the same answer every time, with the buy-box thresholds written down and versioned. When I said the six-figure offer had math I could defend line by line, that is why. There is no "the model estimated" anywhere in the number.
The AI sits at the joints, where language and judgment live: parsing an inbound deal package into structured numbers, updating a deal's status from a pasted email thread, prioritizing the day, drafting the reply. Those joints are exactly where models are magic and exactly where they occasionally get things wrong, so each one has a test set of golden cases, and every time I correct an output, the correction becomes a new test case. I wrote about that discipline in the evals post; this is it running in production on my own deals.
The rule of thumb: if it must be right every time, it is code. If it requires reading, judgment, or tone, it is the model, with a test set watching it.
4. The daily check-in is the interface
I do not manage this system by remembering to ask it things. One command compiles the day: sourcing quotas due, replies I owe people, deadlines approaching, verification queues, and every deal that has gone quiet past its stage's shelf life. The report gets journaled, so there is a written record of what the business looked like every single day.
The important design choice is that deals decay visibly. Every pipeline stage has a service-level clock, and when a deal sits past it, the deal shows up flagged, every morning, until I either act or kill it. Silence has a cost that appears on the board. The 13 deals I killed that afternoon were not a purge I got inspired to do; they were 13 flags the system had been patiently raising, each one resolved with a reason that is now part of the record.
My job in this interface is to make decisions off the report. The assistant's job is to keep the report true. That division has survived every busy week I have thrown at it.
5. The assistant has its own assistants
A chief of staff who personally does every task becomes the bottleneck. Mine delegates. When the pipeline needed a full review, it sent read-only worker sessions to sweep all 17 active deal folders in parallel and returned one triage table. Bulk email screening runs in an isolated sandbox. Long scrapes run as background jobs while the main conversation keeps moving. The main session keeps conclusions and throws away the raw dumps, which is the only way to keep a long working session sharp.
The pattern extends past software, which is the part I did not expect. When a human analyst started doing underwriting support for me, we gave her deliverables their own mailbox, and the assistant now triages that mailbox like any other feed: new workbook arrives, gets checked against the deal record, discrepancies get flagged for me. The org chart is: I manage one chief of staff, and the chief of staff manages the feeds, human and machine alike.
6. Autonomy only at the read-only edges
Exactly one piece of this operation runs with no human in the loop: every evening, a scheduled job syncs my deal-flow group chats, has a model summarize what moved, and sends me a digest. Deals surfaced, buyers hunting, market chatter, distilled from hundreds of messages into two minutes of reading.
It earned full autonomy because it meets a simple test: it reads and summarizes, and that is all. It cannot send anything outward, cannot touch the pipeline, cannot spend a dollar. If it has a bad night, I get a mediocre summary, and that is the entire blast radius.
That is the general principle: automation level proportional to blast radius. Plenty of things in my business could be automated end to end. The ones that are, are the ones where the worst possible output is a bad paragraph.
7. Memory that survives the conversation
AI conversations end, and the fanciest context in the world dies with the session unless you engineer against it. I run three layers, none of which required a database. First, a persistent memory directory: one fact per file, one line each in an index the assistant loads at the start of every session. Second, a pair of project indexes, one for code repositories and one for document projects, so any session can find any piece of work in seconds. Third, the deal record itself, because a disciplined comms log and a daily journal are a memory system, and a searchable one.
People assume an operation like this needs vector databases and retrieval infrastructure. Mine does not, yet. Plain-text search over well-organized files gets you much further than the industry wants you to believe. The discipline is in the writing, not the retrieval: if every decision, correction, and conversation lands in a predictable place with a predictable shape, finding it later is trivial.
8. A written persona, so the voice stays mine
There is a document that defines how the assistant behaves in this business: keep the record current, surface decay, draft when asked, recommend freely, never decide for the principal, never send outward. It also carries the style rules, down to formatting details, so that a drafted email is paste-ready without cleanup and reads like me on a good day rather than like a press release.
This sounds cosmetic and is not. In a relationship business, an email that smells automated is expensive. Sellers, agents, and capital partners are deciding whether to trust me, and they are correct to read tone as signal. The persona document is how a system this automated stays personal, and it is also the file a second operator would read first if I ever handed the keys over.
9. Capabilities get built the day the friction shows up
Nothing in this system came from a roadmap. Every integration was built the day the friction was felt, usually inside the same working session. Zip-level market data looked stale, so a scraper got written that afternoon. A wholesaler only operated over WhatsApp, so the chat archive became a queryable feed. A parking question came up mid-negotiation, so aerial imagery got wired in before the call ended.
My favorite example, including the part where it went wrong: a weekly occupancy review had been overdue for two weeks because the numbers lived behind a co-living platform's host-dashboard login. Mid-conversation, the assistant reused the login flow from an existing scraper to go look, discovered the credentials on file belonged to an account with no listings, and asked me for the right ones. I pasted them into a config file it had prepared. Its first fix had a bug and logged into the wrong account again; I sent a screenshot of what I was seeing; it corrected the bug and pulled the full funnel: impressions, views, occupancy, day by day. The overdue review was filled and filed within the hour, and the one-off script is queued to become a permanent weekly command.
Count my contributions to that story: credentials and one screenshot. That is the pattern. Friction becomes capability in one sitting, and anything that recurs gets promoted from one-off script to standing command. The system gets more capable at exactly the rate the business needs it to, and never faster.
10. Everything is reversible or approved
The last pattern is the one that lets me sleep: there is no action in this system that is both irreversible and unapproved. Killing a deal is a stage move with a logged reason, not a deletion; every kill can be reopened with its history intact. Drafts persist until I use them or discard them. Outbound email requires me to name the recipient, and the harness will refuse to send to an address I never named, which has already caught one real mistake before it happened.
This is why I can let the assistant be aggressive everywhere else. It can reorganize, recompute, redraft, and refile as boldly as it likes, because the failure mode of boldness is a revert, not an apology call to a seller. If you get this pattern right, you stop needing to supervise the work and only need to gate the exits.
Why this deliberately stays at Level 3
On the five-level ladder, this whole operation is a Level 3: an assistant wired into everything, run by an operator who shows up every day. It does not pass the vacation test, because I have built it so it cannot. Nothing outward happens without me.
I want to be precise about why, because it is not caution about the technology. Building Level 4 and 5 systems is literally my day job. I could put offer-sending, follow-up sequences, and deal-killing on full autopilot in a week, and I have chosen not to, because of where my edge actually is. My business is relationship-driven deal-making. The product is judgment: the tone of a negotiation, the read on a seller, the decision to stretch on price for the right structure. Level 4 and 5 move authority to exactly the places where my edge is human.
So the human loop is not a limitation I have failed to engineer away. It is the point. The system exists to make sure that when I spend judgment, I spend it on a full brief, on time, with nothing forgotten, and that I spend it on ten decisions an afternoon instead of two hundred emails.
Your ceiling may be different. If your volume is high and your transactions are commodity, the math changes and the upper levels start paying. That is a genuine strategy question, not a technology one, and it is the most valuable question in this whole subject: which parts of your operation are judgment, and which parts only look like judgment because nobody wrote them down?
The stack, for the technically curious
Claude Code on a desktop machine as the session runtime, with background jobs for long work. A local-first pipeline system in TypeScript: deal records, underwriting engine, eval suites, a small dashboard. Connectors into Gmail, Calendar, Drive, Notion, and Slack. A command-line tool for the WhatsApp archive, browser automation for anything behind a login, and a scheduled task for the nightly digest. Plain files and git underneath all of it.
Nothing exotic. The leverage is not in any component; it is in the patterns above, and eight of the ten would work the same way on a different model or a different toolchain. That portability is pattern 2 doing its job.
Want to know what your version of this looks like?
This system fits my business: my deal flow, my edge, my tolerance for tinkering. Yours would look different, and the interesting work is figuring out where. That is what the AI Workflow Audit is: two weeks, fixed price, and a prioritized map of which patterns pay in your operation and which to skip.
New to the series? Start with the map, then Level 1 done properly.
Running something like this yourself? Email me what you kept human and why. I read all of it.