Funal← Journal
AIAgentsAI AssistantService Business

How Much Should an AI Assistant Be Allowed to Do? Hire It Like a Human One

OCT 6, 2026·12 min read·Alex Boquist

An AI executive assistant for a service business should earn autonomy like a human one: suggest, draft, act with undo, then act alone. How we built it, and why the AI never sets the order of work.

How much should an AI assistant be allowed to do? My answer, after building one for a real team and watching them use it: exactly as much as you'd let a new human assistant do in their first week. Suggest, draft, brief. Not decide. Definitely not send. Then it earns more, one kind of task at a time, and it earns it on a record you can actually look at.

That sounds obvious written down. It is not how most AI agent products behave, and it wasn't my first instinct either.

What does a great executive assistant actually do?

I keep coming back to this because it's the clearest bar I've found for what "AI assistant" should mean.

A great EA reads everything first. Most of it they handle, hand off, or ignore, and you never hear about it. What does reach you shows up as a decision that's already been worked out: "A or B? I'd go A, she mentioned Saturdays twice." You answer in one word. There's a brief before every call. After every call, every promise you made gets tracked and chased until it's actually done. Replies go out in your voice, after you glance at them. They remember who's touchy about what. They guard your calendar. The really good ones see things coming: "payroll closes Monday and three requests are still sitting there, want me to chase Jordan?"

Most of the value is the stuff that never reaches you. That's the whole job.

It's also, if you squint, exactly what every AI agent demo is promising. So when I sat down a few weeks ago with the admissions team at a coaching company that runs on Funal, that was the bar in my head. Not "can the model do this." Can it be that.

What does an admissions team's day look like without one?

Three people. Their entire day lived on one page: four work queues and a scoreboard. Call the new leads. Text back the families waiting on you. Confirm tomorrow's booked calls. Log what happened on the calls that just ended.

Watching them for a while, one thing jumped out. The button they hit more than anything else, by a mile, was Talked. Recording what happened was the actual job. And the page made them do the assistant's job before they could do their own: scan four queues, figure out which one mattered right now, find the row, then act. Nothing was triaged. Everything reached them.

So we built the dumbest possible thing. One screen, one item. What it is, why now, the family's text thread right there next to it, and the two or three buttons that finish it. Press a key, the next one shows up. You never pick a tab.

Which means the first thing the "assistant" does is decide what you see next, and in what order. And that's where it got interesting.

Should the AI decide the order of work?

No. And I want to be honest that I wanted it to.

The obvious 2026 move is to let a model build the order. Give it the queues, the calendar, each family's history, ask what the rep should do next. It gives a good answer. I built the demo. The demo was great.

We didn't ship it. Here's the thing I hadn't thought through: the order isn't a one-time ranking, it's a plan, and the plan gets rebuilt every time anything changes. Rep closes out a call, rebuild. Family texts back, rebuild. New lead applies, rebuild, and that lead has to be at the top right then, because your odds of reaching someone fall off a cliff within minutes of them filling out the form. Put a model in that path and:

None of that gets better with a smarter model. It gets worse, because you lean on it harder.

What decides the order instead?

One pure function. No I/O, tested like arithmetic. It asks one question: which of these loses the most by waiting another minute?

A new lead starts costing five minutes after they apply. A text starts costing fifteen minutes after it lands. A booked call needs confirming a day out. A call that just ended needs its outcome logged while the rep still remembers it. Stakes come from what the family actually did, not a priority number somebody typed: booked and replied beats a cold applicant. The calendar is a constraint, not a score; a call in ten minutes interrupts everything. And every rank explains itself in one line under the title. Applied 3 min ago. Call in 2 hrs, no reply yet.

An order nobody can explain is an order nobody trusts. That line ended up on the wall.

The rule's job is to be fast, cheap, boring, and always up. It's not where judgment goes.

So where does the AI assistant go?

Everywhere else, honestly.

It runs on events, not in the path that builds the plan. Lead comes in, text comes in, call ends. On each one it does what a good assistant does before they hand you anything:

So the assistant absolutely shapes what you see first. It just does it out in the open, through a field a person can read and a rule a person can check. It writes, the rule reads on the next rebuild. The assistant can be a few seconds behind. If it's slow, items are a bit dumber for a moment. If it's down, the plan still rebuilds and the day goes on.

That split turned out to be the whole design. The rule owns everything that has to be right every single time. The model owns everything that benefits from judgment and can afford to be a little late.

And this is the part I'm most excited about. "Text back Diane" is a to-do. "Diane asked about Saturdays. I drafted: we've got 10am open, want it? Send or edit" is a decision. The queue stops being a list of things to do and starts being a list of things the assistant needs from you. That's the EA.

How does the AI assistant learn from feedback?

A human assistant gets better because they watch what you do with their work. You rewrite the email, they notice. You send the next one untouched, they notice that too. A few months in the drafts need no edits, and that's when you stop reading them before they go out.

We're building the same loop, and the trick is the signal was already there. We'd just been throwing it away.

Every draft ends one of four ways: sent as is, lightly edited, rewritten, or dismissed. Log each one against the draft that produced it, with the text that actually went out. Every dismiss asks why. Every note the rep types on an item lands on the family's record (not the queue item, which is gone in a minute), so it shows up on the next item about that family and the assistant reads it before the next brief. Nobody's filling out a feedback form. They're doing their job, and the job is the training data.

Then the misses do the talking. Pull up the rewrites and they fall into piles. The assistant didn't know the answer: that's a missing entry in what it's allowed to say. It knew the answer and said it wrong: that's voice. It answered the wrong question: that's context it didn't have. Each pile is a specific fix. None of them is "make the model smarter." Fix the pile, and the sent-as-is rate moves.

And that rate is what unlocks the next rung.

When is an AI assistant allowed to act on its own?

When it's earned it, and only for that one kind of thing.

A new human assistant doesn't get your inbox and your signature on day one. They earn it, per task, over weeks, and you can feel exactly when it happens. We're making the AI climb the same ladder:

  1. Suggest. "I'd text her back about Saturdays."
  2. Draft. The text is written. You press Send.
  3. Act with undo. "Sent. Undo for ten minutes."
  4. Act quietly. It shows up in the done count.

An action type moves up one rung only on the record from that loop (say twenty drafts in a row sent without an edit), and the assistant asks before it climbs: "want me to just send these?" Your hand stays on the ladder. One confidently wrong text to a family costs more than fifty good ones earn, so the climb starts where mistakes are cheap and reversible: prep notes, internal stuff, close-outs. Texts to families stay at Draft for a long time. Probably longer than feels necessary. That's fine.

It can fall back down, too. The loop keeps running at every rung, and if the edit rate on some kind of action creeps back, it drops a rung and says so. That's what makes the top rung safe to reach at all: it's not a setting somebody flipped once, it's a state the evidence keeps earning.

The top rung is the whole prize. The assistant that handles things inside agreed limits without asking every time is the one whose value is mostly what never reaches you. It's also the rung that has to come last, because it's the one where a mistake is invisible until it isn't.

What I'd tell you if you're building this

Hire the AI the way you'd hire a person. Give it all the context and all the tools on day one; that part's easy now. Hold back the judgment, and let it earn each kind of decision on a record you can see, built from what people already do with its work. Keep it out of the one path that has to be right every time, and let it make everything around that path smarter.

The model was never the hard part. Deciding what it's allowed to decide is.

Frequently asked questions

How much should an AI assistant be allowed to do on its own?

Start with suggesting, then drafting for a person to send, then acting with an undo window, then acting quietly. Each kind of task climbs one rung only on evidence, such as twenty drafts in a row sent without an edit, and the assistant asks before it climbs. Mistakes that are cheap and reversible (prep notes, internal items) climb first. Messages to clients stay at draft for a long time.

Should an AI model decide what a rep works on next?

No. The order of work should come from a deterministic rule that rebuilds in milliseconds whenever something changes and explains each rank in one line. A model in that path makes speed to lead depend on a network call, shuffles the order between runs, costs tokens on every change, and goes stale when the model is down. The AI should enrich each item (a brief, a draft, a read on stakes) on events, not sort the queue.

What does a good AI executive assistant actually do?

The same things a great human one does: read everything first and handle most of it, bring you decisions already worked out, brief you before every call, track every promise until it is done, draft replies in your voice for approval, remember who is sensitive about what, protect your time, and anticipate what is coming.

How does an AI assistant learn from feedback without retraining?

Every draft ends one of four ways: sent as is, lightly edited, rewritten, or dismissed. Each outcome is logged against the draft with the text that actually went out, and every dismiss asks why. Rewrites sort into piles (missing answer, wrong voice, wrong context), each a specific fix. The sent-as-is rate unlocks the next level of autonomy, and it can fall back a level when the edit rate returns.

Why is speed to lead the reason not to put AI in the ranking path?

The odds of reaching a new lead fall off within minutes of them applying, so a new lead has to land at the top the moment it arrives. If a model builds the queue, that moment depends on an inference call to someone else's servers. A pure ranking function puts it at the top instantly, every time, whether or not the model is up.