AI-generated editorial illustration for 'Working with Intelligent Machines'

Working with Intelligent Machines

Categories: Adaptivity, Agency, AI

I’m writing Agency Is All You Need in public — every chapter ships as an essay, gets argued with by readers, then revised into the final book (Q1 2027).

Read the series introduction.

The pilot team sits on the fourth floor, in a corner of the building that used to house the print shop. Eight desks, a kanban board nobody updates, and a name that appeared fully formed in a steering-committee deck: Claims Excellence Lab. Sonja — mid-thirties, insurance economics degree, ten years of doing everything right at a large insurer whose restructuring memos have started deleting the career ladder above her — was “seconded” here in March, in a conversation that lasted eleven minutes and never once mentioned that her old unit is shrinking by a third.

The work is this: the insurer’s new claims assistant — a large language model wrapped in workflow software — drafts coverage decisions for property claims. Sonja reviews them. Thirty to forty a day, each one arriving as a neat package: claim summary, policy check, recommended decision, and a customer letter, fluent and courteous, ready to send.

The drafts are good. That is the first thing that unsettles her. They are better than what half her old team produced on a Friday afternoon — no typos, no missed deadlines, no mood. For two weeks she catches herself clicking approve, approve, approve, her attention thinning by mid-morning into something closer to supervision of a conveyor belt.

Then, on a Thursday, claim 4471: water damage in a rowhouse in Leverkusen, a burst washing-machine hose, homeowner away for a long weekend. The draft recommends paying, minus a 40 percent reduction for gross negligence — the tap left on, the standard clause, the standard deduction. The letter explaining it is a small masterpiece of firm politeness. Sonja is reaching for approve when something older than the software tugs at her: the policy number has the format of the Komfort line. She opens the actual policy document and finds it in section four — an endorsement, sold with exactly this product since 2019, waiving the gross-negligence reduction entirely.

The model missed it. Not loudly — confidently. The draft doesn’t say I am uncertain about endorsements. It says, in flawless German, that the reduction applies, and cites the general clause as if the endorsement did not exist.

She sits back. Forty drafts a day. She has been fully reading, what, six of them? She pulls up her approval log for the past two weeks: 412 approvals. How many 4471s are in there?

She spends her lunch break not eating, going back through a sample of thirty. She finds two more — one harmless, one that shorted a customer €3,100. She reports all three; the lab lead, a friendly man from group IT, says “great catch” and adds a ticket to the backlog.

That evening she writes one line in her notebook: The machine didn’t fail today. I almost did.

What’s left when intelligence is cheap

For your entire working life until now, intelligence was the scarce input. Analysis, drafting, summarizing, calculating, first-pass judgment — these were expensive because they came bundled with a human, and the bundle took decades to produce. Your education, your credentials, your career were all priced against that scarcity.

That pricing is collapsing. Machine intelligence is becoming abundant — not perfect, not trustworthy the way a calculator is trustworthy, but abundant: competent-sounding cognitive work on demand, at near-zero marginal cost. And when any input becomes abundant, value migrates to whatever remains scarce around it. So the question that matters for your working life is not “will AI take my job” — a framing that keeps you passive, waiting for a verdict. The question is: in a process where machine intelligence is plentiful, what is still scarce?

Five things.

Judgment. The machine produces answers; it does not know which answers can be lived with. Judgment is the capacity to evaluate output against reality — against the policy document, the customer’s actual situation, the way this will read in court or in the newspaper. Kahneman and Tversky spent careers showing how unreliable human judgment is under uncertainty, and the lesson cuts both ways: judgment is hard, biased, effortful — and still the thing that decides whether fluent output is correct output. Fluency and correctness have been decoupled. Judgment is what re-couples them.

Taste. Adjacent, but distinct: knowing what good looks like in your domain. The machine can produce twelve versions; taste is knowing which one is right, and why the other eleven are subtly off — the tone that will land with this client, the analysis that answers the question actually being asked. Taste is compressed experience. You acquired yours slowly, through years of seeing good and bad work side by side. It is now worth more than the production skill it used to ride along with.

Framing. A machine answers the question it is given. The scarcest move in most knowledge work has quietly become deciding what the question is — turning a vague situation (“customers are complaining more”) into a tractable problem (“has our denial rate shifted in these two segments since the rollout, and is the letter wording driving escalations?”). Bad framing plus brilliant machine execution yields brilliant answers to the wrong question, at scale and with confidence.

Orchestration. Real work is not one prompt. It is a chain: decompose the problem, route the pieces to the right tool or person, set quality bars, integrate the results, notice when a step went wrong before the error propagates. This is management — except the workers are tireless, instant, occasionally wrong, and never push back. Someone who can run five machine-assisted workstreams and keep the whole coherent does the work of what used to be a team. That someone is a role, and it is being staffed right now, mostly by whoever steps into it first.

Responsibility. The machine cannot sign. When the decision is wrong, someone owns the consequences — legally, professionally, morally — and no deployment contract transfers that to the model vendor. This sounds like a burden and is actually the foundation of the other four: the willingness to say this goes out under my name is exactly what forces judgment, taste, framing, and orchestration to stay sharp.

Notice what these five have in common. None of them is a new technical skill — and every one of them is a form of agency, the capacity to direct your own course, applied here to work. Which brings us to the trap.

The tool that thinks for you

In Education is Broken I described a pattern among students given AI tutors. Used with agency, the tool deepened understanding: the student interrogated it, made it explain, used it to test themselves. Used without agency, the tool simply did the thinking, produced the essay, solved the problem — and the student, holding polished output, genuinely believed they had learned something. Same tool, same access, opposite outcomes, and the divide invisible from outside.

The identical mechanism now operates on working adults, and it is more dangerous there, because working adults are busy, evaluated on output, and long past the age where anyone checks whether they understand their own deliverables.

Here is the mechanism, slowed down. Every time you delegate a piece of thinking to a machine, two things happen: the output arrives, and your own capacity for that kind of thinking goes unexercised. The first is visible and immediately rewarding. The second is invisible and compounds. Skills you don’t use decay quietly — and with them decays the very judgment you need to evaluate the machine’s output. This is the cruel recursion: verification depends on the capability that delegation erodes. Sonja could catch claim 4471 because fifteen hundred manually processed claims live in her fingertips. The reviewer hired after her, trained only on reviewing machine drafts, will have no such reservoir. What happens when nobody in the loop has ever done the underlying work?

And the loss announces itself as its opposite. Your output improves — faster, more polished, more of it. Your metrics look better. You feel more productive. The hollowing shows up only in the moments the machine can’t cover: the novel case, the client who asks a probing question, the day the tool is down or wrong, the interview for your next job. By then the gap between your apparent output and your actual capability has been widening for two years, and you may be the last to see it, precisely because the tool has been papering over it daily.

So the same instrument deepens or hollows, and the difference is not the instrument. It is whether you remained the one doing the directing. That difference can be made concrete.

Delegation versus abdication

Delegation is what a good manager does with a capable report: define the task, set the standard, review the work, own the result. Abdication is handing the task over and looking away. With human colleagues the difference is obvious. With machines it blurs, because the output always looks reviewed — clean, confident, complete.

Four questions draw the line, and they work as a habit, not a policy. Could I have done this myself, even slowly? Delegating work you understand is leverage; delegating work you couldn’t evaluate is faith. Some faith is unavoidable — you can’t master everything — but know which mode you’re in, and never confuse the two under your own name. Did I specify, or just accept? Delegation begins with your intent: the goal, the constraints, the standard. If your prompt was vague and you took the first plausible answer, the machine framed the problem — which means the thinking that mattered was done by a system that doesn’t know your situation. Did I verify at a depth proportional to the stakes? A brainstorm needs a sniff test; a number going to a customer or a regulator needs to be checked against the source. If everything gets the same glance, you are not reviewing — you are ritualizing. And last: would I sign it? Not “does it look fine,” but: would I defend this line by line if challenged? If the honest answer is no, the review isn’t finished, whatever the deadline says.

Abdication usually fails all four at once, and feels like efficiency while it does.

Verification as a craft

Verification sounds like drudgery — the tax on the magic. Reframe it: verification is where your judgment gets its exercise now, the deliberate practice that keeps the reservoir full. A few habits make it a craft rather than a chore.

Learn the failure modes of your tools the way a sailor learns local weather. Current systems fail in characteristic ways, weakest exactly where confidence and correctness diverge: plausible fabricated specifics — citations, clause numbers, figures — silently dropped exceptions (Sonja’s endorsement), smooth interpolation across gaps, and agreement with whatever premise your question smuggled in. Generic distrust is useless; calibrated distrust — knowing where to poke — is fast and effective.

Spot-check on a schedule, not on suspicion, because suspicion arrives only after fluency has already lulled you. Decide in advance: everything above a stakes threshold gets full verification; below it, a fixed sample — one in ten, say — gets traced back to sources completely. The sample is not mainly about catching errors. It is about keeping you calibrated on the tool’s current error rate, which changes with every model update.

Keep a manual core: a slice of your craft you continue to do unassisted — one analysis from scratch a month, one document written cold. Athletes lift weights they will never lift in competition. Same logic. And log the catches: every error you find goes in a note — what, where, what pattern. Ten catches in, you have something valuable and rare, an empirical map of your tools’ weaknesses in your domain. Sonja’s notebook is becoming exactly this, and it is worth more on the market than another certificate would have been.

One more habit, easily forgotten: keep the whole relationship portable. You will get better with these tools the way you get better with a colleague — accumulated context, refined instructions, a sense of blind spots. Build that deliberately, but own it: keep your prompts, checklists, and workflows in plain text that belongs to you, and every few months run your core tasks on a rival system, to stay calibrated on what is general skill and what is platform habit. The platforms want your fluency to live inside their walls, so that leaving costs you your accumulated skill. This is optionality — keeping multiple viable next moves open at low cost — applied to your toolkit. Depending on AI is a decision you can make with open eyes. Depending on one company’s AI is a dependency someone else designed for you.

The edge, and the shadow of the edge

By June, something has shifted in the Claims Excellence Lab, and other people notice it before Sonja names it herself.

Her error log has forty-one entries and a taxonomy — categories, with rates. Her verification routine has become the team’s de facto standard; the lab lead photographed her checklist off the whiteboard and put it in a deck. When the model update in May quietly changed how the assistant handled multi-policy households, Sonja flagged the behavior shift on day two, from her one-in-ten sample, before the vendor’s release notes mentioned it. In the steering meeting, the department head — the same one who “seconded” her in eleven careful minutes — introduces her to a board member as “the person who keeps the machine honest.”

She is, she realizes on the tram home, good at this. Better than she was at the job the machine took. The orchestration, the calibrated distrust, the judgment calls at the edge of the rules — it uses everything she knows about claims and adds a layer that feels, for the first time in two years, like ground rising under her instead of sinking.

And she distrusts the feeling, because she can count. The pilot she is helping to succeed is the business case for shrinking her old department. Thirty of her former colleagues process claims the old way; the steering deck with her checklist in it projects that number at eight. Her new edge exists because the disruption is real — the same wave, and she is learning to surf it while people she trained are still standing where it will land. Her friend Meike asks her on Sunday whether that makes her feel guilty or safe. Sonja thinks about it honestly.

“Both,” she says. “Neither one all the way. I’m the person who checks the machine. For now the machine needs checking. I’m not going to pretend that’s a career. It’s a position. There’s a difference.”

She’s not sure yet what the difference demands of her. But she has stopped clicking approve without reading, in every sense, and she knows the notebook comes with her wherever she goes next.


Sonja is a constructed figure — a persona built to carry real mechanisms, not a case study of a real person. The book’s method note explains how the personas are used.

Scroll to Top