Agentic Attention Surface
A dozen agents running at once. One is finished, one is stuck, one has been quietly wrong for an hour. They all want something, and you cannot read every diff to work out which one needs you.
I spent a day building a dashboard for it. Scores, quadrants, the whole apparatus. It was wrong, and the way it was wrong was the answer.
The board sorted work by how open the decision was and what being wrong would cost. Someone pointed out that two corners meant to be opposites looked identical. They were β both measures were asking do I need a human, one question wearing two labels.
The pair that does come apart: can a machine do it, and can a machine check it? A database migration is easy to do and hard to be sure about. A flaky test is miserable to do and trivial to verify. Between them they draw the only line that matters β everything a machine can both do and check leaves your day entirely.
What is left is the agentic attention surface. It is the only part worth an interface, and much smaller than the dashboard I spent the day on.
ELI5 tells it through a pair of glasses. Technical walks the day, and says where the idea is still soft.
More of the work: Portfolio
ELI5
The glasses on my face
Picture a dozen helpers working in other rooms. Every few minutes one of them shouts. Some are finished, some are stuck, one has been confidently doing the wrong thing since breakfast and does not know it yet.
You cannot go and inspect all of their work. That is the whole reason you have helpers. So you need some way of deciding, from outside the room, which shout to walk toward.
I spent a day building a wall of dials for this. One dial per helper, a number on each, so I could lean in and see who needed me. It felt like the obvious thing to build. It was the wrong thing, and I found that out three times in one day, each time by discovering I already had the answer.
The first one was in a note I had written myself
Two days earlier I had written down what this thing was allowed to be. One page. It said that if a feature does not come back to something a person actually does, it is not part of the product.
That note was open on the same screen all day, and I did not read it. Someone else asked, in about ten seconds, why what I was building looked nothing like what I had described. The note answered it immediately.
A test you never run is not a test. It is a note.
The second one was in an old script
My dashboard had one way of getting my attention: when a helper finished, it interrupted me.
That afternoon I happened to open a little script I had written months earlier, for something completely unrelated. In it was a comment reminding me: only interrupt when a helper is stuck. Never when one has finished β the screen already shows that.
Tapping someone on the shoulder to point at something already in front of them is not a notification. It is an interruption with good manners.
The third one was on my face
By evening the dials had become a map. Every job placed by two measurements, so I could see at a glance which ones needed me. The two measurements were how open is this decision and how bad is it if we get it wrong.
Someone looked at the map and said two of the corners were identical, and they were supposed to be opposites.
They were identical. Both measurements were quietly asking the same question β do I need a person for this β so the map had one direction pretending to be two. I had spent the day drawing a second axis that was a reflection of the first.
The fix came out of changing a single word on a label. Not a machine can finish this, but this needs a person to check it. Which is not a wording change at all. It is a different question, and it had been sitting inside the first one the entire time.
There were always two questions, and they come apart cleanly:
- Can a machine do it?
- Can a machine check it?
Changing a database is easy to do and hard to be sure about. A flaky test is miserable to do and trivial to check. Neither question predicts the other, which is exactly what the first pair of measurements failed to be.
And together they draw the only line worth drawing. Everything a machine can both do and check disappears from your day. It does not need a dial, a card, a column or a dashboard. What is left over β the work a machine can do but nobody can verify, and the work no machine can do at all β is the only thing a person has to look at.
That leftover pile has a name now. It is the agentic attention surface, and it is a great deal smaller than the wall of dials.
I had been searching my desk all day for something that was on my face.
Thatβs the kind of problem I like β the one where the answer was already on the machine and the job was noticing.
See my portfolio β Iβd love to work with you βTechnical
How the second axis turned out to be the first one
The problem is triage under a load you cannot personally inspect. Run enough agents in parallel and the constraint stops being how fast they work and becomes which of their outputs is worth your attention β decided without reading the diffs, because reading the diffs is the cost you were trying to avoid.
I spent a day building the obvious answer to that and got it wrong in a way that was worth more than getting it right.
Building from the wrong reference
I built from screenshots of other tools, and inherited their shape without noticing. Their organising principle is where work is in a process: a card moves through columns, and you read the board to see flow.
That is a task tracker. A task tracker answers what is happening. The question here is what deserves me, and those are not the same question β the first is a property of the work, the second is a relation between the work and a person.
I had written that distinction down two days earlier, in a planning note that was open the whole time. It went unread until someone asked why what I was building did not look like what I had described.
A check you donβt run is a check you donβt have.
The notification, stated plainly
The thing had one way of getting my attention: it interrupted me when an agent finished.
Then I opened an old script of mine, written months earlier for something unrelated, and found a comment saying: only interrupt when an agent is stuck. Never when it finishes β the screen already shows that.
Interrupting someone to point at something already in front of them is not a notification. Two independent lines arriving at the same place β a design judgement made in advance by someone solving a neighbouring problem, and a measurement made without knowledge of it β is about as much agreement as a small idea ever gets.
Two axes, one axis
By evening the surface had become a map: every unit of work placed by two scores, so the ones deserving a person would separate visually from the ones that did not.
The scores were how open is the decision and what does being wrong cost.
Two quadrants that were supposed to be opposites rendered identically. They were identical. Both scores reduced to do I need a human, so the map was one-dimensional wearing two axes, and the quadrant labels papered over it. When the axis was later corrected, all four labels stayed inverted relative to the geometry, because nobody re-derived them β the arithmetic was right, the logic was right, and every visible word was backwards.
The repair came from rewording one label: not a machine can finish it but this requires human verification. That is not a rewording. It is a different axis, and it had been folded inside the first one all along.
- Can a machine do it?
- Can a machine check it?
These are independent in both directions, which is the test the original pair failed. A schema migration is cheap to execute and expensive to verify. A flaky test is expensive to execute and trivial to verify. Neither answer constrains the other.
What the pair actually gives you
Cross them and one cell empties itself. Work a machine can do and check needs no interface at all β no card, no column, no dial. It should leave the human system entirely rather than being displayed more efficiently, which is what every dashboard I had been copying was doing.
What remains is two kinds of thing: work a machine can perform but not verify, and work no machine can perform. That remainder is the agentic attention surface. Its inverse β the part machines absorb completely β is the automation surface, and every unit of work you move onto one comes off the other.
The practical consequence is that the interface got much smaller. Most of what I built that day was a more efficient way of showing work that should never have reached a person, and the correct move was deletion rather than better layout.
Rocks, briefly
The metaphor that stuck arrived late and is not decoration. Work is not cards. Work is rocks: a lump with some mass, faceted where it breaks cleanly, foggy where nobody can see inside. You hit one and it fractures into smaller ones.
One thing fell out of that picture which a card would never have produced. Draw the break lines as facets, and a task with no natural seams renders as a smooth blob with nowhere to cut β which is exactly how an underspecified task feels to work on. Nobody designed that; it came out of the physics of the drawing.
The discipline underneath
The reason any of this got caught is a habit rather than a tool: every decision written down carries the specific observation that would prove it wrong, recorded at the moment of deciding, by the person who wants it to be right.
That habit caught several defects in a single day. It also explains why the failures above are embarrassing rather than unlucky β the checks existed, in writing, and nothing scheduled the reading of them.
Where this is still soft
The claim is that the do/check pair is the axis that matters, and that a machine-checkable task should leave the human system entirely. Two things would falsify it.
If checkability turns out to be a spectrum rather than a property β a test suite that catches most regressions but not the interesting ones β then the clean cell I am deleting is not clean, and the design becomes a threshold problem rather than a boundary. I think this is the likeliest failure and I do not have a good answer to it yet.
And the more uncomfortable one: I have not yet had a week where the corrections arrived from outside rather than from something already on my machine. Until then, βthe answer was already in reachβ is a claim about one day of my own work, not a general finding about building with agents. It is worth what one day is worth.
Thatβs the kind of problem I like β the one where the answer was already on the machine and the job was noticing.
See my portfolio β Iβd love to work with you β