Is Anybody Actually Reading This?

Published:

Is anybody actually reading this? It is a fair question to ask of most of what a software team produces now. The ticket, the pull request description, the Slack reply, the review comment: each arrives promptly, reads well, and looks as though a careful person made it. Increasingly, an AI tool made it. Whether a careful person then read it is harder to tell than it used to be.

It used to be easier. A tidy ticket or a considered reply took real effort, and that effort made it a reasonable proxy for understanding. If someone had written it well, they had probably worked it out. Generating a fluent, well-formatted, plausible version now costs almost nothing.

As far as I can see, many organisations still run on the old assumption: if the ticket is closed, the pull request approved and the reply sent, then somebody understood something. That assumption is what I want to argue with. Not the people who hold it, because it was a sensible shortcut for a long time, but the shortcut itself.

My hypothesis is that it no longer holds, and that the first casualty is the feedback loop. Every loop is made of communication, and a loop with nothing real passing through it looks exactly like one that is working. If that is right, the useful question for any practice your team follows is whether it creates or checks understanding, or only produces another artefact for nobody. Plenty of practices will pass that test and some will not.

What follows is a test you can apply, some things I have seen, and an honest account of what is still open.

None of this is an argument against AI tooling. It is useful, and most capable teams are right to use it. The trouble lies elsewhere: in leadership, incentives, process fit, and the quiet removal of time to think, read and own the work.

When Effort Stopped Meaning Anything

Economists have a tidy way of describing this. In Michael Spence’s model of job-market signalling, as summarised on Wikipedia, a signal can be trusted when it is cheaper to produce for a capable person than for a less capable one. A credential works because the wrong people cannot afford to fake it.

Everyday work runs on informal versions of the same idea. A careful reply showed you had read the question. A detailed ticket showed you had thought about the edge cases. A thoughtful pull request description showed you knew what you had changed and why. Nobody designed this as a system. It was simply true that producing a good artefact was hard, and so a good artefact was evidence.

Generative AI has lowered that cost for everyone, including people who do not understand the thing being described. Some of the clearest early measurements come from hiring rather than engineering. Anaïs Galdin and Jesse Silbert studied an online freelance platform around the launch of an AI writing tool, and found that employers became less able to pick out high-ability workers from their tailored applications (arXiv preprint, November 2025). A separate preprint by Jingyi Cui, Gabriel Dias and Justin Ye saw the link between cover letter alignment and callbacks weaken by about half after a similar tool launched, with employers leaning more on workers’ prior history (arXiv, September 2025). Both come from online labour platforms and neither is about software teams, so I treat them as an illustration of the mechanism, not proof that it holds in your engineering organisation.

In my experience, real communication is being displaced by generated responses in Slack, email, Jira tickets, code reviews and merge requests. A reply arrives quickly, reads well, and does not quite answer what was asked. A review comment is articulate and generic. A ticket is complete in every section and says very little.

Plenty of people use these tools for good reasons. Non-native English speakers, people who find writing hard, and people who simply want to tidy their own thinking are all using them legitimately, and none of them should stop. The line I draw is this: using AI to express your own thinking is fine, and using it to replace your thinking is the problem. A good test is whether the sender could explain and defend the message unaided. If they could, the tool saved them time. If they could not, nothing has been communicated, however good it looks.

The same logic makes output metrics harder to trust. DORA’s own guidance warns that when a measure becomes a target, teams will tend to game it, which is Goodhart’s law in action. A metric that counts tickets closed, pull requests merged or documents produced now has a very cheap way of going wrong, and it needs no bad intent to get there.

Friction Used to Do Some of the Work

This is my own framing, and a hypothesis, not a finding.

AI does not create incompetence so much as remove the friction that used to expose it early. When writing a design note or a pull request took real effort, gaps in understanding surfaced quickly, to the author and to their colleagues. Now fluent output arrives without the effort, and the gap stays hidden until something breaks in front of a customer. Discovery takes longer, and the blast radius is bigger by the time it comes.

What I see here is a problem of systems and incentives, not of people. An organisation that rewards visible output and allows no time for reading will get unexamined output from capable people just as readily as from anyone else. If you are a leader who has noticed this, the useful question is what your system makes easy, not who is at fault.

The research is suggestive at best. The 2025 DORA report, built on nearly 5,000 survey responses and funded by Google, describes AI as an amplifier of whatever a team already is, which fits. At the individual level, Anthropic ran a randomised study of 52 relatively junior engineers learning an unfamiliar Python library. The group using AI scored 50% on a comprehension quiz against 67% for the hand-coders, with the largest gap in debugging, and the ones who asked for explanations kept more of what they learned (Anthropic, January 2026). The study is small, the quiz came straight afterwards, and Anthropic is an AI vendor studying its own category, so weigh it accordingly.

A Fast Loop and a Starved One

The original insight behind agile was about feedback loops. Dave Thomas put it more plainly than most in a 2014 essay: work out where you are, take a small step towards where you want to be, learn from what happened, and repeat. Short iterations existed to shrink the time between a decision and finding out whether it was right. I have written elsewhere about how much of the practice has drifted from that, in Are We Rearranging AI Deck Chairs?, so I will not repeat it here.

That insight is still the right lens. What has changed is which loops get the attention. I think many organisations are optimising one loop hard, the one from intent to generated code, and starving the ones that carry understanding:

  • Read and comprehend. Someone reads the change, understands it, and can say why it is shaped the way it is.
  • Validate with users. Real people try the thing and tell you whether it solves their problem.
  • Observe in production. You find out what the system does, as opposed to what you intended it to do.
Four horizontal loops. The generate-code loop is drawn thick, solid and green. The read-and-comprehend, validate-with-users and observe-in-production loops are drawn thin, dashed and amber, to show they often receive far less time and attention.

One loop is heavily optimised while the three that carry understanding are left thin. Illustrative, not data.

Every one of these loops is made of communication: a question asked, an answer given, a review comment, a reply. If the question is generated, the answer is generated, and the person in the middle pastes one into the other, the loop still runs. Messages move, tickets close, approvals are recorded. But nothing flows through it, because no understanding was put in at either end.

The DORA 2025 findings fit this picture. AI adoption is now associated with higher delivery throughput, a change from the year before, and still with lower delivery stability, with the same survey-based caveats as above. DORA’s own capabilities catalogue lists working in small batches among the practices that counter the risk of instability as AI accelerates development, because it shortens lead times and speeds up feedback.

One engineer’s anonymous account, which I cannot verify, will sound familiar to many teams. They described joining a large company where specs, code, tests, tickets and reports were all produced with an AI coding agent, and engineers at every level spent their days prompting and approving. Management said that pushing code was not the bottleneck while asking why delivery was slow, and nobody had time to read what was being produced.

I do not know how typical that is, so I take the shape and not the detail: a generate loop that is extremely fast and a comprehension loop given no time at all. The human cost is real too. Losing ownership of the work, and the pride that comes with understanding it, is not a small thing to lose.

Then there is the verification asymmetry. Producing plausible output is now nearly free. Checking it still costs a person’s time and attention, and that part does not scale. I think review depends on craft: the judgement that tells you what good looks like is much the same judgement you would need to write the thing, and it is hard to review well what you could not have written yourself. Alberto Bacchelli and Christian Bird’s study of code review at Microsoft found that although developers expected review to find defects, it also delivered knowledge transfer and team awareness, and that understanding the change was the central difficulty. If review becomes approval of something nobody understood, it loses the part that mattered.

Review as comprehension is not complicated to try. The author explains the change in their own words, the reviewer explains it back before approving, and nobody is asked to review more in a day than they can actually read.

The resulting gap has been given names. Margaret-Anne Storey calls it cognitive debt, building on Peter Naur’s idea that a program is a theory held in its developers’ heads, though her evidence is anecdotal, drawn from a student team. Jason Gorman’s comprehension debt describes what accumulates when code is generated faster than anyone can understand it.

Who Reads This, and What Does It Change?

Much of the process we inherited was built when implementation was the scarce, expensive resource. Story points, velocity and sprint capacity planning are all ways of managing that scarcity. When a change takes hours rather than weeks, the scarcity has moved: to understanding the problem, deciding what to build, verifying the result, and learning from production. Ron Jeffries, who thought he might have invented story points, wrote in 2019 that he was sorry, and recommended slicing work thin instead.

What I use is a pair of questions you can ask of any practice or artefact:

  1. Who or what consumes its output? A human, an agent, a test, a board, or nobody.
  2. What decision does it change?

A practice earns its place when it helps one of those consumers understand something or act correctly. It is the same idea as asking whether a sender could defend their message unaided, applied to a practice instead of a person. If nobody consumes it, or consuming it changes nothing, it is a candidate for deletion or redesign. The answers will differ from team to team, which is the point. Here is how the questions land on a few common practices. Treat the right-hand column as a starting position for a conversation, not a verdict.

Practice Who or what consumes it Decision it changes Starting position
Story point estimates and velocity Planning meetings, managers Sprint scope, sometimes a forecast Often a number nobody acts on, and easy to game once it is compared. Rethink it, but see the next section before deleting it.
Standups The team Who helps whom today Earns its place for coordination. Weak when it only relays status the tooling already shows.
Tickets written as human-readable requirements A person, or an agent through the prompt What gets built and in what order Check where the real working spec lives. If it is the prompt, the ticket may have no reader.
Long pull request descriptions that restate the diff A reviewer Nothing the diff does not already show Rethink. Say why the change exists and what was considered, not what changed.
Tests as executable specification Agents, CI, reviewers Whether a change can merge, and whether behaviour is right Likely to matter more. In my experience it is the main way to check what an agent produced.
Architecture decision records and written rationale New humans and agents Whether to repeat or revisit a past choice Likely to matter more. See documentation as decision history.
Small batches with real user feedback Users, product owners What to build next Likely to matter more. It feeds the loop that is most often starved.
Code review Reviewer, author, team Whether to merge, and who understands the change Reframe as comprehension and ownership, not approval.

Some warning signs that a team’s loops may be starved:

  • Approvals arrive faster than anyone could plausibly have read the change.
  • Nobody can explain, unprompted, why a recent change is shaped the way it is.
  • Tickets and pull request descriptions get longer while the questions asked about them get fewer.
  • Throughput looks healthy, while time to user feedback or incident frequency does not improve.
  • Questions are answered with fluent text that does not quite address them.
  • Leadership asks why delivery is slow while the output metrics look fine.

No single one of these proves anything. Several together suggest the loop is running without anything flowing through it.

Check the Fence Before You Pull It Down

It would be easy to read the table above as a deletion list. Please don’t.

G.K. Chesterton described the problem in The Thing (1929). A reformer finds a fence across a road and says he cannot see the use of it, so let us clear it away. The wiser reply is that if you cannot see the use of it, you should not clear it: go away and think, and come back when you can say what it is for.

Estimation is the example I keep returning to. I have argued before that agile estimation mostly fails, and I still think the numbers mostly do. What I would add is that, on good teams, the number was often never the valuable part. The argument about whether something was a three or an eight was often where someone pointed out that it depended on a system nobody had checked yet, and an unknown surfaced before anyone had built anything. Delete the points and you may delete the only moment the team talked about what could go wrong. If you remove a ritual, find out what it was doing and put that somewhere on purpose.

Telling load-bearing practices from habit takes experience, and it is hard to do from inside a team, because everything looks equally normal when you live with it. That is the honest case for an outside perspective, and for mentoring people who have not yet seen enough teams to tell the difference.

It is also why I distrust universal lists, including any I might write. A standup that is vital for one team is dead weight for another. A ticket template that forces useful thinking in one place produces filler in the next. Every team is different, and the test above is meant to be run by the people who live with the answer.

What I Have Seen Work, and What Is Still Open

These are observations from my own experience, and your teams may differ.

Templates and agent skills that generate important documentation, such as product definitions, user stories and epics, and then break the result into granular work items tracked on a board, have been working well. There is a caveat that I think matters more than the result. They earn their place only when they force a human to supply intent, and when their output is consumed by an agent, a test or a board. Otherwise they just produce more cheap artefacts that nobody reads, which is the problem this whole post is about.

I am also seeing Kanban work more effectively with AI agents than Scrum has. I have some hypotheses about why. Continuous flow and work-in-progress limits may suit work that completes in hours, and sprint boundaries may simply add latency. Yuval Yeret makes a related argument: that WIP limits keep agent output from piling up in front of the humans who still have to review it. He is reasoning from experience and offers no data, and I could not find solid evidence either way. This is not “Kanban good, Scrum bad”. A team that does Scrum thoughtfully, with its loops intact, may be doing better than one that does Kanban as a label.

And then there is what I do not know. I do not know what estimates should become when implementation takes minutes. I do not know how junior engineers build judgement when the first draft is free, though the Anthropic result above suggests how they use the tools matters a great deal. I do not know what review should look like at agent volume, or how much of a team’s communication can be generated before it stops being communication. I do not think anyone knows yet, including me, and I am wary of anyone who says they do.

Questions to Ask This Week

If you lead a team, or work closely with one, these cost nothing to ask:

  1. Pick one artefact your team produces, such as a ticket, a pull request description or a status update. Ask the person it is written for when they last made a different decision because of it.
  2. Of the last ten changes merged, how many had a person who could explain them unaided and defend them? Not approve them: explain them.
  3. How long is the gap between a change merging and anyone seeing it in a user’s hands or in production?
  4. Who has time on their calendar to read code they did not write?
  5. Where does the spec an agent actually works from live, and did a human have to supply the intent in it?
  6. If you stopped one recurring ritual for a fortnight, what would break, and who would notice first?

None of these needs a tool. They need someone to ask them honestly and listen to the answers.

A Second Pair of Eyes

If it would help to have someone look at this with you, I am happy to talk.

I work with organisations of up to about 100 people, helping teams and leaders work out how to operate now, through process and communication consulting and through mentoring. The first step is a free 15 to 30 minute introductory conversation about the problem you are seeing, where I give you my initial thoughts and proposals. There is no obligation to go any further.

If it makes sense to continue, there are two usual shapes. One is a fixed-price audit of one or two weeks, depending on team size, which looks at process and communication: which feedback loops are broken, which artefacts have no readers, and which practices are doing real work and which are habit. The other is a monthly retainer for ongoing mentoring and advisory work. Even a lightweight second opinion tends to be valuable.

If any of this sounds like your team, get in touch.

The artefacts are cheap now. Understanding is not, and it is worth protecting on purpose.

If you read this far, thank you. That is one loop closed.


About the Author

Tim Huegdon is the founder of Wyrd Technology, a consultancy that helps engineering teams and leaders work out how to operate in the AI era, through process and communication consulting and mentoring. With over 25 years of experience building and maintaining software and software engineering teams, he works with organisations of up to about 100 people.

Tags:Agile Methodology, AI, AI Adoption, Code Review, Continuous Improvement, Engineering Leadership, Feedback Loops, Future of Work, Human-AI Collaboration, Knowledge Management, Operating Model, Team Communication, Technical Leadership, Testing Discipline