The short answer
Five limits hold for AI agents in 2026: they cannot exercise open-ended judgment without review, cannot be trusted alone with high-stakes irreversible actions, cannot succeed at work with no definable success condition, waste money on low-volume chaos, and cannot be accountable. None of this argues against agents. It argues for assigning them correctly, which is where the canceled 40 percent went wrong.
We run an agent in production and sell agent engagements, which makes this article a strange business decision and an easy editorial one: the limits are where the incidents live, and buyers who know them make better clients than buyers who discover them.
Why the limits matter commercially
The agent market is running ahead of agent capability. Gartner's forecast has over 40 percent of agentic projects canceled by 2027, on costs, unclear value, and inadequate risk controls, with roughly 130 genuine vendors among the thousands wearing the label. Most of those cancellations are not technology failures. They are assignment failures: agents given work from the list below, then blamed for being what they are.
The five limits
- –Open-ended judgment without review. Agents handle judgment at the edges of structured work superbly. Judgment as the core deliverable (what should our refund policy be, is this deal worth taking) produces confident, plausible, unaccountable answers. Plausible is the failure mode, not the feature.
- –High-stakes irreversible actions, alone. Money out the door, contracts accepted, data deleted, commitments made. An agent can prepare all of it; the click that cannot be unclicked stays human, by design and not by nostalgia.
- –Work without a success condition. If nobody can define what a good outcome looks like, no eval can check it, and an uncheckable agent is an incident with a delay timer. Evals are not bureaucracy; they are the boundary of what agents can safely hold.
- –Low-volume chaos. Twelve wildly different cases a month is a human's job description, not an automation target. Agents repay volume and structure; without them the build cost never returns.
- –Accountability. An agent cannot own an outcome, sit in the post-mortem, or carry the pager. Someone answers for the system: name them before launch, or the first incident names them for you.
Most agent failures are assignment failures: work from this list, handed over anyway.
Limits as design inputs
Every limit converts into a design rule, which is how our own agent stays boring in the right ways: the refusal list handles the judgment boundary (it will not quote prices, ever), escalation handles the stakes boundary (exceptional prospects route to a human with context), the brief format defines success so the golden conversations can check it, and a named owner reads the transcripts. The general pattern is in the agents guide: boundaries in the prompt, checks on the output, humans on the calls that matter.
What is moving
The limits are a snapshot, honestly dated. Models improve, and the structured-work boundary keeps expanding outward: more exception classes handled, longer task chains completed, better calibration about when to escalate. What has not moved, and shows no sign of moving on any commercial timescale, is the back half of the list: irreversibility, undefinable success, and accountability are not model-capability problems. They are reasons the assignment question, which machine for which work, outlives every model release.