AI agent PoCs: how to set them up so they reach production
40% of agentic projects will be cancelled by 2027. How to frame an AI agent PoC with success criteria, realistic costs and AI Act compliance built in.
An agent that books appointments, an agent that answers customers, an agent that reconciles invoices. In the boards I sit on the question is always the same: should we run a PoC? The right answer exists, but it depends on how the PoC is framed, on what it has to prove, and on who decides what happens next.
Because the problem is rarely getting started. 71% of large Italian companies have active AI projects, only 9% have structured governance (Osservatorio Artificial Intelligence, Politecnico di Milano, February 2026). PoCs get done, plenty of them. Then they sit there, in that limbo we call pilot purgatory: demonstrations that worked once, in front of the right people, and that nobody ever carried further.
40% of agentic projects will be cancelled by 2027
The estimate comes from Gartner, June 2025: over 40% of agentic AI projects will be cancelled by the end of 2027 due to escalating costs, unclear business value or inadequate risk controls. Anushree Verma, an analyst at Gartner, states that most agentic projects today are experiments or proofs of concept driven by hype, and often misapplied.
MIT, through the NANDA project, reached an even harsher conclusion in the summer of 2025: 95% of generative AI pilots produce no measurable impact on the P&L.
Well, these numbers describe something precise, something we see every week in our assessments: PoCs fail when they are born to impress someone instead of informing a decision. And a PoC that serves no decision is a cost dressed up as innovation.
A PoC without upfront criteria is a demo
The difference between a demo and a proof of concept lives entirely in a document almost nobody writes: the success criteria, fixed before a single line of code.
For an AI agent those criteria cannot stop at "it works". An agent makes decisions autonomously, touches real systems, produces output someone will use, so it has to be measured on dimensions a demo never shows: task completion rate on real cases rather than hand-picked examples, behaviour on edge cases, cost per run (tokens, API calls, human supervision time), how often a person has to step in, what happens when the agent gets it wrong.
In our model we formalise it this way: the PoC starts with real data and real context, with success criteria defined upfront, and it produces three possible outputs, scale, fix, stop. The prototype is not a goal, it is a decision-making tool. If at the end of the PoC you cannot say what production would cost and under which conditions it would be sustainable, the PoC gave you applause and no answers.
There is also a preliminary test we always recommend, because it filters out half of the requests: does this use case actually need an agent? Many processes are solved with traditional automation or with a well-integrated assistant, at far lower cost and risk. Gartner notes that a large share of use cases positioned as agentic do not require agentic implementations at all. Choosing an agent where a workflow would do is the first economic mistake, and it gets made before anything starts.
Which framework does your company actually need?
AI Rating measures maturity across the four areas of the model and shows where to start, with priorities and estimated effort.
Start your AI RatingThe economic mistake: the PoC that drifts into production
The typical scene: the PoC lands well, the sponsor wants speed, and the team ships the same architecture the PoC ran on. Nobody has estimated steady-state costs, nobody has verified the integration with ERP and CRM, nobody has planned the monitoring.
Six months later operating costs sit at two or three times the estimates, because a PoC optimises for speed of demonstration rather than efficiency of operation. An agent that costs a few cents per run in a demo changes order of magnitude in production, once you add retries, wide context windows, multi-model orchestration and human supervision.
A well-built PoC already carries the antidote: among its success criteria sits a realistic estimate of production cost and complexity. If the PoC proves the use case works but would cost more than the process it replaces, that is a successful PoC with a negative outcome. It prevented a bad investment, which is exactly its job. Our value proposition includes saying "stop", whenever the numbers say so.
The regulatory mistake: treating the PoC as a free zone
Here we see the most widespread naivety. "It's just a PoC" gets used as an implicit exemption: real customer data inside a prototype, no register, no risk assessment, vendors picked without checks.
A PoC using real personal data is data processing in every legal sense, and the GDPR grants no discount for experiments. And under the AI Act risk classification belongs at the beginning, at the end it is too late: if the use case you are prototyping falls among high-risk systems (recruitment, credit, access to essential services), the choices made during the PoC, from datasets to human oversight, determine what compliance will cost afterwards. The updated calendar set by Regulation 2026/1744 moves the obligations for Annex III high-risk systems to 2 December 2027, with the new prohibitions applying from 2 December 2026: decision windows, for anyone prototyping right now, far more than administrative deadlines.
For agents the issue amplifies, because an agent acts: it writes to real systems and makes micro-decisions on your behalf. The human oversight the regulation requires for relevant cases has to be designed into the prototype, together with action logs and the ability to intervene, because adding it later means redesigning. In our AI Rating this is the Risk dimension, and it contains what we call critical gates: non-compensable parameters, where technical excellence elsewhere never offsets a gap in compliance or ethics. A company with class-A delivery and no AI system register stays in class C, and the same logic applies, identically, to PoCs.
Rating first, prototype second
All of this explains the sequence of our method, which at first looks slow to some and then turns out to be the fastest route to actual production.
First you measure: the AI Rating captures maturity across four dimensions (Readiness, Delivery, Risk, Confidence) and tells you whether the organisation has the conditions to carry an agent beyond the prototype. An agent PoC inside a company with inaccessible data, unclear ownership and no monitoring practice will produce a nice demo and nothing else, and knowing that beforehand costs a fraction of discovering it afterwards.
Then you decide: which use case, with which success criteria, with which stop threshold. Then you validate, and that is the job of protot.ai, our fast prototyping unit: prototypes in one to three weeks, MVPs in four to eight, on real data, with the explicit mandate of putting the hypothesis to the test before production. protot.ai tests, ZeroFive builds in production and governs: the separation is deliberate, because whoever validates must have no interest in selling the implementation, otherwise every PoC miraculously ends up "working".
Across the 34-plus assessments and projects we have run, the sharpest correlation is this one: the PoCs that reach production are the ones born inside a decision path, with an executive sponsor who has already established what to do with the outcome, whatever it turns out to be.
The question to ask, before funding your next agent PoC, comes well before choosing a model or a vendor: if the prototype worked, would you be able to take it to production and govern it? If the answer is uncertain, the first PoC to run is the one on your own organisation.