Back to blogApprofondimenti

    AI agents in the enterprise: what they are, where they work, where they don't

    AI agents plan and execute actions, not just answers: the criteria for seeing where they make sense, where they don't, and what changes for governance when systems act.

    ZeroFive.AI July 1, 2026Updated on September 18, 2026 4 min

    The word of the year in the boardrooms we frequent is "agents", and like every word of the year it arrives wrapped in a fog worth clearing before signing budgets. An AI agent, stripped of the marketing, is a system that receives an objective instead of a question, decomposes the work into steps, uses tools, access to systems, search, code execution, corporate APIs, and carries the sequence forward correcting itself along the way, with variable degrees of autonomy from human intervention. The difference from the chatbot everyone knows lies in one verb: it does not merely answer, it acts. And it is that verb, not the technology itself, that changes the terms of the evaluation.

    The criterion separating the good cases

    Anyone selling agents will show impressive demos, and the demos will be real: the technology has made genuine strides, and production cases exist. The useful question for a decision maker is not whether agents work, it is where their properties meet the properties of the process, and the happy encounter has four recognisable conditions.

    Bounded perimeter: the agent performs well where the domain is circumscribed, the tools are few and known, the exceptions have a defined way out, and it performs badly in the open spaces where common sense must substitute for rules. Verifiable outcome: the agent's work must be checkable at low cost, by a machine or a human, because an output that costs more to verify than to produce erases the gain. Reversible error: the decisive question is what happens when it fails, and the answer must be "little", because the agent classifying tickets produces at worst a re-sort, while the one touching payments or customer communications produces damage with the same autonomy with which it produces value. Human checkpoints at the irreversible points: the designs that work place approval where the action stops being cancellable, not everywhere out of anxiety nor nowhere out of faith.

    The resulting map bears little resemblance to the proclamations: the first solid returns sit in the back office and operations, ticket triage and enrichment, reconciliations, case-file preparation, data and documentation maintenance, the same territories where the MIT research (2025) found the best ROI and the smallest budgets. At the opposite extreme, the agent with a broad mandate over high-stakes processes and irreversible actions is, at the state of the art, a bet no serious production case yet justifies.

    When the system acts, governance changes grade

    There is a point the enthusiastic plans skip, and it is the one the project's survival depends on: a system that acts is not governed like a system that answers. Identities and permissions become the heart of the design, because the agent operates on systems with credentials of its own, and the least-privilege principle stops being IT hygiene to become the difference between a contained incident and a structural one. The log of actions, not just of answers, becomes the evidence audit and regulators will ask for, and the stopping conditions, dear to readers of this series, here turn literal: thresholds that halt the agent on their own, action budgets per session, switches somebody knows where to find.

    The regulatory classification also needs redoing with fresh eyes, because autonomy shifts the risk profile: the same task that was harmless when assisted can, made autonomous, touch different categories and obligations, and the systems register must say who answers for the agent's actions with the same clarity it does for an employee.

    The sequence stays the same, with one extra test

    For the rest, agents deserve no special method: they deserve the usual method applied without discounts. The use case goes through prioritisation like the others, the PoC answers feasibility on the golden set, the MVP verifies real usage, with one additional test that for agents is mandatory, the red teaming of the perimeter: what the system does when facing hostile inputs, failing tools, ambiguous instructions, and how long a human takes to notice. In our chain this lives between PROTOT.AI for validation and AI Shift for controlled production, with perimeters and checkpoints signed before the first day of autonomy: calendly.com/fabiolalli/zerofive, or hello@zerofive.ai.

    The test for the proposals on your table fits in three questions to the proposer: what is the most damaging action the agent can take on its own, who would notice and how fast, how is it undone. If the answers come back precise, you are talking to someone who has designed an agent. If they come back vague, you are talking to someone who has seen a demo.

    Want to discuss this for your company?

    30 minutes with us to figure out where to start, or an AI Rating to measure your starting point.

    #enterprise AI agents#agentic AI#automation with AI agents#AI agent governance#AI agents use cases
    Share

    Keep reading