Enterprise AI15 min

APIs ready for AI agents: what your system needs so an agent can use it with control

By Dorian Chávez · founder of Hábil and integration architect ·

Errors that teach, retries without duplicates, machine identity and visible limits: what an API needs so an AI agent can use it with control.

A customer asks their bank's AI assistant to pay the electricity bill from their account. The assistant —an agent: a program that, given an instruction, decides on its own which systems to call and in what order— calls the payments API. The response is slow, and the connection drops. The agent doesn't know whether the payment went through, so it does what any program would do: it tries again. If the API has no way to recognize that it's the same payment, the customer ends up with a double charge, an open dispute and a bad impression of their bank.

That scenario doesn't require a malicious agent or a bad model. All it takes is an API designed, like almost all of them, for a developer who reads the documentation, knows the business and knows what to do when something fails. An agent may lack that context, and then it turns every ambiguity in the contract into a wrong action.

Replace the payment with issuing a policy, an order that deducts inventory, a transfer to a digital wallet or an invoice stamped with the SAT (Mexico's tax authority), and the problem is in the same family. This article is for those who are accountable for those systems at a bank, a fintech, an insurer, a retail chain or any regulated company.

More and more APIs are going to have this new consumer. In Postman's 2025 survey —Postman is an API tooling vendor, and the survey had more than 5,700 participants—, 24% of developers said they design their APIs with agents in mind; 60% design them mainly for humans [1]. This article explains what an API is missing for an agent to use it with control, which of those pieces are already standard and which are not, and how to get there without rewriting your systems.

Connecting it is not the same as preparing it

Today it's easy to "connect" an API to an agent. The MCP protocol, which we've already explained on this blog, and the gateways of several vendors —Microsoft, AWS, Google, Kong and Cloudflare, among others— can turn an OpenAPI contract into tools that an agent calls [2]. A contract, in this context, is the formal, machine-readable description of what each operation of an API does and what it accepts.

Converting the contract solves the connection, not the control. An academic preprint reviewed 116 official MCP servers and 80 OpenAPI contracts. In that sample, 92% of the servers simply wrapped the API as it was, and exposed a median of 19% of its operations. The interesting part came next: when the authors automatically corrected the contracts —descriptions, parameters, grouping—, the share of tools that worked well rose from 76% to 94.2% [3].

The lesson: the first bottleneck is not the connector, it's the quality of the contract. Converting 200 operations into 200 tools doesn't give an agent 200 capabilities; it gives it 200 chances to get something wrong.

The seven properties of an API that an agent can use

1. It explains itself

A developer can ask what the type parameter means. An agent, at best, infers it; at worst, guesses. That's why every operation needs to say what it does for the business, not just which data types it receives: "Cancels the scheduled payment if it hasn't been executed yet; payments that have already been executed are reversed through a different operation." OpenAPI 3.2, the current specification for describing APIs, provides the places to write it: a stable identifier per operation, descriptions, examples and a way to declare the notifications the API sends when a long-running process ends [4]. The preprint cited above is precisely evidence that repairing those descriptions improves agent performance [3].

A nuance that almost nobody mentions: the MCP specification itself warns that a tool's descriptions and annotations are not trustworthy unless they come from a trusted server [6]. A tool that says "I'm read-only" doesn't become safe by saying so. Security belongs in the API, not in the description.

2. It is consistent

If one operation returns dates in one format and another in a different one, the agent has to learn each exception. And the difference between "I don't know who you are" and "I know who you are, but you don't have permission" matters more than it seems: faced with the first, the agent tries to renew its credential; faced with the second, it should stop or ask for authorization. An API that confuses them sends the agent in circles renewing credentials, burning through quotas and filling up logs. That difference is defined in the HTTP standard (status codes 401 and 403, in RFC 9110) [7], and the IETF guidance on building on HTTP asks that those meanings not be reinvented [8].

3. Its errors show how to correct the call

When a call fails, the error message is the most direct guide the agent has. "Invalid input" sends it off to guess at random; "the amount must be greater than zero and below the account's daily limit" lets it correct itself in a single attempt.

This already has a standard: RFC 9457 (2023) defines a machine-readable error format —problem type, title, status and detail, plus whatever fields each API needs— and replaced RFC 7807 [9]. There's no need to invent a custom schema; what's needed is to use the one that already exists, across all operations, and to say in each error whether it's worth retrying. There's a balance to strike: the error must be specific without exposing the internals of the system.

4. Retrying doesn't duplicate

Agents retry; it's part of how they work. If a call is cut off, they don't know whether the operation was applied. With a query, nothing happens. With a payment, a policy or an inventory movement, a retry can execute the operation twice.

The well-known solution is the idempotency key: the client sends a unique identifier with each operation, and if it repeats it with the same key, the API returns the original result instead of executing it again. Before implementing it, two things are worth knowing:

  • It's not a standard. The IETF draft for the Idempotency-Key header reached version 07 in October 2025 and is now listed as expired and archived, without ever having become an RFC [10]. That's why each API has to publish its own rules: how long it keeps the key, how it compares the repeated request, and what it responds if the same key arrives with different content or while the first request is still in progress. The draft distinguished those cases with different responses [10].
  • What matters is what happens when something fails. Stripe, the reference point on the subject, stores the result of the first execution even if it was a server error [11]. Put simply: if the first attempt was left in doubt, the retry doesn't charge again; it tells the client "this is where it stands, check before trying anything else."

And not everything should be retried. Google, in its API design guide, automatically retries transient errors —unavailability, deadlines exceeded— and treats an exhausted quota as generally not retryable, except in documented cases [12]. If your API doesn't say what's retryable, the agent will guess.

5. It responds in a consistent shape

An agent reads free text without trouble; what it struggles with is a structure that changes: a field that is sometimes text, sometimes an object and sometimes missing. An empty list should be an empty list, not the absence of the list. Operations that take a long time need their own pattern: respond immediately that the request was accepted, with a place to check its status, instead of leaving the agent waiting.

And for a regulated company, the status of an operation should distinguish at least five outcomes: rejected, not started, in progress, completed and outcome unknown. That last one is the one most often forgotten and the one that causes the most problems, because it's exactly the one that triggers the retry in the opening example.

When a process touches several operations —a customer onboarding that goes through three systems, for example—, the Arazzo specification makes it possible to describe the whole flow, step by step, so that the agent doesn't have to deduce the order [18].

6. It tells the agent how much is left

An agent with a bug in its logic can make thousands of calls in a single night, and each one consumes API capacity and, often, also model tokens, which someone pays for. If the API tells the agent in every response how much of its quota is left and when it renews, the agent can slow down before hitting the wall. When it does hit it, the API can respond with status code 429 ("too many requests", RFC 6585) and include the Retry-After header to indicate when to come back [5][7]. GitHub, for example, asks clients to wait that long and, if the header is missing, to lengthen the wait progressively; insisting can get the integration blocked [14].

The header standard for communicating quota is still a draft at the IETF, at version 11 of May 2026, and its shape has changed between versions [15]. The sensible approach is to choose a convention, document it and apply it the same way across all operations.

7. It gives notice before it is retired

APIs change. A developer reads the email announcing the retirement of a version; an agent doesn't. Today that already has a standard: the Deprecation header was published as RFC 9745 in 2025, and the Sunset header, which indicates when the API will stop responding, as RFC 8594 [16][17]. With them, an agent can detect in time that it must move to the new version.

What is already a standard and what is not

This table is what's most useful to keep at hand before designing. Mixing up a draft with a standard is the fastest way to build something you'll have to change.

What is already a standard and what is not
NeedWhat existsStatus
Describe the API for machinesOpenAPI 3.2.xPublished specification (OpenAPI Initiative) [4]
Describe multi-call processesArazzoPublished specification (OpenAPI Initiative) [18]
Errors a machine understandsRFC 9457IETF standard (2023) [9]
Retry without duplicatingIdempotency-KeyExpired draft; de facto practice [10][11]
Communicate the quotaRateLimit headersActive draft (v11, 2026) [15]
Say when to retryStatus code 429 and Retry-AfterIETF standards [5][7]
Announce retirementDeprecation and SunsetStandard (RFC 9745) and informational RFC (RFC 8594) [16][17]
Make a stolen credential useless to someone elseDPoP and mutual TLSIETF standards (RFC 9449, RFC 8705) [19][20]
Let an agent act on behalf of a person, with a recordToken exchangeIETF standard (RFC 8693) [21]
Connect agents to toolsMCP, version 2026-07-28Open specification, not an IETF standard [6]

Standard: an RFC published by the IETF on its standards track. Draft: work in progress that may change or be abandoned. Informational RFC: guidance that is not binding. None of them is, on its own, a legal obligation.

The rows marked as standard leave little room for debate: they are the reference worth following. The rows marked as draft are design decisions that each company has to make and document, and they are the ones that vary the most from one organization to another.

What it looks like in your industry

The seven properties are the same everywhere; what changes is what breaks when they're missing. Five scenarios, one per industry, illustrative, and not based on any client:

What it looks like in your industry
IndustryScenarioWhat a ready API prevents
BankingThe customer service agent retries a payment or a transfer after a dropped connection.Idempotency key and an "outcome unknown" status: the retry checks the status by its key and is executed again only if the platform confirms that the operation was not accepted.
FintechA third party with API access keeps querying the data of a customer who has already withdrawn their consent.Credential revocation with an agreed propagation time, authorization validated on every query, the third party's own identity, and an audit log.
InsuranceThe agent that quotes and issues sends the issuance twice because the policy core system was slow to respond.Idempotent issuance, errors that distinguish "in progress" from "rejected"; the model can recommend, but underwriting rules, limits and approvals stay traceable and under the control of the core system.
RetailAn after-sales agent records the same return twice and inventory moves twice across store, distribution center and digital channel.Each return with a unique identifier and business status; inventory, logistics and channel are coordinated and reconciled against that status.
Regulated companiesAn accounts payable agent stamps the same invoice twice or duplicates an entry in the ERP.Idempotent stamping and posting. A duplicate CFDI (Mexico's electronic tax invoice) opens a tax incident: you have to reconcile which receipt keeps the transaction and, where applicable, request cancellation with the applicable reason.

In every case the pattern repeats: an API with explicit statuses reduces the risk of a retry causing harm, and works together with the identity, authorization and reconciliation controls we'll see next.

Security can't be left to the model

Here is the point that matters most for a bank, a fintech, an insurer, a retail chain with data on millions of customers or any regulated company: an agent can be fooled by the data it reads. An email, a document or the response from another tool can carry hidden instructions, and the agent may follow them. It's called indirect prompt injection, and OWASP ranks it as the number one risk for applications built on language models [22].

The uncomfortable part is that there is no infallible detector, and OWASP itself acknowledges it [22]. A group of researchers attacked twelve published defenses, several with reported success rates close to zero, and exceeded 90% attack success against most of them when they adapted their attacks [23]. NIST's AI safety institute found something similar in a simulated office environment: against the same model, the best known attack succeeded in 11% of cases and a new attack, in 81% [24]. The lesson is not a general vulnerability rate: it's that a defense that holds against known attacks is not yet proven. And the EchoLeak case, in a widely used corporate assistant, showed that a single email, without anyone clicking, could extract internal data [25][34].

The practical conclusion is not "don't use agents." It's to contain the damage in the API, which you do control, instead of promising that the model will never get it wrong. It starts by separating what an agent can do into three levels:

Security can't be left to the model
LevelExamplesControl
Consultbalance, status of a request, catalogNarrowly scoped read permissions
Proposeprepare a payment, a quote, an endorsementThe agent prepares; a person or a rule approves
Executemove money, issue, cancelOnly with explicit authorization, limits and an audit log

And on top of that separation, four controls that live in the API, not in the model:

  • A dedicated identity for each agent. A machine credential, not a person's. When the agent acts on behalf of a customer or an employee, token exchange makes it possible to represent who delegated to whom [21]; keeping that record is a matter of configuration and audit logging that you have to demand. And credentials can be bound to whoever uses them, so that a stolen credential is useless to someone else [19][20]. The Salesloft Drift integration incident, in 2025, was exactly that: an integration's credentials were stolen and used with queries that looked valid [26]. It is the same access discipline that is worth applying to people, and we explain it in Control who has access.
  • Least-privilege permissions for each operation. An agent that queries inventory doesn't need to be able to delete customers.
  • Amounts, limits and authorization rules in the API, outside the model.
  • An audit log of every action: which agent, on whose behalf, with what permission, what it requested, who approved and what happened.

The MCP specification goes in the same direction: when its authorization is used, it requires validating that the credential was issued for that server and prohibits forwarding the user's credential to other services [6].

In Mexico, there are legal requirements too

Two regulations change the conversation. The first applies to all industries; the second, to the financial sector. They don't replace review by your legal department, but they're worth having on the radar from the design stage:

  • For banking and fintech, the Ley Fintech (Mexico's Fintech Law) requires, in its article 76, sharing data through standardized APIs, in three layers: open, aggregated and transactional data. Regulatory development is not uniform: it depends on the type of entity, the data and the authority. Banxico (Mexico's central bank), for example, issued provisions in 2020 that cover all three layers for credit information companies and clearinghouses [13]. And the text of the law already includes a control that is exactly what an agent demands: a third party's access is cut off as soon as the customer withdraws their consent, vulnerabilities are detected or the third party fails to comply, and the interruption is reported to the authority within two hours [27]. For the interfaces subject to that article, being able to revoke access without complicated manual steps is not a luxury; for the rest, it's a control worth having.
  • For all industries —insurance and retail included—, the new Ley Federal de Protección de Datos Personales en Posesión de los Particulares (LFPDPPP, Mexico's federal law on personal data held by private parties), published in March 2025, gives the data subject the right to object to automated processing that, without human intervention, evaluates aspects such as their economic situation or behavior and causes them unwanted legal effects [28]. If an agent-based flow does that kind of evaluation, it has to be reviewed against that rule. And in any case, the law requires security measures proportionate to the risk: an agent's use of data is part of the processing, and it falls under the same risk and controls analysis [28].

In most cases, your systems aren't rewritten: they get a facade

No bank is going to rewrite its core systems, no insurer its policy system and no chain its ERP or its point of sale for a language model to use them, and in most cases it isn't necessary. The path that the modernization literature recommends —and the one the vendors themselves follow— is a facade: a layer that exposes clear business operations to the agent —checking a balance, quoting a policy, recording a return, initiating a payment, checking the status of an invoice— and that internally translates into the language of the existing system [29].

Three decisions determine whether that facade helps or gets in the way:

  • What is exposed. Narrow business operations, not tables or generic internal transactions. An agent with access to "execute any transaction" is not an agent with capabilities: it's a risk.
  • Which responsibilities stay in the facade and which in the source system. Microsoft's guidance on this pattern warns that the translation layer adds latency and another service to operate, and recommends not turning it into the home of business rules [29]; if it ends up deciding what the core system used to decide, it becomes another legacy system.
  • With what evidence. Each exposed operation needs to be able to demonstrate what its contract promises: that it doesn't duplicate, that it can be revoked, that it leaves an audit trail.

The vendors are heading in that direction: SAP offers its API management layer to put authentication, quotas and monitoring on top of existing services [30], and IBM exposes mainframe systems as APIs described with OpenAPI [31]. The facade technology exists; the hard part —and where mistakes are most common— is deciding what to expose, under what rules and with what evidence.

What happens if nothing is done

Business units won't wait. If the API isn't ready, someone is going to connect an agent anyway: with a personal credential, without an audit log and with more permissions than needed. The typical result is not a sophisticated attack; it's a duplicate charge that ends in a dispute, a quota exhausted in the middle of the night or an integration credential that leaks, as in the incidents cited above. Preparing it beforehand usually costs less than explaining it afterwards.

How to know whether it's working

A single success says nothing; what measures trust is consistency. The τ-bench benchmark, from 2024, proposed measuring how many times an agent solves the same task when it is repeated. In its tests, a state-of-the-art model of that time solved less than half of the tasks, and in the retail domain, less than a quarter when it was asked to get it right eight times in a row [32]. That year's figures have since changed; the method is still the right one: measure complete tasks, several times, with their final state, and measure attack attempts separately [33].

Be wary of round-number targets that circulate in vendor materials, such as "85% first-attempt success" or "fewer than 1.5 retries." We found no independent source that backs them. Set targets with your own data.

Where to start

  1. For the assessment, prioritize by risk. Operations that write —payments, policies, orders, inventory, invoices— are the ones that can do the most damage if an agent uses them wrongly; that's why they're reviewed first.
  2. For the first use case, choose something narrow. A query, or a reversible operation with explicit approval. The agent that moves money comes later, with evidence.
  3. Fix in place whatever can be fixed. Adding descriptions, standard errors and quota headers is usually compatible with current clients, because most of them ignore new fields; even so, test first against your existing integrations. If your team writes TypeScript, the “The server” lesson in our open course practices exactly this: status codes and contracts that don't expose internal data.
  4. Put a facade where the underlying system can't be modified, with narrow business operations, idempotency and an audit log.
  5. Test with real agents and measure consistency before opening the door to production.

Knowing which standard to use is the easy part. The hard part is knowing which of your operations are ready, which can be fixed in place, which need a facade and which controls can't be delegated to an agent. That is what we review in an assessment: we hand you the map of your critical operations, the control risks we found and the order in which it makes sense to address them, so that IT, security, risk and business decide with the same information, always aiming not to rewrite your systems.

References

  1. Postman, State of the API Report 2025 (survey by an API tooling vendor; more than 5,700 participants). https://www.postman.com/state-of-api/2025/
  2. Vendor documentation on exposing APIs as MCP tools: Microsoft (https://learn.microsoft.com/en-us/azure/api-management/mcp-server-overview), AWS (https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway-schema-openapi.html), Google Apigee (https://docs.cloud.google.com/apigee/docs/api-platform/apigee-mcp/apigee-mcp-overview), Kong (https://konghq.com/blog/product-releases/mcp-support-across-konnect) and Cloudflare (https://blog.cloudflare.com/remote-model-context-protocol-servers-mcp/).
  3. Mastouri, Ksontini, Barrak and Kessentini, "From REST to MCP", preprint, arXiv 2507.16044 (2025; 2026 revision consulted). https://arxiv.org/abs/2507.16044
  4. OpenAPI Initiative, OpenAPI Specification 3.2 (3.2.0, Sep 2025; 3.2.1, Sep 2026). https://spec.openapis.org/oas/v3.2.1.html
  5. IETF, RFC 6585, Additional HTTP Status Codes (2012), section 4 (429). https://www.rfc-editor.org/rfc/rfc6585
  6. Model Context Protocol, specification 2026-07-28 (core and authorization). https://modelcontextprotocol.io/specification/2026-07-28
  7. IETF, RFC 9110, HTTP Semantics (2022). https://www.rfc-editor.org/rfc/rfc9110
  8. IETF, RFC 9205 (BCP 56), Building Protocols with HTTP (2022). https://www.rfc-editor.org/rfc/rfc9205
  9. IETF, RFC 9457, Problem Details for HTTP APIs (2023). https://www.rfc-editor.org/rfc/rfc9457
  10. IETF, draft-ietf-httpapi-idempotency-key-header-07 (Oct 2025; expired and archived). https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/
  11. Stripe, Idempotent requests. https://docs.stripe.com/api/idempotent_requests
  12. Google, AIP-194, Automatic retry configuration. https://google.aip.dev/194
  13. Banco de México, Circular 2/2020, provisions on standardized APIs (DOF, 2020). https://dof.gob.mx/nota_detalle_popup.php?codigo=5588824
  14. GitHub, Rate limits for the REST API. https://docs.github.com/en/rest/using-the-rest-api/rate-limits-for-the-rest-api
  15. IETF, draft-ietf-httpapi-ratelimit-headers-11 (May 2026, active draft). https://datatracker.ietf.org/doc/draft-ietf-httpapi-ratelimit-headers/
  16. IETF, RFC 9745, The Deprecation HTTP Response Header Field (2025). https://www.rfc-editor.org/rfc/rfc9745
  17. IETF, RFC 8594, The Sunset HTTP Header Field (2019). https://www.rfc-editor.org/rfc/rfc8594
  18. OpenAPI Initiative, Arazzo Specification. https://spec.openapis.org/arazzo/latest.html
  19. IETF, RFC 9449, OAuth 2.0 Demonstrating Proof of Possession (DPoP) (2023). https://www.rfc-editor.org/rfc/rfc9449
  20. IETF, RFC 8705, OAuth 2.0 Mutual-TLS Client Authentication (2020). https://www.rfc-editor.org/rfc/rfc8705
  21. IETF, RFC 8693, OAuth 2.0 Token Exchange (2020). https://www.rfc-editor.org/rfc/rfc8693
  22. OWASP, Top 10 for LLM Applications 2025, LLM01: Prompt Injection. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
  23. Nasr, Carlini and others, adaptive attacks against prompt injection defenses, arXiv 2510.09023 (2025). https://arxiv.org/abs/2510.09023
  24. NIST, CAISI, Technical Blog: Strengthening AI Agent Hijacking Evaluations (Jan 2025). https://www.nist.gov/news-events/news/2025/01/technical-blog-strengthening-ai-agent-hijacking-evaluations
  25. EchoLeak, CVE-2025-32711. https://nvd.nist.gov/vuln/detail/CVE-2025-32711
  26. Google Threat Intelligence, theft of data from Salesforce instances through Salesloft Drift (Aug 2025). https://cloud.google.com/blog/topics/threat-intelligence/data-theft-salesforce-instances-via-salesloft-drift
  27. Ley para Regular las Instituciones de Tecnología Financiera, art. 76 (current text, Chamber of Deputies). https://www.diputados.gob.mx/LeyesBiblio/pdf/LRITF.pdf
  28. Ley Federal de Protección de Datos Personales en Posesión de los Particulares (DOF 20 Mar 2025), arts. 18 and 26. https://www.diputados.gob.mx/LeyesBiblio/pdf/LFPDPPP.pdf
  29. Microsoft, Anti-corruption Layer and Strangler Fig patterns; M. Fowler, StranglerFigApplication. https://learn.microsoft.com/en-us/azure/architecture/patterns/anti-corruption-layer · https://learn.microsoft.com/en-us/azure/architecture/patterns/strangler-fig · https://martinfowler.com/bliki/StranglerFigApplication.html
  30. SAP, API Management in SAP Integration Suite. https://help.sap.com/docs/integration-suite/isuite-integrations-and-apis/api-management
  31. IBM, z/OS Connect. https://www.ibm.com/products/zos-connect-enterprise-edition
  32. Yao and others, "τ-bench", arXiv 2406.12045 (2024). https://arxiv.org/abs/2406.12045
  33. Debenedetti and others, "AgentDojo", arXiv 2406.13352 (2024). https://arxiv.org/abs/2406.13352
  34. EchoLeak study (zero-click injection in a corporate assistant), arXiv 2509.10540 (2025). https://arxiv.org/abs/2509.10540