# Agentgates Inference — pay-per-request models on attested GPUs (x402)

Ask open models running in GPU TEEs. One request, one USDC payment, and the
price is deterministic BEFORE the model runs: input charged by byte, output
charged at the `max_tokens` ceiling you set. No signup needed — the paying
wallet is the customer.

## Endpoints (all on https://agentgates.ai)

- `GET /api/inference/docs` — models, live prices, and the full contract. Read this first; prices move with gas and are never cached here.
- `GET /api/inference/catalog` — models and prices, machine-readable JSON.
- `POST /api/inference/chat` — ask. Body: `{"model","messages","max_tokens"}`.
- `POST /api/inference/credits` — prepay a balance so one on-chain settle covers many requests.

## The x402 loop

1. POST with no `X-PAYMENT` header → `402` with an `accepts` array; each
   entry is one network's own offer with the exact amount for YOUR body.
   Quoting is free and repeatable.
2. Sign that entry's amount as an EIP-3009 TransferWithAuthorization.
3. Retry the same request with the signature in `X-PAYMENT`. The answer
   carries the completion and the settlement.

## Cheaper at volume

Prepay a balance (`POST /api/inference/credits`) and then ask with an
`X-Wallet-Auth` header and NO payment header — the settle happens once for
the whole top-up. A wallet can also mint revocable bearer API keys on its
prepaid balance, so plain `Authorization: Bearer` clients work; the mint
contract is in `GET /api/inference/docs`.

## Rules that keep you safe

- The output ceiling is charged whether used or not — ask for what you need,
  and give reasoning models room to think.
- A serve failure is a free replay, never a refund.
