# Inference — machine-paid inference, confidential and open tiers

Open models over one paid API, one request one payment, USDC over x402.
No signup, no API key needed: the paying wallet is the customer and the only
identity (a wallet CAN mint bearer keys on its prepaid balance — below).

TWO TIERS, and every model on this page names its tier:

- CONFIDENTIAL: runs in attested hardware (GPU TEEs / SEV-SNP enclaves). The
  attestation is public and served beside the price, and a model is served
  only while its provider's own live catalog affirms it (fail closed).
- OPEN: served by an open provider with NO attestation. Cheaper; pick it for
  price, never for privacy. Nothing on the open tier is sold as confidential:
  the tier is in this page, in the catalog JSON, and in the 402's own
  description, so what you approve at signing names what you get.

## Models (live prices — reread before paying)

- `gpt-oss-120b` (GPT-OSS 120B) — input $0.155/M tokens, output $0.618/M tokens. Context 131,072 tokens; prompt up to 412,876 bytes, output ceiling up to 16,384 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=gpt-oss-120b`.
- `qwen3-5-397b` (Qwen3.5 397B) — input $0.464/M tokens, output $3.09/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-5-397b`.
- `glm-5-2` (GLM 5.2) — input $1.288/M tokens, output $4.069/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=glm-5-2`.
- `kimi-k3-phala` (Kimi K3 on Phala) — input $3.09/M tokens, output $15.45/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=kimi-k3-phala`.
- `kimi-k3` (Kimi K3) — input $3.09/M tokens, output $15.45/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=kimi-k3`.
- `kimi-k3-chutes` (Kimi K3 on Chutes) — input $3.09/M tokens, output $15.45/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=kimi-k3-chutes`.
- `kimi-k3-near` (Kimi K3 on NEAR) — input $3.399/M tokens, output $16.995/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=kimi-k3-near`.
- `deepseek-v4-flash-0731` (DeepSeek V4 Flash 0731) — input $0.454/M tokens, output $1.36/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=deepseek-v4-flash-0731`.
- `llama-3-3-70b-instruct` (Llama 3.3 70B Instruct) — input $2.06/M tokens, output $2.06/M tokens. Context 131,072 tokens; prompt up to 412,876 bytes, output ceiling up to 16,384 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=llama-3-3-70b-instruct`.
- `qwen3-8-27b` (Qwen3.8 27B) — input $0.248/M tokens, output $2.266/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-8-27b`.
- `muse-glimmer-30b` (Muse Glimmer 30B) — input $0.309/M tokens, output $1.133/M tokens. Context 131,072 tokens; prompt up to 412,876 bytes, output ceiling up to 16,384 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=muse-glimmer-30b`.
- `qwen3-5-27b` (Qwen3.5-27B) — input $0.309/M tokens, output $2.472/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-5-27b`.
- `glm-5-3-flash` (GLM 5.3 Flash) — input $0.155/M tokens, output $0.515/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=glm-5-3-flash`.
- `nemotron-3-5-lightning` (Nemotron 3.5 Lightning) — input $0.073/M tokens, output $0.206/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=nemotron-3-5-lightning`.
- `glm-5-3` (GLM 5.3) — input $1.442/M tokens, output $4.532/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=glm-5-3`.
- `qwen3-6-27b` (Qwen3.6 27B) — input $0.33/M tokens, output $3.348/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-6-27b`.
- `qwen3-8-27b-uncensored` (Qwen3.8 27B Uncensored (Aggressive)) — input $0.309/M tokens, output $1.545/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-8-27b-uncensored`.
- `gemma-4-26b-a4b-uncensored` (Gemma-4 26B-A4B Uncensored (Heretic)) — input $0.155/M tokens, output $0.721/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=gemma-4-26b-a4b-uncensored`.
- `kimi-k2-6` (Kimi K2.6) — input $0.515/M tokens, output $2.936/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=kimi-k2-6`.
- `gemma-4-31b-it` (Gemma 4 31B) — input $0.155/M tokens, output $0.474/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=gemma-4-31b-it`.
- `qwen3-32b` (Qwen3 32B) — input $0.108/M tokens, output $0.429/M tokens. Context 40,960 tokens; prompt up to 129,024 bytes, output ceiling up to 5,120 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-32b`.
- `qwen3-6-35b-a3b` (Qwen3.6 35B A3B) — input $0.176/M tokens, output $1.133/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-6-35b-a3b`.
- `glm-5-1` (GLM 5.1) — input $1.01/M tokens, output $3.173/M tokens. Context 202,752 tokens; prompt up to 638,668 bytes, output ceiling up to 25,344 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=glm-5-1`.
- `deepseek-v3-2` (DeepSeek V3.2) — input $1.03/M tokens, output $1.03/M tokens. Context 163,840 tokens; prompt up to 516,096 bytes, output ceiling up to 20,480 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=deepseek-v3-2`.
- `qwen3-vl-30b-a3b-instruct` (Qwen3 VL 30B A3B Instruct) — input $0.155/M tokens, output $0.567/M tokens. Context 128,000 tokens; prompt up to 403,200 bytes, output ceiling up to 16,000 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen3-vl-30b-a3b-instruct`.
- `qwen-2-5-7b-instruct` (Qwen2.5 7B Instruct) — input $0.103/M tokens, output $0.206/M tokens. Context 32,768 tokens; prompt up to 103,219 bytes, output ceiling up to 4,096 tokens (default 1,024). Confidential (TEE-attested). Attestation: `GET https://agentgates-backend.vercel.app/api/inference/attestation?model=qwen-2-5-7b-instruct`.

Open tier (cheaper, NOT confidential: these run on ordinary provider
infrastructure with no attestation; pick them for price, never for privacy):

- `openai/gpt-5.5-pro` (GPT-5.5 Pro) — input $32.445/M tokens, output $194.67/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-sol` (GPT-5.6 Sol) — input $2.163/M tokens, output $10.815/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-plus-02-15` (Qwen3.5 Plus 2026-02-15) — input $0.282/M tokens, output $1.688/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3-flash-preview` (Gemini 3 Flash Preview) — input $0.541/M tokens, output $3.245/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `x-ai/grok-4.5` (Grok 4.5) — input $2.163/M tokens, output $6.489/M tokens. Context 500,000 tokens; prompt up to 1,684,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.7-flash` (Gemini 3.7 Flash) — input $0.812/M tokens, output $4.056/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `moonshotai/kimi-k3:batch` (Kimi K3 (batch)) — input $2.466/M tokens, output $12.33/M tokens. Context 1,048,576 tokens; prompt up to 3,715,891 bytes, output ceiling up to 16,384 tokens (default 1,024). Open tier: no hardware attestation.
- `mistralai/mistral-large-2512:batch` (Mistral Large 3 2512 (batch)) — input $0.271/M tokens, output $0.812/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `deepseek/deepseek-v4-flash-0731` (DeepSeek V4 Flash 0731) — input $0.015/M tokens, output $1.385/M tokens. Context 1,310,720 tokens; prompt up to 4,000,000 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-sol-pro` (GPT-5.6 Sol Pro) — input $2.163/M tokens, output $10.815/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.1` (GPT-5.1) — input $1.352/M tokens, output $10.815/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.1-codex` (GPT-5.1-Codex) — input $1.352/M tokens, output $10.815/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.1-codex-mini` (GPT-5.1-Codex-Mini) — input $0.271/M tokens, output $2.163/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.1` (GLM 5.1) — input $1.045/M tokens, output $3.284/M tokens. Context 204,800 tokens; prompt up to 645,120 bytes, output ceiling up to 25,600 tokens (default 1,024). Open tier: no hardware attestation.
- `deepseek/deepseek-v4-flash` (DeepSeek V4 Flash 0423) — input $0.007/M tokens, output $1.385/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-sonnet-5:batch` (Claude Sonnet 5 (batch)) — input $1.082/M tokens, output $5.408/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-sonnet-5` (Claude Sonnet 5) — input $2.163/M tokens, output $10.815/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `deepseek/deepseek-v4-pro` (DeepSeek V4 Pro 0423) — input $0.322/M tokens, output $0.644/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.2-chat` (GPT-5.2 Chat) — input $1.893/M tokens, output $15.141/M tokens. Context 128,000 tokens; prompt up to 403,200 bytes, output ceiling up to 16,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-luna-pro:batch` (GPT-5.6 Luna Pro (batch)) — input $0.109/M tokens, output $0.649/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `meta-llama/llama-4-maverick` (Llama 4 Maverick) — input $0.203/M tokens, output $0.706/M tokens. Context 1,048,576 tokens; prompt up to 3,715,891 bytes, output ceiling up to 16,384 tokens (default 1,024). Open tier: no hardware attestation.
- `meta-llama/llama-4-scout` (Llama 4 Scout) — input $0.109/M tokens, output $0.325/M tokens. Context 1,310,720 tokens; prompt up to 4,000,000 bytes, output ceiling up to 16,384 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5` (GLM 5) — input $0.649/M tokens, output $2.077/M tokens. Context 204,800 tokens; prompt up to 645,120 bytes, output ceiling up to 25,600 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-astra:batch` (GPT-6 Astra (batch)) — input $5.408/M tokens, output $27.038/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-astra-pro` (GPT-6 Astra Pro) — input $10.815/M tokens, output $54.075/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-astra-pro:batch` (GPT-6 Astra Pro (batch)) — input $5.408/M tokens, output $27.038/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `moonshotai/kimi-k3` (Kimi K3) — input $0.541/M tokens, output $12.978/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5v-turbo` (GLM 5V Turbo) — input $1.298/M tokens, output $4.326/M tokens. Context 202,752 tokens; prompt up to 638,668 bytes, output ceiling up to 25,344 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.7-flash:batch` (Gemini 3.7 Flash (batch)) — input $0.406/M tokens, output $2.028/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.8-flash` (Gemini 3.8 Flash) — input $0.812/M tokens, output $4.056/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.8-flash:batch` (Gemini 3.8 Flash (batch)) — input $0.406/M tokens, output $2.028/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-fable-5.1` (Claude Fable 5.1) — input $10.815/M tokens, output $54.075/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-fable-5.1:batch` (Claude Fable 5.1 (batch)) — input $5.408/M tokens, output $27.038/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `x-ai/grok-4.20-multi-agent` (Grok 4.20 Multi-Agent) — input $1.352/M tokens, output $2.704/M tokens. Context 2,000,000 tokens; prompt up to 4,000,000 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `deepseek/deepseek-v4.1-flash` (DeepSeek V4.1 Flash) — input $0.325/M tokens, output $1.298/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.3:batch` (GLM 5.3 (batch)) — input $0.487/M tokens, output $2.163/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-astra` (GPT-6 Astra) — input $10.815/M tokens, output $54.075/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-luna:batch` (GPT-5.6 Luna (batch)) — input $0.109/M tokens, output $0.649/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-terra-pro:batch` (GPT-5.6 Terra Pro (batch)) — input $1.082/M tokens, output $6.489/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-terra:batch` (GPT-5.6 Terra (batch)) — input $1.082/M tokens, output $6.489/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-sol-pro:batch` (GPT-5.6 Sol Pro (batch)) — input $1.082/M tokens, output $5.408/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-sol:batch` (GPT-5.6 Sol (batch)) — input $1.082/M tokens, output $5.408/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5-pro` (GPT-5 Pro) — input $16.223/M tokens, output $129.78/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `x-ai/grok-4.3:batch` (Grok 4.3 (batch)) — input $1.082/M tokens, output $2.163/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.5-pro:batch` (GPT-5.5 Pro (batch)) — input $16.223/M tokens, output $97.335/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.5:batch` (GPT-5.5 (batch)) — input $2.704/M tokens, output $16.223/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4-nano:batch` (GPT-5.4 Nano (batch)) — input $0.109/M tokens, output $0.676/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4-mini:batch` (GPT-5.4 Mini (batch)) — input $0.406/M tokens, output $2.434/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4-pro:batch` (GPT-5.4 Pro (batch)) — input $16.223/M tokens, output $97.335/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4:batch` (GPT-5.4 (batch)) — input $1.352/M tokens, output $8.112/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.2-pro:batch` (GPT-5.2 Pro (batch)) — input $11.356/M tokens, output $90.846/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.2:batch` (GPT-5.2 (batch)) — input $0.947/M tokens, output $7.571/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.1:batch` (GPT-5.1 (batch)) — input $0.676/M tokens, output $5.408/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5-pro:batch` (GPT-5 Pro (batch)) — input $8.112/M tokens, output $64.89/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5:batch` (GPT-5 (batch)) — input $0.676/M tokens, output $5.408/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `x-ai/grok-4.20` (Grok 4.20) — input $1.352/M tokens, output $2.704/M tokens. Context 2,000,000 tokens; prompt up to 4,000,000 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-oss-120b:batch` (gpt-oss-120b (batch)) — input $0.033/M tokens, output $0.148/M tokens. Context 131,072 tokens; prompt up to 412,876 bytes, output ceiling up to 16,384 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-opus-5` (Claude Opus 5) — input $5.408/M tokens, output $27.038/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-opus-5:batch` (Claude Opus 5 (batch)) — input $2.704/M tokens, output $13.519/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `deepseek/deepseek-v4-pro-0813` (DeepSeek V4 Pro 0813) — input $0.714/M tokens, output $2.142/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5-mini:batch` (GPT-5 Mini (batch)) — input $0.136/M tokens, output $1.082/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5-nano:batch` (GPT-5 Nano (batch)) — input $0.028/M tokens, output $0.217/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4-nano` (GPT-5.4 Nano) — input $0.217/M tokens, output $1.352/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4-mini` (GPT-5.4 Mini) — input $0.812/M tokens, output $4.867/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5-turbo` (GLM 5 Turbo) — input $1.298/M tokens, output $4.326/M tokens. Context 202,752 tokens; prompt up to 638,668 bytes, output ceiling up to 25,344 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-oss-120b` (gpt-oss-120b) — input $0.041/M tokens, output $0.184/M tokens. Context 131,072 tokens; prompt up to 412,876 bytes, output ceiling up to 16,384 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-9b` (Qwen3.5-9B) — input $0.109/M tokens, output $0.163/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4-pro` (GPT-5.4 Pro) — input $32.445/M tokens, output $194.67/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-27b` (Qwen3.5-27B) — input $0.282/M tokens, output $2.812/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.3-prime` (GLM 5.3 Prime) — input $3.029/M tokens, output $9.518/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.4` (GPT-5.4) — input $2.704/M tokens, output $16.223/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.6-flash` (Gemini 3.6 Flash) — input $0.812/M tokens, output $4.056/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.1-flash-lite-preview` (Gemini 3.1 Flash Lite Preview) — input $0.271/M tokens, output $1.623/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.5-flash` (Gemini 3.5 Flash) — input $1.623/M tokens, output $9.734/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-35b-a3b` (Qwen3.5-35B-A3B) — input $0.087/M tokens, output $0.812/M tokens. Context 262,144 tokens; prompt up to 884,736 bytes, output ceiling up to 16,384 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-opus-5.5` (Claude Opus 5.5) — input $4.326/M tokens, output $21.63/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-opus-5.5:batch` (Claude Opus 5.5 (batch)) — input $2.163/M tokens, output $10.815/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.5-flash:batch` (Gemini 3.5 Flash (batch)) — input $0.812/M tokens, output $4.867/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-122b-a10b` (Qwen3.5-122B-A10B) — input $0.282/M tokens, output $2.25/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-flash-02-23` (Qwen3.5-Flash) — input $0.071/M tokens, output $0.282/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.1-pro-preview-customtools` (Gemini 3.1 Pro Preview Custom Tools) — input $2.163/M tokens, output $12.978/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `mistralai/mistral-large-2407` (Mistral Large 2407) — input $2.163/M tokens, output $6.489/M tokens. Context 131,072 tokens; prompt up to 412,876 bytes, output ceiling up to 16,384 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.3-flashx` (GLM 5.3 FlashX) — input $0.401/M tokens, output $1.352/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5` (GPT-5) — input $1.352/M tokens, output $10.815/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5-mini` (GPT-5 Mini) — input $0.271/M tokens, output $2.163/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `deepseek/deepseek-v4.1-flash:batch` (DeepSeek V4.1 Flash (batch)) — input $0.122/M tokens, output $0.364/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5-nano` (GPT-5 Nano) — input $0.055/M tokens, output $0.433/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.3-codex` (GPT-5.3-Codex) — input $1.893/M tokens, output $15.141/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-luna-pro` (GPT-6 Luna Pro) — input $0.109/M tokens, output $0.541/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-luna-pro:batch` (GPT-6 Luna Pro (batch)) — input $0.055/M tokens, output $0.271/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-luna` (GPT-6 Luna) — input $0.109/M tokens, output $0.541/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-luna:batch` (GPT-6 Luna (batch)) — input $0.055/M tokens, output $0.271/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-sol-pro` (GPT-6 Sol Pro) — input $2.163/M tokens, output $10.815/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-sol-pro:batch` (GPT-6 Sol Pro (batch)) — input $1.082/M tokens, output $5.408/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-sol` (GPT-6 Sol) — input $2.163/M tokens, output $10.815/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-6-sol:batch` (GPT-6 Sol (batch)) — input $1.082/M tokens, output $5.408/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.6-flash:batch` (Gemini 3.6 Flash (batch)) — input $0.406/M tokens, output $2.028/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `x-ai/grok-4.7` (Grok 4.7) — input $2.163/M tokens, output $6.489/M tokens. Context 500,000 tokens; prompt up to 1,684,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.1-pro-preview` (Gemini 3.1 Pro Preview) — input $2.163/M tokens, output $12.978/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.1-pro-preview:batch` (Gemini 3.1 Pro Preview (batch)) — input $1.082/M tokens, output $6.489/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-fable-5:batch` (Claude Fable 5 (batch)) — input $5.408/M tokens, output $27.038/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.1-flash-lite` (Gemini 3.1 Flash Lite) — input $0.271/M tokens, output $1.623/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.1-flash-lite:batch` (Gemini 3.1 Flash Lite (batch)) — input $0.136/M tokens, output $0.812/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.5-flash-lite` (Gemini 3.5 Flash Lite) — input $0.325/M tokens, output $2.704/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `x-ai/grok-4.3` (Grok 4.3) — input $1.352/M tokens, output $2.704/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3.5-flash-lite:batch` (Gemini 3.5 Flash Lite (batch)) — input $0.163/M tokens, output $1.352/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `deepseek/deepseek-v4-flash-vision-exp` (DeepSeek V4 Flash Vision Exp) — input $0.234/M tokens, output $0.70/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.3` (GLM 5.3) — input $0.044/M tokens, output $5.192/M tokens. Context 1,310,720 tokens; prompt up to 4,000,000 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `x-ai/grok-4.6` (Grok 4.6) — input $2.163/M tokens, output $6.489/M tokens. Context 500,000 tokens; prompt up to 1,684,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.3-flash` (GLM 5.3 Flash) — input $0.163/M tokens, output $0.541/M tokens. Context 1,310,720 tokens; prompt up to 4,000,000 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `anthropic/claude-fable-5` (Claude Fable 5) — input $10.815/M tokens, output $54.075/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.2-codex` (GPT-5.2-Codex) — input $1.893/M tokens, output $15.141/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-plus-20260420` (Qwen3.5 Plus 2026-04-20) — input $0.325/M tokens, output $1.947/M tokens. Context 1,000,000 tokens; prompt up to 3,484,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-luna-pro` (GPT-5.6 Luna Pro) — input $0.217/M tokens, output $1.298/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `mistralai/mistral-large` (Mistral Large) — input $2.163/M tokens, output $6.489/M tokens. Context 128,000 tokens; prompt up to 403,200 bytes, output ceiling up to 16,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-luna` (GPT-5.6 Luna) — input $0.217/M tokens, output $1.298/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-terra-pro` (GPT-5.6 Terra Pro) — input $2.163/M tokens, output $12.978/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `google/gemini-3-flash-preview:batch` (Gemini 3 Flash Preview (batch)) — input $0.271/M tokens, output $1.623/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `qwen/qwen3.5-397b-a17b` (Qwen3.5 397B A17B) — input $0.595/M tokens, output $3.786/M tokens. Context 262,144 tokens; prompt up to 828,518 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.5` (GPT-5.5) — input $5.408/M tokens, output $32.445/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.2-pro` (GPT-5.2 Pro) — input $22.712/M tokens, output $181.692/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.2` (GPT-5.2) — input $1.893/M tokens, output $15.141/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.6-terra` (GPT-5.6 Terra) — input $2.163/M tokens, output $12.978/M tokens. Context 1,050,000 tokens; prompt up to 3,664,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.2` (GLM 5.2) — input $0.065/M tokens, output $4.543/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `z-ai/glm-5.3-flash:batch` (GLM 5.3 Flash (batch)) — input $0.065/M tokens, output $0.217/M tokens. Context 1,048,576 tokens; prompt up to 3,659,673 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.
- `openai/gpt-5.1-codex-max` (GPT-5.1-Codex-Max) — input $1.352/M tokens, output $10.815/M tokens. Context 400,000 tokens; prompt up to 1,324,800 bytes, output ceiling up to 32,000 tokens (default 1,024). Open tier: no hardware attestation.

…and 293 more open models, every one priced live in the catalog: `GET https://agentgates-backend.vercel.app/api/inference/catalog` (tags mark `frontier`, `reasoning`, `multimodal`).

Prices are read from the catalog at request time: `GET https://agentgates-backend.vercel.app/api/inference/catalog`.

## How a request is priced (the ceiling before the model runs, the tokens after)

- YOU PAY FOR TOKENS: the prompt tokens and the completion tokens the
  provider counts, at the model's input and output rates, plus one settle.
  The receipt says both counts, the rates, and the amount.
- THE CEILING is known before inference and is what a draw reserves or an
  x402 payment settles: the prompt's bytes (`utf8ByteLength(JSON.stringify(messages))`)
  divided by the bytes-per-token this model has shown on its served turns
  (`observed.reserveBytesPerToken` in the catalog; one byte per token until
  it has any, and no more than 3 until it has 20), plus `max_tokens` (the model's row default if you omit it)
  at the output rate. The charge never exceeds the ceiling.
- WHAT THE TOKENS DID NOT COST comes back: a draw releases it within reach
  the same moment; an x402 turn adds it to your balance
  (`charged.refundedToBalanceUsd`), spendable on the next turn, with no
  second chain write. Reasoning models spend part of the ceiling on a
  thinking pass before the answer; a ceiling too small can return
  `finishReason: "length"` with a `reasoning` field and little or no
  `content` — that is a served completion, billed at its tokens, so give
  reasoning models room.
- ONE ON-CHAIN SETTLE is carried by every x402 request, at the same margin as
  the tokens — from $0.001597 right now; a
  turn drawn from a balance writes nothing on chain and carries none.
  It is read from live gas at the moment you ask, never a constant, so the
  quote moves with the chain. Paying us costs us a transaction; that cost is
  in the price instead of coming out of it. The settle costs a different amount on each chain, so THE PRICE IS PER
  NETWORK: `base` $0.001597, `polygon` $0.002964, `solana` $0.010000.
  What another chain's gas is doing never moves the price on yours.
- The total snaps up to the next $0.001 and floors at $0.001000 per
  request. You never need to compute any of this: the 402 quotes the exact
  amount for your exact body.

## Prepay once, then ask many times

🚨 THIS IS A PREPAID BALANCE. You send USDC FIRST and spend it down. Nothing
is loaned in either direction: the transfer settles on chain before your
balance moves, and an empty balance serves nothing.

Every request you pay for individually carries one on-chain settle. Top a
balance up instead and that settle happens ONCE, for the whole amount.

1. `POST https://agentgates-backend.vercel.app/api/inference/credits` with `{"amountUsd": 5}`
   (`1`, `5`, `25` or `100`). The `402` asks for the amount plus the one
   settle that delivers it; `GET /api/inference/catalog` prices every rung.
2. Pay it the same way you pay for a request (below). The reply carries your
   balance.
3. From then on, `POST /api/inference/chat` with an `X-Wallet-Auth` header
   and NO `X-PAYMENT`. The ticket is drawn from the balance and the answer
   comes back with no chain write at all.
4. `GET https://agentgates-backend.vercel.app/api/inference/credits` with `X-Wallet-Auth` reads the
   balance and the ledger.

A balance never expires and is not refundable. It belongs to the wallet that
paid for it and cannot be transferred. There is a ceiling on how much unspent
balance the lane will hold at once, and on how much it takes in a day: read
`sale.availableUsd` off the catalog or your own balance and pick a rung
inside it, rather than being refused after you have signed. A balance short of a request's ticket is
never partly drawn: you get the ordinary `402` and can pay for that one
request or top up. A request that fails to serve spends nothing, so there is
nothing to replay — ask again.

## API keys (optional, for software that cannot sign wallet messages)

The wallet stays the only account. A key is a bearer handle the wallet mints
on its own prepaid balance — for an OpenAI-style client, a cron job, or a
teammate's script that holds no wallet code. Mint one on
`https://agentgates-backend.vercel.app/inference` (connect the wallet, one
signature), or over the API:

- `POST https://agentgates-backend.vercel.app/api/inference/keys` with `X-Wallet-Auth` and
  `{"name": "my-agent"}` — mints a key; the answer carries the secret ONCE
  and it is never stored or shown again.
- `GET https://agentgates-backend.vercel.app/api/inference/keys` with `X-Wallet-Auth` — the wallet's
  active keys (names and prefixes, never secrets).
- `POST https://agentgates-backend.vercel.app/api/inference/keys` with `X-Wallet-Auth` and
  `{"revoke": "<key id>"}` — kills a key instantly.

Then any request is one header:

`POST /api/inference/chat` with `Authorization: Bearer ag_…` draws the
wallet's balance exactly like an owner-signed draw — no chain write, no
`X-PAYMENT`. `GET /api/inference/credits` with the same header reads the
balance the key spends. An empty balance answers the ordinary `402`.

Several keys per wallet are fine (one per agent, up to 20 active); each is
named, listed, and revocable on its own. A key can ONLY draw the balance its
wallet prepaid — it signs no payments and never touches the wallet, so the
unspent balance is the most a leaked key can ever spend.

## Pay in one pass

1. `POST https://agentgates-backend.vercel.app/api/inference/chat` with JSON
   `{"model": "<model id>", "messages": [{"role": "user", "content": "..."}], "max_tokens": <int, optional>}`
   (`temperature`, `top_p`, `stop`, `tools` and `tool_choice` are
   passed through; a message may carry content parts, a `tool` role with its
   `tool_call_id`, or an assistant turn's `tool_calls`).

   `"stream": true` returns Server-Sent Events, the gateway's own chunks
   verbatim, and a last frame carrying the invoice and what it charged. A
   streamed turn is served from a PREPAID BALANCE, never from a per-request
   x402 leg, since that leg settles on chain before the model runs; a streamed
   request paid the x402 way refuses with `stream_unsupported_payer`.

   The tool sheet counts toward the prompt the ticket is priced over, and the
   prompt cap is per model: every row on this page names its own.

   TOOLS TAUGHT. A server that runs a model without tool parsing refuses
   `tools`; the lane then writes the tools into the system message in the
   format the model was trained on and returns its calls as ordinary
   `tool_calls`. Each model's format is measured, and a model whose own calls
   measured worse is taught on every turn. `"x_tool_format"` sets it for one
   request: `"native"` is the server's own tool parsing only; `"hermes"`,
   `"glm"`, `"llama"` or `"pythonic"` teaches the tools in that format. The
   last frame's `x_inference.taughtTools` names the format a turn was taught in.

   CACHED INPUT. An agent resends its whole session every turn, and the
   provider serves the repeated part from its cache. When the provider reports
   it (`usage.prompt_tokens_details.cached_tokens` in its own frames), a
   balance-drawn turn charges those tokens at the row's cached-input price,
   the provider's cached rate at the row's margin, and the rest of the prompt
   at the input price. The last frame's `x_inference.usage.cachedInTokens`
   and `charged.perMTokens.cachedInUsd` say both. A cached token is never
   charged more than a fresh one; a row with no cached rate read charges
   every input token in full.

   THE OPEN TIER'S HOSTS. An open row is served only by the hosts that sell
   it at the row's own price, always in the same order, so a session's turns
   share one cache and none lands on a dearer host. Its turn is charged what
   that host billed for the tokens it counted (cached input at its cached
   rate, input written into its cache at its write rate, a long prompt at its
   long-prompt rate), at the row's margin. `"reasoning_effort"` (`"low"` to
   `"max"`) is asked of that host, and `"prompt_cache_key"` names a session
   so its turns land on one cache. The last frame's
   `x_inference.usage.cacheWriteTokens` counts the written tokens.
2. The reply is `402` with an `accepts` array (x402, scheme `exact`).
   Networks offered: `base` (USDC 0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913, pay to 0x151c7cafb427264e34952c04273e019e9dc89ff6), `polygon` (USDC 0x3c499c542cEF5E3811e1192ce70d8cC03d5c3359, pay to 0x151c7cafb427264e34952c04273e019e9dc89ff6), `solana` (USDC EPjFWdd5AufqSSqeM2qN1xzybapC8G4wEGGkZwyTDt1v, pay to EhYCDMf6dQBAPphYNGRfW6zvPUG5gba2cttQzjJtPu2q).
3. Sign an EIP-3009 `TransferWithAuthorization` for `maxAmountRequired`
   USDC to `payTo` (EIP-712 domain in `extra`), and retry the SAME body with
   `X-PAYMENT: base64({"x402Version":1,"scheme":"exact","network":"<network>","payload":{"signature":"0x…","authorization":{"from","to","value","validAfter","validBefore","nonce"}}})`.
   On a `solana` / `solana-devnet` leg, build a transaction of exactly
   [ComputeBudget SetComputeUnitLimit, ComputeBudget SetComputeUnitPrice,
   spl-token `TransferChecked` of `maxAmountRequired` USDC (`asset`, 6
   decimals) to `payTo`'s associated token account, optional Memo], set the
   fee payer to `extra.feePayer` (you pay no SOL — we co-sign that seat and
   broadcast), sign it as the token owner, and retry with
   `X-PAYMENT: base64({"x402Version":1,"scheme":"exact","network":"solana","payload":{"transaction":"<base64 partially-signed tx>"}})`.
   Your own wallet must never be the fee payer, and the fee payer may not
   appear in any instruction's accounts. Sign against a FRESH blockhash —
   the payment is only good for ~60 seconds.
   The price is recomputed from your body and must match the quote, so do
   not change the body between the two calls.
4. Settlement is confirmed on-chain, then the model runs. The reply is
   `200 {"id", "model", "content", "tool_calls", "finishReason", "usage": {"inTokens", "outTokens"}, "charged": {"amountUsd", "amountMicro", "chainId", "txHash"}, "attestation", "docs"}`,
   `tool_calls` present when the answer called a tool (the OpenAI shape, `arguments` a JSON string).
   `usage` is the provider-reported actual token count — an accounting
   fact; the charge was fixed before the model ran.

## If you lose the answer

A retry with the SAME `X-PAYMENT` answers `409 payment_already_spent`, never a
fresh price: you cannot be charged twice for one authorization. The completion
itself is gone (no prompt and no answer is stored, ever), so recovery means the
receipt, not the text. Find it with the owner call below. A new question needs
a new authorization.

If we settled and the model did NOT serve, you get `serve_failed` with an
`invoiceId`: replay that invoice free, above. Nothing is owed twice.

## Wallet-signed calls (no payment)

Reads and replays are wallet-signed, no session. Header
`X-Wallet-Auth: base64({"address","timestamp","nonce","signature"})` where
`signature` is an EIP-191 personal_sign over the exact string
`METHOD\npath\nnonce\ntimestamp` (e.g. `GET\n/api/inference/invoices\n<nonce>\n<unix seconds>`).
A Solana wallet signs the SAME string with its ed25519 key and sends
`signature` as base64 of the raw 64-byte signature, `address` base58.
Nonces are single-use; timestamps must be within the server skew.

- `GET https://agentgates-backend.vercel.app/api/inference/invoices` — your settled requests: sizes,
  actual usage, amounts, transaction hashes.
- REPLAY: if a paid request answers `502` with `{"code": "serve_failed",
  "invoiceId": "..."}`, the ticket stands and the delivery is on us —
  re-send the SAME request body plus `"replayInvoice": "<invoiceId>"` under
  `X-Wallet-Auth` (signed by the paying wallet) to the same endpoint. No
  new payment. A served invoice never replays.

## Claude Code

`curl -fsSL https://agentgates.ai/claude-code | sh`

Claude Code speaks the Anthropic Messages API and nothing else, so the
installer puts a translator on YOUR machine, takes a bearer key you minted
here, and adds a row per model to its `/model` picker. Every turn is a
streamed call to this lane on your prepaid balance, with its own invoice; a
confidential row runs in attested hardware and says so on the receipt. A model
this catalog does not carry is relayed to Anthropic with your own
authorization header untouched, so one session holds both.

Read it before you run it: https://agentgates.ai/claude-code-source

## Verify the hardware (confidential tier)

Every CONFIDENTIAL model's attestation is public:
`GET https://agentgates-backend.vercel.app/api/inference/attestation?model=<model id>` re-serves the
serving provider's attestation exactly as published (a per-model TDX report
with workload keyset and receipt signing keys, or the enclave proxy's
SEV-SNP document, depending on the provider), plus this lane's own live
read of the model's confidential status. The report is re-served verbatim
and names its own origin — check it against what it claims. An OPEN-tier
model answers the same endpoint with `confidential: false` and no report,
because none exists.

## What we can see (stated plainly)

Your prompt and the completion transit this service in memory over TLS on
the way to the provider. They are not stored and not logged by us, on either
tier; the prompt's byte count and the provider's token counts are the only
metering facts kept, on your invoice. The chain sees the payment. On the
CONFIDENTIAL tier the model runs inside attested hardware — verify it
yourself, above. On the OPEN tier the model runs on ordinary provider
infrastructure: we make no claim about what that provider can see, which is
exactly why the tier is named everywhere it is sold.

## Rules

Sanctioned wallets are refused at every settlement (fail-closed screen).

Two daily walls, both ahead of the payment gate, so nothing is ever charged
on a refusal. `503 at_capacity` means the lane has SERVED its requests for
this UTC day — drawn and paid alike, because either one is a call we make
for you. `503 sales_at_capacity` means it is not TAKING more money today; a
balance topped up earlier still serves, because those dollars were counted the
day they were paid. A top-up is one of the day's sales, so
`GET /api/inference/credits` publishes `sale.availableUsd` — how much the
lane will add to balances right now — before you pick a rung.

Payments are final: a ticket settles a ceiling, the served tokens are the
charge and the rest goes to your balance, and a serve failure is
re-served free rather than refunded.
