---
name: compute-state
description: Find which compute can actually run a job right now — LLM inference or GPU hours — from official vendor prices, paid per call over x402.
---

# Compute State

Ask one question: Which compute can actually run this job right now?

Base URL: https://compute-state-api.replit.app
Brand mark: https://compute-state-api.replit.app/brand/hummingbird-logo.png
No API key. No signup.

## What this is, and is not

Compute State ranks current published list and effective rates from official vendor pages for the query you sent. It is a decision aid, not a quote, invoice, or promise the vendor will accept the job or bill that amount. Confirm capacity, region, and the vendor's own price at purchase time.

## First call

1. GET /healthz — if book.stale or payments.ok is false, stop. Do not pay.
2. GET /v1/example — shape only, example: true.
3. GET /v1/cheapest — unpaid 402. Pay from that accepts list only (Base or Solana).
4. Resend with PAYMENT-SIGNATURE. 200 or paid 404 is the answer.

One paid route. 0.005 USDC. No key.

## Provider fleet

The current registry covers 16 companies in 17 feeds: 11 inference and 6 GPU.

| provider | work | source |
|---|---|---|
| openai | inference | https://platform.openai.com/docs/pricing |
| anthropic | inference | https://docs.claude.com/en/docs/about-claude/pricing |
| google | inference | https://ai.google.dev/gemini-api/docs/pricing |
| groq | inference | https://console.groq.com/docs/models |
| together | inference | https://www.together.ai/pricing |
| deepseek | inference | https://api-docs.deepseek.com/quick_start/pricing |
| mistral | inference | https://mistral.ai/pricing/api/ |
| fireworks | inference | https://docs.fireworks.ai/serverless/pricing |
| xai | inference | https://docs.x.ai/developers/models |
| runpod | gpu | https://www.runpod.io/pricing |
| lambda | gpu | https://lambda.ai/service/gpu-cloud |
| fireworks | gpu | https://fireworks.ai/pricing |
| crusoe | gpu | https://www.crusoe.ai/cloud/pricing |
| minimax | inference | https://platform.minimax.io/docs/guides/pricing-paygo |
| cerebras | inference | https://api.cerebras.ai/public/v1/models |
| nebius | gpu | https://nebius.com/prices |
| hyperstack | gpu | https://www.hyperstack.cloud/gpu-pricing |

## Free vs paid

These endpoints are free:

    GET /v1/example
    GET /openapi.json
    GET /.well-known/x402
    GET /llms.txt
    GET /skill.md
    GET /healthz
    GET /privacy
    GET /privacy.txt

Only GET /v1/cheapest is paid: 0.005 USDC per call. It uses x402
v2 on Base or Solana with the PayAI facilitator.

## When to use this

Use it when you need to choose where to run a workload and the choice turns on
price: picking a model for a known token budget, or picking a GPU cloud for a
known number of GPU-hours.

Do not use it for:

- Pricing State — https://pricing-state-api.replit.app/ — What is this vendor's seat / list ladder?
- Status State — https://status-state-api.replit.app/ — Is the vendor up?
- Changelog State — https://changelog-state-api.replit.app/ — Did the SDK ship a break?

## Learn the shape for free

    curl https://compute-state-api.replit.app/v1/example

That returns a fixed sample of the paid body, flagged example: true.

## Ask a real question

Embeddings, priced on input alone:

    curl "https://compute-state-api.replit.app/v1/cheapest?work=inference&model_tier=embed&est_input_tokens=100000"

An overnight batch job where most of the prompt is a cached prefix:

    curl "https://compute-state-api.replit.app/v1/cheapest?work=inference&batch=true&est_input_tokens=200000&cached_input_tokens=180000&est_output_tokens=5000"

Both return 402 the first time. The challenge body lists both chains under
accepts; the same JSON is base64-encoded in the PAYMENT-REQUIRED header. Sign
the payment payload for whichever chain you hold USDC on and resend the request
with it base64-encoded in PAYMENT-SIGNATURE. The 200 carries the settlement
receipt in PAYMENT-RESPONSE.

## Paying

Price is 0.005 USDC per settled call, on Base (eip155:8453) or
Solana, x402 v2 scheme exact, settled by the PayAI facilitator. Build your
payment from the challenge's own accepts list — do not hardcode the token's
name or a blockhash, both change over time. You need a wallet that can
complete that challenge: EIP-3009 transferWithAuthorization on Base USDC, or a
Solana USDC transfer built against the challenge's extra.feePayer and
recentBlockhash.

Use any x402 v2 client pointed at this challenge. PayAI is the facilitator.
There is no API key for this product. If no payout wallet is configured on
this deployment, GET /v1/cheapest answers 503, not a 402.

A listing probe of GET /v1/cheapest with no query string is a 402, not a 400.
work is required to rank. It is not required to see the challenge.

## Reading the answer

- recommendation.est_cost_usd is the total for the job you described, not a rate.
- recommendation.components shows how that total was assembled: each entry has
  the published rate (rate_usd), its unit, its contribution (usd), and a tag.
- A tag of unknown never appears on the top recommendation. If a SKU had an
  unreadable component it is in disqualified with the component named.
- as_of is the oldest input timestamp. Anything older than 900s
  cannot be used in an answer.
- disqualified is not noise. It is the list of things that matched your question
  and could not be ranked, with the reason on each.

## Failure modes worth handling

- 400 — your query was invalid. The message says what. Unpaid. Fix and retry.
  Sent only once a signed payment verifies; an unsigned call gets the 402.
- 402 — pay, or your signature was rejected at verify or settle. Unpaid. Retry
  with a fresh payload — unless the body says settlement is pending, in which
  case resend the identical PAYMENT-SIGNATURE.
- 503 — the book is stale or empty. Unpaid. Honour Retry-After.
- 404 — wrong path (unpaid), or a verified call whose book had no match once
  ranked, then settled (paid); disqualified explains what came close.
- 500 — rare: the process faulted before settlement. Operator fault. You were
  not charged. Not something a retry with different parameters fixes.

No body, no settle; no settle, no charge. You are charged only when settlement succeeds and we have
an answer in hand (200 or a no-match 404). 400, 402, unpaid 503, and 500 all
happen before money moves. A 500 is our crash, not a stale book — we do not
settle if we cannot build the body. If the HTTP response is lost after
settlement, retry the same signature; we will not charge twice.

The book we rank is the one that passed the pre-pay freshness check, even if
the clock crossed 900s during the wallet step.

## Constraints to respect

- work is required and must be exactly one of inference or gpu. Sending
  inference parameters with work=gpu (or the reverse) is a 400 — this API
  answers one question per call and will not optimise a mixed stack.
- availability=spot is a 400. Spot is not ingested in v1.
- region is accepted as a hint and does not filter in v1.
- batch and cached_input_tokens are inference-only; sending either with
  work=gpu is a 400, as is cached_input_tokens greater than est_input_tokens.
- batch=true and cached_input_tokens > 0 rank on rates the vendor itself
  publishes. A SKU without the rate your query needs is disqualified
  (no_official_batch_rate, no_official_cached_input_rate), never silently
  priced at its interactive rate. The exception is a vendor that publishes
  that it charges the standard input rate for cached tokens (Cerebras), where
  the cached share is priced at that published input rate and named
  cached_input_at_input_rate.
- model_tier=embed is input-only: est_output_tokens is ignored rather than
  rejected, and only official embedding SKUs compete.
- DeepSeek is billed by the clock — peak (01:00-04:00 and 06:00-10:00 UTC, Monday to Friday) and off-peak are
  separate published rates, so the same query can cost different amounts at
  different hours. recommendation.why names the window used.
- gpu_model matches a GPU family, not a product string: h100 covers H100 SXM
  and H100 PCIe as separate SKUs; h200 never matches an H100.
- model_tier buckets come from published output price: cheap is
  <= $1.00/1M, frontier is >= $10.00/1M. Prices between the two
  are in neither bucket and will not be returned under a tier filter.
