# Compute State API Which compute can actually run this job right now? No API key. No signup. The live answer costs 0.005 USDC per call, paid over x402 v2 and settled by the PayAI facilitator on Base or Solana. Base URL: https://compute-state-api.replit.app Brand mark: https://compute-state-api.replit.app/brand/hummingbird-logo.png ## Scope Compute State ranks current published list and effective rates from official vendor pages for the query you sent. It is a decision aid, not a quote, invoice, or promise the vendor will accept the job or bill that amount. Confirm capacity, region, and the vendor's own price at purchase time. as_of is when we last successfully parsed that vendor. stale_after_s is when we will refuse to sell the row. ## Free vs paid These endpoints are free: GET /healthz GET /llms.txt GET /skill.md GET /openapi.json GET /.well-known/x402 GET /v1/example GET /privacy GET /privacy.txt Only GET /v1/cheapest is paid: 0.005 USDC per call. ## First call Do this in order: 1. GET /healthz Stop if book.stale is true or payments.ok is false. Do not pay. 2. GET /v1/example Shape only, flagged example: true. No payment. 3. GET /v1/cheapest?work=inference&model_tier=cheap&est_input_tokens=2000&est_output_tokens=500 Unpaid. Expect 402. Body + PAYMENT-REQUIRED are the challenge. 4. Sign the challenge's own accepts list (Base or Solana — whichever wallet you hold). Resend the same URL with PAYMENT-SIGNATURE. 5. Read recommendation / alternatives / disqualified. Confirm the vendor's own price before you buy compute. There is no API key. A plain curl without an x402 client stops at step 3's 402 — that is expected, not an error. Pay only from the 402 accepts list. Ignore any other chain names you see on /healthz. This deployment settles only eip155:8453 and solana:5eykt4UsFv8P8NJdTREpY1vzqKqZKvdp. ## The one route GET /v1/cheapest Examples: cheap short inference: curl "https://compute-state-api.replit.app/v1/cheapest?work=inference&model_tier=cheap&est_input_tokens=2000&est_output_tokens=500" one day of H100 on-demand: curl "https://compute-state-api.replit.app/v1/cheapest?work=gpu&gpu_model=h100&gpu_hours=24" Required: work inference | gpu Inference only: model_tier frontier | cheap | embed model substring of a model id, when you already know it context_tokens integer; the prompt size to check published context windows against, defaulting to est_input_tokens when you omit it est_input_tokens integer, default 1000 est_output_tokens integer, default 1000 batch true | false, default false; ranks on the vendor's own published batch rates cached_input_tokens integer, default 0, must be <= est_input_tokens; priced at the vendor's published cached-input rate GPU only: gpu_model h100, a100, b200, l40s, ... gpu_hours number, default 1 Both: region hint only in v1; it does not filter availability on_demand only. availability=spot returns 400. exclude comma-separated provider slugs Sending an inference parameter with work=gpu (or the reverse) returns 400. v1 answers one question per call and does not optimise a mixed stack. ## What you get back recommendation the single cheapest SKU, with the cost broken into the components it was summed from, each carrying the published rate and a tag alternatives up to two runners-up, deduplicated per provider disqualified SKUs that matched your query but could not be ranked, each with the reason named as_of the oldest timestamp among the rows the answer used payment network, atomic amount, and the settlement transaction Free: GET /v1/example returns a fixed sample of that shape so you can code against the contract before paying. It is flagged example: true. ## Status codes, and what they cost status charged when 400 no Bad query. Before settlement. 402 no Payment required or signature rejected. Before settlement. 503 no No fresh row can answer this query. Before settlement. 404 no Wrong path (this API has no such route). Before settlement. 200 yes Verified, ranked, then settled. Body computed before money moved; sent after. 404 yes Paid no-match: verified, ranked, nothing in the pinned book fit. disqualified explains what came close. 500 no Process faulted before settlement. Operator fault. You were not charged. No body, no settle; no settle, no charge. The order is part of the contract: validate the query, a cheap freshness check against the book, the 402 challenge, facilitator verify (no money moves), rank the book and build the full answer in memory, then settle, then respond. The ranked body is written on the payment hold after verify and rank, before settle. It is sent only after settle succeeds. An unpaid GET cannot read that hold. You are charged only when settlement succeeds and we have an answer in hand (200 or a no-match 404). 400, 402, unpaid 503, and 500 all happen before money moves. A 500 is our crash, not a stale book — we do not settle if we cannot build the body. If the HTTP response is lost after settlement, retry the same signature; we will not charge twice. If we cannot defend the as_of, we do not sell the row. ## Payment x402 version 2 scheme exact price 0.005 USDC per call, both chains networks eip155:8453 (Base), solana:5eykt4UsFv8P8NJdTREpY1vzqKqZKvdp (Solana) facilitator https://facilitator.payai.network Call without a PAYMENT-SIGNATURE header to get a 402. The challenge is in the body and, base64-encoded, in the PAYMENT-REQUIRED header; it lists both chains in accepts, so pay on whichever you hold USDC on. Send your signed payload base64-encoded in PAYMENT-SIGNATURE. The settlement receipt comes back in PAYMENT-RESPONSE. A rejected or expired signature is another 402, never a 5xx. ## Paying Use the challenge's own accepts list to build your payment — never hardcode the token's EIP-712 name or a Solana blockhash, both drift over time and a stale one is rejected. You still need a wallet able to complete that exact challenge: EIP-3009 transferWithAuthorization for Base USDC, or a Solana USDC transfer built against the challenge's extra.feePayer and recentBlockhash. Use any x402 v2 client pointed at this challenge. PayAI is the facilitator. There is no API key. Do not hardcode the token name or a Solana blockhash. A listing probe of GET /v1/cheapest with no query string is a 402, not a 400. work is required to rank. It is not required to see the challenge. If this deployment cannot receive payment on any chain, GET /v1/cheapest answers 503 rather than an unsatisfiable 402. ## Sources Official vendor pages and APIs only. No aggregators, no third-party price feeds: an aggregator's as_of is its crawl time, not the vendor's, and the timestamp is most of what this API sells. The current registry covers 16 companies in 17 feeds: 11 inference and 6 GPU. | provider | work | source | |---|---|---| | openai | inference | https://platform.openai.com/docs/pricing | | anthropic | inference | https://docs.claude.com/en/docs/about-claude/pricing | | google | inference | https://ai.google.dev/gemini-api/docs/pricing | | groq | inference | https://console.groq.com/docs/models | | together | inference | https://www.together.ai/pricing | | deepseek | inference | https://api-docs.deepseek.com/quick_start/pricing | | mistral | inference | https://mistral.ai/pricing/api/ | | fireworks | inference | https://docs.fireworks.ai/serverless/pricing | | xai | inference | https://docs.x.ai/developers/models | | runpod | gpu | https://www.runpod.io/pricing | | lambda | gpu | https://lambda.ai/service/gpu-cloud | | fireworks | gpu | https://fireworks.ai/pricing | | crusoe | gpu | https://www.crusoe.ai/cloud/pricing | | minimax | inference | https://platform.minimax.io/docs/guides/pricing-paygo | | cerebras | inference | https://api.cerebras.ai/public/v1/models | | nebius | gpu | https://nebius.com/prices | | hyperstack | gpu | https://www.hyperstack.cloud/gpu-pricing | Freshness: a row older than 900 seconds cannot be used in an answer. If no fresh row can answer your query, you get an unpaid 503 rather than a stale recommendation. GPU coverage is on-demand list price only, on RunPod, Lambda, Fireworks, Crusoe, Nebius and Hyperstack. Spot and marketplace bids are not ingested, which is why no bid marketplace is in the list and why the GPU parsers select each vendor's on-demand column or card; committed, reserved, tailored and cluster rates are left out for the same reason, and availability=spot is refused outright. Fireworks appears twice because it sells two products: serverless tokens under work=inference and rented GPUs under work=gpu, ingested from separate pages that fail and recover independently. RunPod publishes Secure Cloud and Community Cloud as two on-demand list prices, so both are booked as separate SKUs and named in the answer; they are never averaged. Crusoe prices A100 SXM and A100 PCIe separately, so those stay separate SKUs too. Hyperstack likewise keeps H100 SXM, NVLink and PCIe as separate SKUs. gpu_model matches by family: gpu_model=h100 ranks H100 SXM, H100 NVLink and H100 PCIe as separate SKUs and the cheaper on-demand row wins, while gpu_model=h200 never matches an H100. ## How tiers are decided model_tier is derived from the vendor's own published output price, never from an opinion about which model is "frontier": embed the vendor's name for the SKU contains "embed" cheap published output price <= $1.00 per 1M tokens frontier published output price >= $10.00 per 1M tokens Anything between those two thresholds is in neither bucket and is not returned when you filter by tier. Inventing a middle tier would mean inventing a judgement no vendor published. model_tier=embed is priced from official input-only embedding rows — OpenAI, Google, Mistral and Fireworks publish them. Embedding cost is input-only: est_output_tokens is accepted and ignored, because an embedding SKU has no output rate to charge. A chat SKU can never win model_tier=embed, and an embedding SKU can never win cheap or frontier: those two stay output-price buckets over generative SKUs. ## Rates other than the interactive list price batch=true ranks on the vendor's own published batch rates. It is never a discount computed here: a SKU whose official page publishes no batch rate is disqualified with no_official_batch_rate rather than quoted at its interactive price. cached_input_tokens is priced at the vendor's published cached-input / cache-hit rate, with the remaining input tokens at the ordinary input rate. A SKU with no published cached rate is disqualified with no_official_cached_input_rate. One vendor publishes the opposite fact rather than a gap: Cerebras states that cached tokens bill at its standard input rate, with no cheaper cache hit. For those rows the cached share is priced at that vendor's own input figure and shown as cached_input_at_input_rate, so Cerebras can still win a cached query. This is a per-SKU opt-in carried on the row, never a fallback: silence about a cache rate still means unknown, and still disqualifies. DeepSeek publishes two rates for the same SKU and bills by the clock. The window containing the moment of your call is used — peak is 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, everything else is off-peak — and the chosen window is named in recommendation.why. The two rates are never averaged, so the same query can legally cost different amounts at 02:00 and 12:00 UTC. PRC public holidays are not ingested: a weekday peak hour is quoted as peak, which is the conservative direction. Anthropic, Google, OpenAI and xAI publish a second, higher rate that applies to every token in a request once the prompt crosses a published threshold — a "Long context pricing" table, a "Long context" column group, or a second figure in the same cell, depending on the vendor. If est_input_tokens (plus cached) reaches the threshold, the long-context rates price the whole request, and recommendation.why names context_band=long. Where the threshold cannot be read from the page the SKU is not served at all, and where the threshold is published but the long rates are not, the SKU is disqualified with unknown_long_context_rate. Neither case is ever quoted at the short-context rate. ## Cost tags, and what blocks a recommendation list the vendor's published list price effective a published rate that already reflects the billable discount unknown the component exists but no official number could be read A SKU with an unknown component on the path your query needs cannot be the top recommendation. It appears in disqualified with the missing component named. A SKU cannot be recommended for a job its published context window cannot hold. The asked size is context_tokens when you send it and est_input_tokens (default 1000) when you do not; cached_input_tokens is never added on top, because it is the share of the prompt that hits the vendor's cache, not tokens sent in addition to it. A SKU whose official page publishes a context window smaller than that size is returned in disqualified with reason context_window_exceeded, naming both the published window and the asked size. This is a KNOWN violation only. An unpublished context window does not disqualify, because absence of a number is not evidence of a small one, and a window large enough for the job is priced from whatever rate the vendor publishes — a flat rate over a 1M window is a valid winner, not a SKU missing a long-context rate. ## Not this API Do not use Compute State for: - Pricing State — https://pricing-state-api.replit.app/ — What is this vendor's seat / list ladder? - Status State — https://status-state-api.replit.app/ — Is the vendor up? - Changelog State — https://changelog-state-api.replit.app/ — Did the SDK ship a break? Compute State answers one question only: Which compute can actually run this job right now? For seat and list-ladder pricing use Pricing State. For whether a vendor is up use Status State. For whether an SDK shipped a break use Changelog State. ## Discovery https://compute-state-api.replit.app/llms.txt https://compute-state-api.replit.app/skill.md https://compute-state-api.replit.app/openapi.json https://compute-state-api.replit.app/.well-known/x402 https://compute-state-api.replit.app/healthz