Inputs
Adjust the hardware, duty cycle, power draw, and comparison assumptions used in the estimate.
Estimate hardware, power, throughput, and cloud cost on the same basis.
Hardware amortization, electricity, estimated token output, and a direct cost-per-million comparison against the selected cloud default. It produces rough financial and throughput estimates from the values you enter and a small set of simplifying assumptions.
Model quality, latency variance, cooling, networking, setup time, support burden, compliance requirements, credits and negotiated vendor pricing are all outside the model. So is every one of the following:
Cloud per-token prices and "typical" cloud throughput values shown here are illustrative defaults captured at the time of authoring. Published prices, model availability, rate limits and streaming speeds change frequently, and the direction is not always down. Treat the cloud defaults as placeholders. Verify current pricing and benchmark real performance before making any purchasing, contracting, capacity-planning or budgeting decision.
A comparison between a number you measured and a number this page remembered is not a comparison.
Independently verify all figures with vendor invoices, measured power draw at the wall rather than a datasheet TDP, and current published pricing. Benchmark the model you would actually run, at the quantisation, context length and concurrency you would actually use, on the hardware you would actually buy.
Adjust the hardware, duty cycle, power draw, and comparison assumptions used in the estimate.