AI Signal 252
Show HN: CostPerPrompt – Live AI API pricing and real-workload cost calculators
Most AI cost estimates miss prompt caching (up to 90% input cost reduction) and batch processing (~50% off), leading to projections 2–3× too high or low. Engineers planning chatbot, agent, RAG, or voice AI deployments can now model real usage patterns—growing context, retries, multi-step loops—against actual token rates instead of naive per-token math.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Tracks pricing for 232+ models with automatically refreshed input/output token rates and context windows.
Provides workload-specific calculators for chatbots, agents, RAG, voice AI, and image APIs that simulate real patterns like conversation history growth, tool-use retries, and separate indexing versus generation phases.
Accounts for prompt caching and batch discounts, which can cut input costs by up to 90% and overall costs by roughly 50%, respectively—discounts most other cost estimators ignore.
THE CLUSTER