Choose models by cost per usable answer, not token price
On this page
This check is due for a source refresh. Confirm the current documentation before you rely on provider-specific details.
Choose the lowest-cost model that passes the same quality and latency tests. Otherwise, token price can hide rejected answers and leave you paying more for each usable result.
What you need first
- Tool
- Use the worksheet in a text editor or spreadsheet to review existing evaluation results. Start with a shared task set, scoring criteria, application logs and provider usage exports for one recorded time window. No cloud commands are needed.
- Access
- Ask an authorized cloud colleague to provide read-only evaluation and usage results for the named models, deployments, regions and time window. Ask your finance partner for applicable rate assumptions.
- If you do not use that tool
- Give the application owner the candidate list, task-set version and start/end times with timezone. Request task-level pass/fail results, application response times and usage totals split by billing unit. Have the cloud owner identify missing telemetry.
Why this is worth a look
Skipping this can make a cheap token price look best when too few answers pass your requirements. Amazon Bedrock and Google pricing tables list separate input and output token prices, while cache classes and other billing arrangements can affect the estimate. Match each candidate's rates to measured usage, then divide included cost by accepted results.
Run this check
CHECKLISTComplete once per candidate using existing evaluation reports and authorized usage exports. This read-only review does not run models or change resources. Compare candidates only over the recorded scope and window.
READ-ONLY MODEL COMPARISON
Review existing records in a text editor or spreadsheet.
Do not run evaluations or change cloud resources with this worksheet.
Scope
Evaluation ID: ____________________
Start/end time and timezone: ____________________
Task-set version: ____________________
Provider, model/version, region and deployment: ____________________
Prompt version and output limits: ____________________
Retry policy: ____________________
Included cost components: ____________________
Excluded components and reason: ____________________
Comparability
[ ] Same tasks and scoring criteria for every candidate
[ ] Quality, safety and latency limits recorded before comparison
[ ] Prompt, output-limit and retry differences recorded
[ ] Scoring rubric and evaluator recorded where used
[ ] Missing records marked unavailable, not zero
Task-level results from existing evaluation records
Task ID | pass/fail | reason | application response time (milliseconds)
_______________________________________________________________
Accepted result means a task answer that passes the agreed criteria,
not merely a successful API request.
Usage and rate record for the same window
[ ] Input and output tokens recorded separately, where applicable
[ ] Cache usage classes kept separate where exposed
[ ] Other applicable billing units recorded with their unit names
[ ] All measured billable usage included, not just passing answers
[ ] Each usage class matched to its applicable rate and unit
[ ] Rate basis records model, modality, region, tier and contract basis
[ ] Usage is not counted twice across overlapping token metrics
Measured quantity | billing unit | applicable rate | estimated cost
________________________________________________________________
For per-million-token rates: cost = token count / 1000000 * rate.
Do not apply this formula to other billing units.
Summary
Tasks assessed: __________
Accepted results: __________
Pass rate (accepted / assessed, percent): __________
Latency limit met: yes / no / unknown
Estimated cost of included components (state currency): __________
Estimated cost per accepted result = included cost / accepted results
If accepted results = 0, record 'no accepted results'; do not divide.
[ ] Candidate passes every required quality and latency limit
[ ] Scope, missing usage and rate assumptions allow comparison
Decision: choose / reject / wait for missing data
Reason: ____________________How to confirm it
- 01
Set the acceptance rule
Define what counts as a usable answer and set quality, safety and response-time limits before comparing costs. Request results for the same tasks and scoring criteria. Record configuration differences rather than treating unlike evaluations as equivalent.
- 02
Request matching records
Ask the application owner for task results and client-side response times covering the evaluation window. Request matching provider usage from the cloud owner. Keep token counts separate from milliseconds. Azure's generic Latency metric must not be used for Azure OpenAI; its provider response-time measures also exclude client-side latency.
- 03
Match usage to billing units
Apply the candidate's applicable rates to each measured billing unit, including cache classes where available. Record included and excluded cost components. Do not use Bedrock Invocations as a count of all attempts: it counts successful requests. Use application records to account for retries, and mark missing usage rather than silently estimating it.
- 04
Choose only among passing candidates
Divide each candidate's estimated included cost by its accepted results, then choose the lowest-cost candidate that meets every required limit. Reject a candidate with no accepted results. Wait if missing usage, unmatched rates or different cost scopes could change the ranking.
Before making changes
Treat the ranking as an estimate for the recorded scope, not a full application bill. Assume the tasks represent the intended workload and the records cover the same window. For Generative AI on Google's Agent Platform, requests returning non-200 responses are not charged for input or output. Do not apply that rule to other Google services. A rejected answer can still come from a billable successful Bedrock request. Repeat the comparison when the model, prompts or workload change.