Check GPU commitment cost against useful work
On this page
This check is due for a source refresh. Confirm the current documentation before you rely on provider-specific details.
You risk paying for unproductive capacity if you renew on GPU activity alone. Compare period cost, reservation use, and accepted workload output before deciding, using one reporting period and a shared scope.
What you need first
- Tool
- Use this manual worksheet with existing cost, reservation, monitoring, and workload records. In Azure, start in the Azure portal, Microsoft's web console, at Reservations, then select the utilization percentage to see history and details.
- Access
- Azure lists reservations where you have the Owner or Reader role. This does not establish access to cost or monitoring records. Ask authorized billing and platform colleagues to provide those read-only results for your selected scope.
- If you do not use that tool
- Give your billing owner and GPU platform engineer one provider, account, subscription, or project, commitment ID, resource scope, and start/end dates. Request period cost, available provider utilization with its definition, and existing GPU metrics. Ask the workload owner for accepted output under one fixed definition.
Why this is worth a look
You can mistake a busy GPU for a well-used commitment or useful output. Keep activity, provider-reported commitment utilization, and locally defined output separate.
AWS CloudWatch metric nvidia_smi_utilization_gpu measures the percentage of the sample period when one or more kernels ran, not accepted results. Use Effective Cost for the period comparison. FOCUS defines it as cost recognized in a charge period, including applicable pricing adjustments and recognized portions of related purchase charges. It can differ from Billed Cost, so keep the invoice comparison separate.
Run this check
CHECKLISTComplete these read-only steps using existing records, not resource commands. Review one commitment and one reporting period. Leave unavailable measures blank as N/A rather than estimating them from GPU activity.
READ-ONLY WORKSHEET
Execution surface: manual review of existing reports and monitoring views.
Access: use authorized colleagues' results if you cannot view a record.
Azure reservation view: Owner or Reader on the reservation; other records
need separate authorized access.
Scope: one provider, account/subscription/project, and commitment.
Time window: enter explicit start/end dates and use them for every input.
Units: cost in billing currency; utilization in its reported unit;
GPU activity in the metric's defined unit; output in one local unit.
1. SET THE COMPARISON
[ ] Provider and account/subscription/project: ____________________
[ ] Commitment program and ID: ____________________
[ ] Period start/end and time zone: ____________________
[ ] Finest resource scope shared by all records: ____________________
[ ] Keep node or job metrics separate if they cannot be matched to that scope.
2. RECORD COST AND RESERVATION USE
[ ] Effective Cost, billing currency, and source: ____________________
[ ] Confirm the cost uses the FOCUS definition; otherwise mark it N/A.
[ ] Billed Cost for separate invoice comparison: ____________________
[ ] Provider-reported commitment utilization, unit, definition: __________
[ ] Utilization period and scope: ____________________
[ ] Azure: Reservations > utilization percentage gives history and details.
[ ] For other programs, request the available measure from the billing owner.
[ ] Do not convert a spend measure or GPU activity into commitment GPU-hours.
3. ADD EXISTING GPU ACTIVITY
[ ] Metric name, definition, unit, sample period: ____________________
[ ] Resource scope, reporting window, aggregation, value: _______________
[ ] AWS: use nvidia_smi_utilization_gpu only where the CloudWatch agent
collects NVIDIA GPU metrics from a Linux server with an NVIDIA driver.
Meaning: percentage of the sample period when one or more kernels ran.
CloudWatch has no Unit by default for these metrics.
[ ] Azure Kubernetes Service: use existing NVIDIA DCGM Exporter metrics
collected by Azure Monitor managed service for Prometheus. Record the
available utilization metric and its unit.
[ ] Google Kubernetes Engine: use existing DCGM monitoring metrics; record
the chosen activity metric's definition and unit.
[ ] Mark missing telemetry N/A. Do not enable collection as part of this review.
4. ADD LOCAL OUTPUT AND RECORD THE DECISION
[ ] Workload owner's definition of accepted output: ____________________
[ ] Accepted output count, local unit, period, and scope: _______________
[ ] Only when cost and output have matching scope and period:
Effective Cost / accepted output count = __________ [currency/unit]
[ ] If output is zero or missing, record unit cost N/A, not zero.
[ ] Record separately: commitment use, GPU activity, cost per accepted unit.
[ ] Decision: investigate / defer renewal decision / retain current plan.
[ ] Evidence gap or question for the workload owner: ____________________How to confirm it
- 01
Choose one comparable scope
Choose one commitment and explicit start/end dates with your billing and workload owners. Use the finest resource scope they can both support. Do not assign reservation-level cost to individual jobs without a documented allocation.
- 02
Separate cost from commitment use
Record Effective Cost and the program's own utilization measure in the worksheet. Keep Billed Cost for invoice comparison. For Azure, use utilization history rather than treating the last known percentage as the value for your whole reporting period.
- 03
Ask what the GPU activity delivered
Ask the platform engineer for existing activity metrics and the workload owner for accepted output. Agree on one output definition before calculating cost per unit. Treat that definition as a local business measure, not a provider metric.
- 04
Decide whether the renewal needs review
Flag low commitment use or high cost per accepted unit against your team's requirements for review before renewal. Ask the workload owner to explain the mismatch before recommending less capacity. Defer the comparison if the records cannot be aligned; do not let a high activity percentage settle the decision.
Before making changes
Treat cost per accepted unit as a local calculation, not a provider guarantee of efficiency. This worksheet assumes existing monitoring, a fixed output definition, and matching cost/output scope and dates. Metrics from different services need not measure the same activity. Missing data is not zero use, and this review does not establish whether a commitment can be changed.