Budget AI search and tool usage separately
On this page
This check is due for a source refresh. Confirm the current documentation before you rely on provider-specific details.
A token-only budget can miss search and tool charges, so budget each operation with its documented billing unit. For Gemini 3 Google Search grounding, count generated queries sent to Search, not prompts, or one prompt that creates several queries can understate cost.
What you need first
- Tool
- Use the checklist as a local worksheet in a text editor or spreadsheet. Start with one workload, its model and tool inventory, existing usage records, and one defined period. It makes no cloud changes.
- Access
- The worksheet needs no cloud permissions. Ask an authorized cloud colleague for read-only usage results covering the chosen account, project or subscription, regions, tools, and period.
- If you do not use that tool
- For AWS browser or code-interpreter usage, ask an authorized AWS colleague to open the Amazon CloudWatch Bedrock AgentCore Observability Console and Built-in Tools Session page. Give the tool, session scope, and time range; request existing CPU and memory results and the AWS billing statement for the same scope.
Why this is worth a look
A prompt count can understate Gemini 3 grounding usage because one prompt may generate several charged Search queries. Model input and output tokens are billed separately. Vertex AI documents a successful-source condition for some grounding pricing, so apply the rule to the exact model and feature. Azure Web Search uses Grounding with Bing and adds costs. AgentCore browser and code-interpreter telemetry reports vCPU-hours and memory GB-hours for monitoring, not authoritative billing.
Run this check
CHECKLISTComplete this local worksheet from existing records. Keep provider units separate, distinguish observations from assumptions, and leave unverified charges unresolved instead of treating them as zero.
AI WORKFLOW BUDGET REVIEW
[ ] Workload: ____________________
[ ] Account, project or subscription and region(s): ____________________
[ ] Period start, end and time zone: ____________________
[ ] Observed usage or planning assumption: ____________________
[ ] Exact provider, model, search feature and tools: ____________________
MODEL USAGE
[ ] Record input and output tokens separately for every model call.
[ ] For Vertex AI, match model, region, context tier and processing mode to the applicable rate entry. Record rates per 1M tokens where listed.
[ ] If using LiveAPI, include accumulated session-context tokens billed again on each turn, not only new tokens.
SEARCH AND GROUNDING
[ ] For Gemini 3 Google Search grounding, record generated queries sent to Search. Do not substitute prompt count.
[ ] Record the charging condition for the exact model and grounding feature. Do not treat a successful-source rule as a universal per-prompt meter.
[ ] For Azure Web Search, request the applicable Grounding with Bing billing unit and terms before estimating additional cost.
[ ] For AgentCore Web Search Tool, request its applicable billing unit and rate. Do not use browser or code-interpreter units for this connector.
AGENTCORE BROWSER OR CODE INTERPRETER, IF USED
[ ] Request existing session usage for the same scope and period: elapsed seconds, vCPU-hours and memory GB-hours.
[ ] Label these values as monitoring data, not billed quantities.
[ ] Verify billing units and rates before using telemetry to estimate charges.
ONE BUDGET LINE PER VERIFIED METER
Provider / model or tool: ____________________
Quantity and unit: ____________________
Charging condition: ____________________
Rate and units covered: ____________________
Rate source or contract and date recorded: ____________________
Estimated line cost = quantity / units covered by rate x rate
Unresolved unit, rate or assumption: ____________________
[ ] Keep model and search/tool lines separate.
[ ] Compare AWS estimates with the AWS billing statement for the same scope and period; investigate differences before accepting the budget.How to confirm it
- 01
Set one scope and period
Record one workload, its account, project or subscription, regions, and time window. List the exact models and tools, including model calls that consume search results. Mark missing quantities as assumptions, not measured usage.
- 02
Count the documented units
Separate model input tokens, output tokens, and search usage. For Gemini 3 grounding, count generated queries sent to Search. If the workload uses AgentCore browser or code interpreter, request existing session usage for the same scope and period.
- 03
Match each line to a rate
Record the billing unit, charging condition, and rate source for each model and tool. Convert token quantities to millions only when applying a per-million-token rate. Leave Azure Grounding with Bing and AWS tool estimates unresolved until their applicable units and rates are confirmed.
- 04
Resolve gaps before approval
Flag missing rates and usage assumptions with the workload owner. For AgentCore, compare telemetry estimates with the AWS billing statement because monitoring totals can differ from metered usage. Do not treat an unresolved tool charge as zero.
Before making changes
Treat this as a usage estimate, not a bill. Quantities and rates must cover the same scope and period, and this review does not model provisioned-capacity commitments. AgentCore telemetry may differ from billing because of aggregation timing, reconciliation, and measurement precision. Before approving Azure Web Search, review Grounding with Bing terms because customer data flows outside the Azure compliance boundary.