‹ Blog

How to Budget for AI Costs: A Framework for Forecasting, Attribution, and Control

Kris Newlin

Forecast volatile AI spending, segment costs by category, attribute usage to teams, and apply FinOps controls to manage token-based pricing at scale.

A token-metered API line can jump an order of magnitude in a quarter because prompts got longer, model mix shifted, an agent looped, or adoption spiked. Consumption-based AI pricing breaks the budgeting model most finance teams run on. The fix is to split the AI budget into five named categories: model API tokens, seat licenses and SaaS fees, compute and hosting, data preparation, and integration and people costs. To convert expected usage into a monthly dollar estimate, multiply tokens per request by requests per user per month by user count, price input and output tokens separately at the provider’s published per-million-token rate, then add the non-token lines on top. A flat-rate SaaS seat costs the same in month twelve as in month one; finance cannot manage a token line the same way. Once you have the bottom-up estimate, add a contingency buffer of roughly 20%, which matches the FinOps Foundation’s maximum budget variance target at Crawl maturity, because early-stage AI consumption is too volatile to forecast tightly. Then install the controls that keep actuals inside the plan: set hard spend caps at both the provider and agent level, tag every request with a cost center or virtual key at the gateway, configure anomaly alerts that fire below the cap so you can act before the limit is hit, and run a rolling forecast reviewed weekly on each consumption line. Treat AI spend like cloud infrastructure rather than software licensing. Forecast it on short cycles and attribute it to owners. Let the weekly review be the decision point.

Viktor is the AI employee that lives in Slack and Microsoft Teams, connects to 3,200+ tools, and does the work. Viktor uses usage-based pricing, so finance includes his costs in the same weekly review as every other consumption line. Segment his budget and set spend caps. Forecast it on short cycles.

Why AI costs behave differently from traditional software spend

Across the market, vendors often stack four distinct billing structures inside a single contract:

  • Per-token metering: OpenAI, Anthropic, Google, and AWS Bedrock bill per million tokens processed. The FinOps Foundation reports that output tokens are “consistently priced at a premium, often costing three to five times more than input tokens.”
  • Per-request fees: OpenAI charges $10.00 per 1,000 web search calls on top of content tokens, and transcription bills by the minute.
  • Prepaid credits: OpenAI’s prepaid API credits expire after one year and are non-refundable; many agent platforms meter in proprietary credits with their own expiry rules.
  • Outcome-based billing: Intercom Fin charges $0.99 per outcome and Zendesk starts at $1.50 per automated resolution. Bain research found only about 10% of hybrid AI meters are outcome-based so far.

A seat license caps your downside at seats times price. A consumption line has no ceiling unless you build one, and the underlying unit is itself unstable. FinOps analysis notes that “the same prompt can lead to various outputs, lengths, and costs.”

Choose the seat when usage per person is predictable and bounded, because the seat price caps your downside by definition. Choose metered pricing with a hard spending cap when work volume swings sharply across users and tasks — a flat seat cannot fairly price a quick chat question and a multi-hour autonomous coding session, which is exactly why GitHub moved Copilot to usage-based billing in June 2026. Anthropic reports Claude Code averages $13 per developer per active day and $150–$250 per developer per month across enterprise deployments, numbers that look reasonable until a single power user runs an overnight agent loop. Before you sign anything, divide your measured monthly token cost by active users and compare that figure with the quoted seat price — that ratio is your decision.

The volumes are moving fast too. Vista Equity Partners’ enterprise spending report puts average enterprise spending on AI model usage at roughly $7 million in 2025, nearly triple the $2.5 million spent in 2024. The Pragmatic Engineer’s token-spend report documented multiple organizations describing token spend increasing roughly 10× in six months. And IDC research found 96% of organizations deploying GenAI reported implementation costs higher than expected, with 71% admitting they have little to no control over where those costs come from.

The core categories every AI budget must include

A fixed license line can live in an annual budget; a token line cannot. Split the AI budget into four categories so each gets the right forecasting treatment:

Category

Examples

Volatility

Fixed SaaS and license fees

Per-seat AI subscriptions, AI add-ons that existing SaaS vendors bundle

Low. Vendors price these per seat; costs move at renewal.

Variable token and API consumption

OpenAI, Anthropic, and Gemini API usage; per-outcome fees; agent platform credits

High. Scales with usage, prompt length, model choice, and retries.

Compute and infrastructure

GPU hours, provisioned throughput commitments (Azure PTUs, Google GSUs), vector stores, self-hosted inference

Medium to high. Commitments smooth spend but idle capacity is pure waste.

Talent and data

ML and platform engineers, data preparation, fine-tuning, change management, compliance work

Medium. Often the largest share; moves with headcount, not usage.

Review the variable consumption line weekly. The fixed lines can keep a quarterly cadence, and the talent line follows your normal headcount planning.

Here is how to convert expected usage into a monthly dollar estimate before you lock in a line item:

  1. Requests per month: 500 users × 20 requests per working day × 21 working days = 210,000 requests
  2. Token volume: 3,000 input tokens and 600 output tokens per request → 630 million input tokens and 126 million output tokens
  3. Cost at GPT-5.6 Terra rates ($2.00/M input, $12.00/M output): $1,260 input + $1,512 output = $2,772/month
  4. Add 20% contingency (the FinOps Foundation’s maximum budget variance target at Crawl maturity): $2,772 × 1.20 = ~$3,326 budgeted line

Price input and output separately for every model in the mix, run the same arithmetic per use case, and set the hard cap at the buffered number.

Calculating total cost of ownership for AI

The API invoice is a minority of the real bill. A McKinsey analysis of 63 gen AI use cases found models account for only about 15 percent of the overall cost of gen AI applications, and for every $1 spent developing a model, plan about $3 for change management. McKinsey separately attributes as much as 70 percent of development effort to wrangling and harmonizing data.

Vista and ICONIQ’s survey of 202 enterprise AI leaders found inference runs 20% of cost pre-launch and 23% at scale, while talent runs 32% pre-launch and 26% at scale. Gartner cost ranges show how wide TCO swings with ambition: a coding-assistant deployment runs roughly $100K–$200K in initial cost, while a custom LLM runs $8M–$20M.

Evaluate build-versus-buy choices through total cost of ownership. Self-hosting open-source models trades API fees for infrastructure and engineering time, and the trade doesn’t always favor self-hosting.

Hidden line items to track

Finance teams routinely omit two categories from AI budgets even at well-run companies.

The first is retrieval. RAG pipelines carry their own meters: Bedrock pricing sets Knowledge Bases at $5.00 per GB of raw data per month for index storage, $1.00 per 1,000 standard retrieval calls, and $4.00 per 1,000 agentic retrieval calls plus $1.00 per 1,000 underlying calls. Every retrieved chunk then re-enters the model as billed input tokens, and vector stores like Pinecone or Qdrant add their own infrastructure line.

The second is energy. Google’s Gemini telemetry puts the median text prompt at 0.24 Wh, and the IEA projection has global data center electricity demand more than doubling from 415 TWh in 2024 to roughly 945 TWh by 2030. No regulation yet mandates an AI-specific energy line item, but EU CSRD captures cloud AI energy through Scope 2 and Scope 3 disclosures for covered entities. California SB 253 requires covered entities with annual revenue over $1 billion to begin Scope 1 and Scope 2 reporting in 2026 and Scope 3 reporting in 2027. The Green Software Foundation ratified the SCI for AI standard in December 2025; the standard defines per-token emissions units, and the FinOps Foundation’s Sustainability capability recommends folding carbon data into cost allocation now.

How to forecast AI spend when usage is unpredictable

Only 11% of organizations can forecast AI spending within ±10%, according to CFO Dive coverage of a Mavvrik/Benchmarkit study, and forecasting misses reach 11–25% for 56% of companies. Gartner is blunter: CIOs who don’t understand how GenAI costs scale could make forecasting errors of 500%–1,000%, so proofs of concept should test cost scaling alongside technical performance.

Four methods hold up against that volatility:

  • Rolling forecasts: The FinOps Foundation’s forecasting guidance recommends weekly or monthly cadences for AI because predictability is low and costs move fast. AFP reports 48% of organizations already run rolling forecasts on an 18–24 month horizon.
  • Driver-based models: The Deloitte formula is: AI cost scenarios equal “(avg. tokens/user × user volume × cost/token) × (some model mix factor).” Forecast the drivers engineers can observe (users, requests, tokens per request, model mix), rather than last year’s ledger.
  • Scenario planning: Forrester’s TEI method models a range from low- to high-impact outcomes, risk-adjusting benefits down 10–15% and costs up 10–15%. Finance should pre-authorize three scenario bands to accommodate variance.
  • Zero-based budgeting, applied selectively: In a ClickUp CFO interview, Dan Zhang argues “AI is forcing companies back to zero-based budgeting” because last-year-plus-20% no longer works. Match Group’s CFO requires a business case with clear cost savings or efficiency gains before any material AI spend. With no reliable prior year to anchor on, finance requires each AI line owner to justify spending from zero.

Attributing and allocating AI costs across teams and products

Finance and engineering teams often fail to account for 20–30% of AI spend because investments fragment across vendors, tools, and commercial models. McKinsey’s fix is to tag usage to business unit, product, use case, workflow, and cost center to “move from aggregate vendor invoices to meaningful unit economics, such as cost per task, case, code review, or customer interaction.”

The mechanics run at three layers. At the gateway layer, LiteLLM virtual keys track spend per key, user, team, and organization automatically, and Helicone attributes per-user costs through a Helicone-User-Id header. At the platform layer, CloudZero dimensions allocate AI costs across teams and products and compute unit economics like cost per customer or per feature, while Datadog cost management ingests OpenAI and Anthropic billing and splits shared costs by team and by project or environment tags. At the standards layer, FOCUS 1.4 models tokens as SaaS virtual currencies so chargeback works across vendors.

Watch for shadow spend while you set this up. The same Mavvrik study found 98% of engineering organizations use AI coding assistants, yet only 42% include developer AI tool spend in their AI cost reporting. Chargeback that misses the fastest-growing category isn’t chargeback.

Viktor’s own usage is usage-based and carries the same tags as every other consumption line, so it lands in allocations by team and by product or use case rather than in an untagged bucket. @Viktor pulled the per-team spend from the connected billing and observability tools, including LiteLLM and Helicone along with CloudZero. He grouped it by tag and posted the allocation in the finance channel, including his own line.

Applying FinOps to AI workloads

AI cost management has become the center of the FinOps discipline: 98% of State of FinOps 2026 respondents now manage AI spend, up from 63% in 2025 and 31% in 2024. On March 19, 2026, the FinOps Foundation published the FinOps Framework 2026 and formalized Executive Strategy Alignment with AI unit measures such as cost per token and cost per inference. The Foundation’s AI Scope maps AI spend to four domains: understand usage and cost, quantify business value, optimize usage and cost, and manage the practice.

Two adaptations matter most for LLM workloads. First, monitoring shifts from infrastructure utilization to virtual currency burn rates across credits, DBUs, tokens, and provider-specific units because usage no longer maps cleanly to servers. Second, review cycles compress: budget variance targets under the FinOps Budgeting capability run 20% at Crawl maturity, 15% at Walk, and 12% at Run, and hitting even the Crawl target requires continuous review rather than quarterly review.

Ownership has to operate at two levels. Finance and engineering share operational ownership: finance owns the scenario bands and variance thresholds, while engineering owns the tagging and routing decisions. Each AI use case also gets a named owner accountable for its unit costs. Put all three in the same weekly review of the consumption line.

Governance controls: spending caps, alerts, escalation, and circuit breakers

One team’s unbounded retry loop on a malformed document ran from Friday 6 p.m. to Monday 9 a.m. and a runaway incident cost $39,847.20. In a LangChain incident, a multi-agent setup where the verifier agent lacked a ‘done’ predicate looped for 264 hours and burned $47,000. A third onboarding incident spawned 5,080 agent sessions in 15 hours, consuming roughly 400 million input tokens for $1,200+.

Agentic workloads make these failure modes structural, because every turn re-sends the full conversation history. One engineering benchmark found a 10-turn agent conversation costs closer to 55× a single turn, rather than 10×, and an eight-model arXiv study found agentic coding tasks consume 3,500× more tokens than code reasoning tasks, with runs on the same task varying up to 30×. GitHub cited exactly this when it moved Copilot to usage-based billing in June 2026, saying it could no longer sustainably charge the same amount for a quick chat question and a multi-hour autonomous coding session. Per-seat budgeting cannot absorb that spread.

Build two groups of control:

  • Prevent runaway spend: Set hard caps at the provider and agent level. Google Gemini enforces spend-based rate limits ranging from $10 to $200 per rolling 10-minute window, with $50 at Tier 2. Anthropic’s Agent SDK exposes max_turns and max_budget_usd, and AWS Bedrock AgentCore defaults to a 75-iteration limit per invocation. Set every one of these before production traffic. Circuit-breaking matters more than it sounds: during one API outage, uncontrolled retries burned about $2 in tokens in 30 seconds; a circuit breaker cut that to roughly $0.01 in a documented circuit-breaker result.
  • Detect and resolve anomalies: Set anomaly alerts below the cap. Helicone alerts start at 50% and end at 95%, including an 80% warning, with configurable windows from 30 minutes to 30 days. Alerts at the cap are too late; alerts at 50% give the owner time to look. Then establish an escalation playbook. Name the person the cost-anomaly alert pages and identify who can kill the workload. Document the retry policy and assign the incident report.

When an anomaly alert fires in Slack or Microsoft Teams, Viktor pulls the spend breakdown by key and model from the connected billing and observability tools, identifies the owner of the flagged workload, and posts a structured summary in the channel. He then follows up directly with that owner, tracking the thread until the owner marks the item closed. Before taking any high-stakes action, such as killing a production workload, Viktor asks for explicit human confirmation, keeping that decision with the person accountable for it.

Optimizing costs through model selection and prompt efficiency

The Pragmatic Engineer documented one organization cutting AI costs 30% by changing its default model. The spread between flagship and small models makes that lever enormous:

Provider

Flagship (per MTok, input/output)

Small model (per MTok, input/output)

OpenAI

GPT-5.6 Sol: $4.00 / $20.00

GPT-5.6 Luna: $0.20 / $1.20; GPT-5 nano: $0.05 / $0.40

Anthropic

Claude Opus 5: $5.00 / $25.00 (Claude Fable 5.1: $10.00 / $50.00)

Claude Haiku 4.5: $1.00 / $5.00

Google

Gemini 2.5 Pro: $1.25–$2.50 / $10.00–$15.00

Gemini 2.5 Flash-Lite: $0.10 / $0.40

Sources: OpenAI pricing and Google pricing pages, plus Anthropic’s official pricing documentation, current as of September 2026.

Five levers exploit that spread:

  • LLM routing: The RouteLLM study cut costs 3.66× on MT Bench while retaining 95% of GPT-4 quality; the FrugalGPT study achieved up to 98% cost reduction while matching GPT-4. In production, Intercom cut costs 20% through a Fin model migration from GPT-4o to GPT-4.1, and Ramp’s routing saved over 25% while reducing error rates.
  • Quantization: An ACL quantization study of the Llama-3.1 family found INT4 quantization delivers 2–3× lower cost per query with at least 98% accuracy retention on the 8B model, and 5–7× reductions at 405B scale.
  • Fine-tuning smaller models: Shopify’s fine-tuned Qwen3-32B replaced a frontier model in its Flow agent at 68% lower cost and ran 2.2× faster. One caveat from the same report: the initial deployment showed a 35% lower workflow activation rate until retraining on production data closed the gap. Budget for that iteration.
  • Prompt caching: Anthropic prices cache reads at 0.1× base input price, and OpenAI’s cached input runs 90% below standard on GPT-5-class models. A PwC agent benchmark across 500+ sessions measured real-world savings of 41–80%. Savings apply to input tokens only, and cache writes cost extra, so break-even needs at least one read within the TTL.
  • Prompt efficiency as behavioral control: Headroom’s JSON context compression saved 30% of tokens immediately and an estimated $700,000 over five months. Shorter system prompts and trimmed, compressed conversation context are engineering habits, and they compound.

One warning before you chase the cheapest per-token rate: the FinOps Foundation’s model-cost guidance cautions that a less capable model may need longer prompts, more retries, more verbose outputs, or added human review. Measure the total cost of each successful outcome.

Viktor is model-agnostic, and the team chooses which models he runs on. Routing him to a different model and caching a stable prompt prefix move his cost line. Trimming conversation context has the same effect. The team reviews his spend in the same weekly consumption review by cost per finished deliverable rather than cost per token.

Negotiating pricing and contract protections with AI vendors

CFO Dive reported that AI software prices rose 20–37% in Tropic’s December 2025 data, and Redress Compliance found vendors opening 2026 GenAI renewals with 30–50% increases while 72% of reviewed contracts carried auto-renewal language. The negotiating room is equally large on the other side: negotiated Anthropic enterprise discounts average 28% with a 15–45% range, and Vertice puts the average AI-category discount at 43% versus 33.8% across all software.

The sources below provide the baseline benchmarks:

  • Commit below run rate: Redress contract guidance recommends committing to 75% of measured usage with quarterly true-ups and requesting 25–30% rollover of unused balance before signature. Redress found early annual commitments overshot actual usage by 20–40%.
  • Cap escalation in writing: The renewal playbook puts effective protection at 0–5% annual increases in years two and three and constrains the next renewal to 3–7%.
  • Lock the unit across model classes: Redress contract guidance recommends per-token rate locks covering model classes, most-favored pricing on successor models, and a minimum 12 months’ notice before a contracted model class loses availability.
  • Unbundle AI add-ons: Procurement teams that evaluate AI features separately from the platform renewal secure prices 20–35% below bundled proposals, according to VendorBenchmark renewal data. Push back on legacy SaaS vendors folding AI uplifts into the base renewal.
  • Bring a credible alternative: An AI procurement advisory found that structured competitive pressure produces a 4–8 percentage-point discount lift and puts the gap between initial offers and achievable pricing at 20–35%.

Metrics and unit economics for tracking AI ROI

The FinOps Foundation defines two unit-economics tiers. Tier 1 measures resource efficiency through cost per token or inference, with cost per API call as another option. Tier 2 measures business outcomes through cost per case resolved and cost per assist, with cost per agent action as another option. Mature from the first toward the second; McKinsey makes the completed business outcome the unit of governance.

OpenAI CFO Sarah Friar frames the target metric as “the full cost of producing a successful outcome, measured against the value that outcome creates,” noting that cheaper tokens may still require repeated attempts and added human-review time. Concrete reference points exist for some workloads; Anthropic reports Claude Code averages of $13 per developer per active day and $150–$250 per developer per month across enterprise deployments. No credible cross-industry benchmark yet exists for cost per active user or cost per feature, so build internal baselines and track trend rather than chasing a published target.

Observability tooling makes these metrics live instead of retrospective. CloudZero AI Signals captures AI spend in seconds, Datadog runs cost anomaly detection across OpenAI, Anthropic, Bedrock, and Vertex AI, and Langfuse attributes cost per generation by user, session, model, and feature. The gap between having this and not having it is stark: only 26% of organizations have real-time cost visibility into AI at scale, according to KPMG, and Gartner ROI research reports 84% of CFOs struggle to measure AI ROI. Gartner analyst Lydia Clougherty Jones attributes AI value erosion to cost creep caused by human factors.

Putting a budget process in place

A weekly finance-and-engineering review anchors the process. Finance and engineering teams can implement the framework in six steps over one quarter:

  1. Segment spend into the four categories: fixed licenses, variable consumption, compute, and talent. Give the variable line its own owner.
  2. Set hard caps and turn limits at every provider and agent framework, then layer anomaly alerts at 50% and 80% of each budget.
  3. Assign attribution before scale: engineering should enforce virtual keys or tags per team, product, use case, and workflow at the gateway so untagged spend can’t exist.
  4. Forecast on a rolling monthly cadence with three scenario bands, using tokens-per-user drivers instead of last year’s actuals.
  5. Pick one Tier 1 metric (cost per request) and one Tier 2 metric (cost per resolved case, per report, per PR, or per agent action) for each use case.
  6. Review the consumption line weekly with finance and engineering in the same room.

The review itself is work someone has to do every week, and it’s work you can hand off. Viktor is the AI employee that lives in Slack and Microsoft Teams, connects to 3,200+ tools, and does the work. For example, on Monday morning, ask @Viktor to pull the month’s AI spend from the connected billing and observability tools. He groups it by team tag, compares it with the approved scenario band, posts the variance report in #finance, and follows up with line owners until the open items close. Viktor runs the workflow on schedule, and his usage-based pricing keeps his spend legible in the same report. For a governed workspace-wide rollout, contact the enterprise team.

FAQ

What are the main cost categories in an AI budget?

Four: fixed SaaS and license fees, variable token and API consumption, compute and infrastructure, and talent and data. The variable line needs weekly review; the others follow normal budget cycles. Remember that the model bill is the minority of TCO, with people, infrastructure, data work, and change management taking the larger share.

What is the 30% rule in AI?

There is no single official “30% rule.” The phrase refers to two separate findings that happen to share the same number. First, on cost structure: Presidio published in 2026 that the AI services bill, the API line itself, is roughly 30% of total AI cost, with the remaining ~70% sitting in orchestration and retrieval, observability, guardrails, and the rereads and redos; related analyst estimates put the model share lower still, at about 15% (McKinsey) to 20–23% (Vista/ICONIQ). Second, on project survival: Gartner predicted that organizations would abandon at least 30% of generative AI projects after proof of concept by the end of 2025 because of poor data quality, weak risk controls, escalating costs, or unclear business value. Both findings point in the same direction. Budget the API line as the minority of the bill and size the rest of the stack, including retrieval, observability, guardrails, and change management, deliberately.

How do you forecast AI spend that changes every month?

Use rolling forecasts revised monthly and driver-based models built on tokens per user times user volume times cost per token. Set three scenario bands rather than a single number. Apply zero-based justification to each new AI use case, since there’s no reliable prior year to extrapolate from.

How should AI costs be attributed across teams?

Enforce tagging at the gateway with virtual keys per team, product, use case, and workflow, then run allocation and chargeback through a FinOps platform. Include developer AI tools in the reporting scope; most organizations currently leave them out.

Should we set hard spending caps on AI workloads?

Yes, at both the provider level (spend-based rate limits) and the agent level (turn limits and budget parameters). Documented runaway incidents ranging from hundreds of dollars to $47,000 trace to missing termination conditions, deduplication, retry bounds, or circuit breakers.

When does it make sense to use smaller or self-hosted models?

Route to smaller models whenever your evals show quality holds; routing research retains roughly 95% of flagship quality at a fraction of the cost, and the displayed small-model rates run roughly 5–80× below flagship rates depending on the provider and token type. Self-host when your infrastructure and engineering costs beat the managed API premium, which is not a given.

What can you negotiate in an AI vendor contract?

Escalation caps of 0–5% in-term and 3–7% at renewal, commitments set at 75% of measured run rate with true-ups, rollover of unused balance, per-token rate locks with successor-model protections, and 12 months’ notice on model deprecation. Unbundling AI add-ons from platform renewals and bringing a credible alternative both move price materially.

Should I pay per seat or per token, and how do I choose?

Pay per seat when usage per person is predictable and bounded — a fixed daily workload where the seat price caps the downside at seats times price. Pay per token when work volume varies widely across people and tasks, because a seat price averaged across a light user and a heavy agent user either overcharges one or bankrupts the vendor.

GitHub moved Copilot to usage-based billing in June 2026 because a quick chat question and a multi-hour autonomous coding session could no longer cost the same. Anthropic reports Claude Code averages $13 per developer per active day and $150–$250 per developer per month across enterprise deployments. Bain found about 10% of hybrid AI meters are outcome-based, with the rest effort- or output-based.

Price the seat against your own measured monthly token cost per user. Take the seat deal when it lands below that number with headroom. Take metered pricing with a hard cap when it does not.

Which metrics prove AI ROI?

Start with cost per request and cost per token, then graduate to outcome metrics: cost per resolved case, per report delivered, per merged PR, or per agent action. Judge each use case by comparing the total cost of a successful outcome with the value that outcome creates.

Is it worth paying for AI?

Evaluate whether AI pays off separately for each use case. McKinsey’s 2026 State of AI found 37% of respondents attribute at least some EBIT impact to AI, yet only roughly 6% qualify as high performers attributing 5% or more of EBIT to AI. Deloitte’s October 2025 survey of 1,854 executives found just 15% of gen AI users report significant measurable ROI and only 6% report payback under one year. Gartner reports 72% of CIOs say their organizations have broken even or lost money on AI. The pattern is consistent: evaluate whether the spend pays off for each use case. Fund use cases with a named owner and a measurable outcome, and kill the ones that miss. Compare the total cost of producing a successful outcome with the value that outcome creates, and stop funding any use case that cannot clear that bar.

What tooling tracks LLM costs in near real time?

Provider consoles set caps; gateways like LiteLLM and Helicone track spend per key and user as requests flow; observability platforms like Langfuse and Datadog attribute cost per trace and detect anomalies; FinOps platforms like CloudZero and Finout roll AI spend into unit economics alongside cloud costs. Pick one at each layer rather than trying to make a single tool do all four jobs.

Can ChatGPT make a budget?

ChatGPT can build the model structure: a token-driver formula (tokens per user × user volume × cost per token × model mix) and scenario bands. It can also build a variance template, and you paste the numbers in yourself. It has no connection to your OpenAI, Anthropic, or Bedrock billing, your gateway spend per key, or your FinOps platform, so the figures are only as current as your last copy-paste, and published token prices change. Viktor is the AI employee that lives in Slack and Microsoft Teams, connects to 3,200+ tools, and does the work: he pulls the month’s spend from the connected billing and observability tools, compares it against the approved scenario band, and posts the variance report in the finance channel. The decision rule is straightforward: use ChatGPT when the job is structuring the model, and use Viktor when the job is producing the actual monthly numbers.

4.9 on G2G2 rating: 4.9 out of 5 stars

One hire. The output of a team.

Viktor works nights, remembers every decision, and connects to 3,200 tools. Start free with $100 in credits, then $50 a month.

Get Started for Free