A CTO enters a quarterly budget review with several versions of the same story. Finance has invoices from AI vendors. The platform team has token and model usage reports. Engineering has adoption dashboards. Product leaders have a list of features that shipped.

Every report is accurate. None of them explains the investment.

AI usage increased. The bill increased. Some teams believe they are moving faster. But nobody in the room can say which work generated the cost, what the organization received in return, or where the next dollar should go.

Token totals tell you what AI consumed. They do not tell you what the organization gained.

Four systems, four answers, and no way to reconcile them

The gap in that budget review is not a reporting gap. Every system in the room is doing its job. Finance sees spend. AI vendors see consumption. Engineering systems see commits, pull requests, reviews, tests, deployments, and defects. Product systems see features, initiatives, and business priorities.

What is missing is the join. No single system knows that a particular agent run, on a particular model, in a particular repository, produced the pull request that shipped the checkout redesign. That relationship is the thing leadership is actually asking about, and it does not exist in any of the four sources.

It used to exist, incidentally, when AI was purchased like software. A company bought seats, assigned licenses, and reviewed utilization at renewal. One tool, one line item, one owner. Attribution was trivial because there was almost nothing to attribute.

That has changed. Engineering organizations now combine seat-based subscriptions, included usage allowances, token-based model charges, usage credits, premium model consumption, direct API costs, and autonomous agent runs – often several of them inside a single delivery cycle, each metered differently. The cost of one engineering task now moves with the model selected, the context loaded, the number of tool calls, repeated attempts, parallel agents, and the human work required afterward.

Those mechanics deserve their own treatment, and we cover them separately in AI token spend management: what efficiency actually means. The point here is narrower: variability is what destroyed attribution. When cost was fixed per seat, the total was the answer. Now the total is the least useful number available.

What attribution means hereAttribution is the join between a cost event and the engineering work it paid for. Without it, AI spend is a number with no subject - defensible as a bill, useless as a decision.

Why invoices and vendor dashboards stop too early

An invoice can tell you which vendor charged the organization and how much was billed. A vendor dashboard may add seats, requests, tokens, credits, models, or individual user activity.

Those views are necessary. They are not sufficient to make an engineering decision.

They cannot tell you which feature generated the cost, which repository was involved, whether the output was accepted, how much review it required, whether the work reached production, or whether the code remained stable after release. The problem is not inaccurate billing data. It is missing operating context.

Illustrative
Vendor view
Seats480
Tokens41M
Model spend$284K
Agent runs1,240
Missing context
What leadership still needs
Which team and feature?Attribution
What reached production?Output
How much human review?Effort
What changed afterward?Outcome
Engineering value
DeliveryCycle time
QualityRework
Human effortReview load
BusinessPriority
The AI spend visibility gap Billing and usage data explain consumption. Engineering context explains whether the investment created value.Figures shown are an illustrative example, not customer data.

Consider what that missing context costs you in practice. A usage dashboard shows that one team consumed twice as many tokens as another. That single fact supports two opposite conclusions: the first team was wasteful, or the first team was doing more complex and more valuable work. Without attribution there is no way to choose between them, so most organizations quietly choose the one that matches an existing bias.

A renewal report shows that a tool has high adoption. Adoption is not a result. It does not reveal whether the tool reduced delivery time, shifted work into review, increased rework, or supported the product areas that mattered most this quarter.

DORA’s 2025 report, State of AI-assisted Software Development, found that the largest returns on AI investment came not from the tools themselves but from a deliberate focus on the underlying organizational system — platform quality, workflow clarity, team alignment. That finding has a practical consequence that is easy to miss: if returns depend on the system around the tool, then a per-tool view can never locate them. You can only see the system if you can attribute across it.

What AI cost attribution actually requires

A usable attribution model connects five layers. Each one answers a question the layer above it cannot.

01

Origin

Identify who or what generated the cost: a developer, team, assistant, agent, model, workflow, API, or integration.

02

Work

Connect the cost to the feature, initiative, ticket, pull request, repository, migration, incident, test, or documentation task it supported.

03

Output

Determine what the activity produced, and whether that output was accepted, reviewed, merged, released, or used.

04

Outcome

Compare cost and output with delivery, review effort, rework, code stability, human effort, and strategic value.

05

Decision

Use the evidence to expand, guide, govern, reallocate, or retire a tool, model, agent, or workflow.

A spend dashboard becomes valuable only when it improves one of those decisions. Everything before the fifth layer is instrumentation.

O
Origin

Tool, model, agent, team

W
Work

Feature, PR, repository

Output

Code, tests, review

Outcome

Delivery, quality, effort

D
Decision

Expand, guide, govern

Worked example
Claude Code agent Checkout initiative 12 PRs and tests Faster cycle, more review Expand tests, guide backend
Five layer attribution model Move from a raw usage event to the leadership decision it should inform.The bottom row is an illustrative worked example.

One warning about the second layer, because it is where most attribution efforts quietly fail. The join is only as reliable as the identifiers that survive the journey from the AI tool to the merge. If an agent commits under a shared service account, if sessions carry no durable identifier, or if pull requests are not linked to tickets, then the cost data and the work data cannot be reconciled no matter how good the reporting layer is. Attribution is a data-model problem before it is a dashboard problem.

How Milestone attributes AI spend to engineering work

Milestone does not treat token cost as an isolated financial metric. It connects AI spend with the engineering signals needed to understand what the organization received in return.

One ledger across the AI portfolio

Milestone AI Spend Hub brings tool, license, seat, token, model, and agent costs into one connected ledger. Engineering and finance leaders can compare spend across tools and teams, identify underused seats, and enter renewal discussions with evidence rather than four separate vendor reports.

Attribution to features and initiatives

AI Feature Spend Attribution connects AI usage with planning data, Git activity, pull requests, reviews, delivery signals, quality fixes, and post-release work. That makes it possible to examine what a feature cost across its full delivery cycle, rather than assigning the bill to a broad technology budget.

Human and AI effort in the same cycle

AI tool cost is only part of the investment. A feature also carries engineering time, prompt preparation, human review, correction, rework, and maintenance. Attribution that stops at the token charge will systematically understate what a feature cost.

Human and Agent Spend connects AI activity with human effort and the features teams ship. It shows whether AI removed work, redirected it, or moved it to another part of the delivery cycle.

The questions attribution makes answerable

Once cost is joined to work and outcome, the conversation changes:

  • Which tools produce the strongest results for each team?
  • Which features justify greater model spend?
  • Where is AI reducing implementation time but increasing review effort?
  • Which agents produce accepted output without creating quality problems?
  • Where should usage expand, change, or stop?

The objective is not to spend less. It is to make AI investment explainable.

An attribution checklist for engineering leaders

Six questions. The answers describe how far your organization can currently trace a dollar.

Do we have one inventory of every AI tool, model, seat, API, and agent?Include formal purchases, team experiments, direct model access, and autonomous workflows.
Can we attribute spend by team, contributor, tool, model, and agent?A company total is not enough to identify where action is needed.
Can we connect spend with features, repositories, pull requests, and workflows?Without this join, usage stays detached from the work it was meant to improve.
Do the identifiers survive from the AI tool to the merge?Shared service accounts, unlinked pull requests, and sessions without durable IDs break attribution at the source.
Do we compare AI cost with delivery, review, rework, and quality?Implementation speed is one part of the engineering result, not the result.
Can we explain which investments should expand, change, or stop?This is the standard that turns reporting into management.

A “no” does not prove that spend is being wasted. It shows that leadership does not yet have the evidence required to know either way.

Conclusion

AI engineering cost is now attached to how work is executed rather than to how many licences were purchased. That makes attribution a leadership capability, not a finance chore.

Invoices still matter. Vendor usage reports still matter. But neither can tell you whether AI accelerated an important feature, created additional review work, protected quality, or returned enough to justify the next increase.

Attribution connects cost to the complete engineering system:

Who or what used the AI.
Which work it supported.
What it produced.
What happened afterward.
What leadership should do next.

The organizations that handle this well will not necessarily have the smallest AI bills. They will be the ones that can explain where each additional dollar creates an engineering advantage.

Milestone AI Spend Intelligence

See what your AI spend is actually producing.

Connect token, tool, and agent costs with engineering contribution, delivery, quality, human effort, and business priorities.

Book a Milestone demo

FAQs

1. What is AI cost attribution?

AI cost attribution is the practice of connecting each unit of AI spend – a token charge, a seat, an agent run, an API call – to the engineering work it paid for, and then to what that work produced. It answers which team, feature, repository, and pull request a cost belongs to, rather than reporting a total by vendor.

2. Why can vendor dashboards not attribute AI spend to features?

Vendor dashboards are scoped to one product and to identities inside that product. They can report seats, requests, tokens, credits, and per-user activity, but they have no view of your planning system, your repositories, your review process, or your release pipeline. The join between a usage event and a feature has to happen outside the vendor.

3. What data do you need to attribute AI cost to engineering work?

At minimum you need usage records that carry a stable identity, and engineering records that share a key with them: contributor identity, repository, branch, pull request, ticket or initiative, and agent or session identifier. Attribution quality is limited by how reliably those identifiers survive from the AI tool through to the merge.

4. How is AI cost attribution different from cloud cost allocation?

Cloud allocation maps spend to infrastructure that already has owners and tags. AI spend attaches to human and agent activity instead, so the unit of cost is a session or a task rather than a resource. It also has to be compared with outcomes such as accepted output, review effort, and rework, which have no equivalent in a cloud bill.

5. How is Milestone different from a vendor usage dashboard?

A vendor dashboard shows activity and cost inside one product. Milestone connects spend across AI tools with engineering systems and outcomes, so leaders can see which features and teams the cost belongs to, what reached production, and where investment should expand, change, or stop.

Sources and further reading

  1. DORA — State of AI-assisted Software Development (2025)
  2. GitHub — About billing for GitHub Copilot in organizations and enterprises
  3. Anthropic — Claude Code: Manage costs effectively
Written by

Sign up to our newsletter

By subscribing, you accept our Privacy Policy.

Related posts

AI Agent Frameworks Comparison: Which One Fits Your Engineering Stack?
High Token Usage Is Not Waste. Unaccepted Output Is.
Aug 12, 2026

High Token Usage Is Not Waste. Unaccepted Output Is.

Your AI Token Bill Shows Spend. It Does Not Show Value
Aug 05, 2026

Your AI Token Bill Shows Spend. It Does Not Show Value

Ready to Transform
Your GenAI
Investments?

Don’t leave your GenAI adoption to chance. With Milestone, you can achieve measurable ROI and maintain a competitive edge.