Agentforce is no longer a promise
When this site first described Agentforce as a black box, the objection was about an announced direction. The scale has changed the question. In its fiscal 2026 fourth-quarter release, Salesforce reported $800 million in Agentforce annual recurring revenue, up 169 percent year over year, with 29,000 deals closed. A quarter earlier it reported 9,500 paid deals and 3.2 trillion tokens processed.
Those figures are vendor-reported, not independent proof of customer value. They are still important because they move Agentforce from an interesting product pitch into deployed digital labor. A system touching customer records, campaigns, service interactions, and pipeline at that scale has to be observable before it can be trusted.
That is also why the ROI question becomes more exact. “Is the AI working?” is too vague. A buyer needs to ask which sessions were handled, which actions were taken, how much the activity cost, whether the agent stayed within its permissions, and whether the resulting business outcome was incremental.
The first four questions are operational. The last one is causal. They belong in related but different systems.
Scale changes the cost of being imprecise. A small pilot can be manually reviewed. A system with tens of thousands of deals and trillions of processed tokens cannot depend on a champion remembering why an audience was changed. At that point, the record of the decision is part of the product, not an after-the-fact compliance exercise.
It also changes the renewal conversation. Procurement can ask whether the agent is available, adoption can report how often it is used, and operations can report whether it is healthy. The CMO and CFO still need a separate answer about whether the activity produced an outcome that the old process would not have produced.
The critique became a product
In June 2025, Salesforce announced Agentforce 3 with the Agentforce Command Center, described in the launch release as an observability solution for monitoring and optimizing agents. The product has since been expanded into a standing Agentforce Observability line.
This deserves a precise reading. Salesforce did build the missing window. Session traces, tool calls, action histories, errors, latency, usage, and health metrics make the machine more inspectable. For an operations owner, that is real progress. A team cannot investigate an agent failure it cannot see.
It also validates the original criticism. The market did not need another promise that autonomous work would be efficient. It needed an instrument panel. Salesforce shipped one because autonomous systems become difficult to sell at enterprise scale when their operators cannot answer basic questions about behavior.
But a window is not a verdict. The platform that runs the agent also defines the fields, summaries, and success views the buyer sees. Cross-check what the vendor dashboard includes against the data you need for renewal, finance, legal, and customer-impact decisions.
Ask the vendor to demonstrate an incident from beginning to end, not just the clean operating view. Can an operator retrieve the exact context the agent saw? Can they distinguish a tool failure from a policy refusal? Can they export the action history with stable identifiers? Can the business connect that action to an outcome outside Salesforce? These questions turn “observability” from a product adjective into a testable capability.
The dashboard is strongest when it helps an owner find the raw event. It is weakest when a polished aggregate replaces the event. Preserve both. Summary charts are for orientation; the underlying record is what lets an investigator reproduce a decision.
Observability is not attribution
Observability answers, “What did the agent do?” It can show sessions handled, actions taken, tools called, errors thrown, escalations, tokens consumed, and time spent. Those facts are necessary. They make an agent legible as a running system.
Attribution asks a different question: “What did the agent cause?” If Agentforce reprioritized a lead and the lead later bought, did the decision create demand, accelerate an existing deal, or merely reorder a pipeline that was already likely to close? If an agent changes a campaign while pricing, seasonality, and another launch change too, which part of the revenue movement belongs to the agent?
A useful boundary
What the system can show
Sessions and actions
Errors and escalations
Usage and latency
The vendor view records activity. The distinction matters because visible activity is not automatically evidence of a business outcome.
The missing object is the counterfactual: what would have happened if the agent had made a different decision, or no decision at all? No dashboard can recover that after the system has touched every audience. Holdouts, control regions, or other designed comparisons have to exist before launch.
This is the same distinction that matters in the broader agentic measurement problem. Participation is not causation. A platform can honestly report that an agent touched a customer journey while still being unable to show that the touch changed the outcome.
Consider a lead-scoring agent. It may surface a lead earlier, route it to a seller, draft a follow-up, and record a conversion. A platform report can attribute all four events to the agent. A defensible ROI analysis asks whether a comparable lead outside the agent's reach received a slower response and converted less often. Without that comparison, the report describes sequence, not lift.
There is a second boundary around ownership. Salesforce may know what happened inside its objects, but the buyer knows whether the opportunity was already in motion, whether a discount changed the deal, and whether another channel created the demand. Attribution crosses those boundaries. It cannot be delegated simply because the activity began in one platform.
The consumption wrinkle
Agentforce also changes the cost shape of marketing automation. Salesforce reports token processing at trillion-scale, and the broader Agentforce model uses consumption-linked economics. Cost therefore follows activity: more sessions, actions, and tokens can mean a larger bill whether or not the extra activity produces incremental value.
That creates a subtle reporting hazard. A dashboard that shows rising agent activity can make rising spend look like evidence of progress. It is not. Cost per conversation or cost per action describes throughput. The buyer needs cost per incremental outcome, measured against an outcome that would not have happened without the agent.
Consumption pricing is not inherently bad. It can align spend with use and make waste visible. It becomes a problem when the system's own activity metric is used as the denominator for its own success story. More work is not the same as more value.
Track the trend over a consistent window. If agent efficiency improves, cost per incremental outcome should fall or the outcome value should rise. If usage climbs while the control comparison stays flat, the system is buying motion.
Separate platform cost from outcome value in the ledger. Credits and tokens are inputs. Gross pipeline, attributed revenue, and resolved cases are outputs. Incremental revenue or avoided cost is the narrower result that can support an ROI claim. Keeping those layers apart prevents a high-usage month from looking like a high-return month.
Make the denominator durable. If one quarter uses “influenced pipeline” and the next uses “closed revenue,” the trend is a change in vocabulary rather than a change in economics. Decide which outcome matters, document the attribution window, and keep the calculation stable long enough for a skeptical reader to reproduce it.
The buyer measurement stack
The vendor observability layer should be your floor, not your complete measurement stack. Build four buyer-owned capabilities above it.
1. Exportable decision logs
Move agent decisions and actions into a warehouse you control. Preserve context, identity, permissions, alternatives, confidence, and the downstream object touched. Join those records to revenue, pipeline, retention, or service outcomes. If decision data cannot leave the platform, treat that as a procurement finding.
2. Fenced counterfactuals
Agree on untouched audiences, geographies, or time windows before the agent launches. Lock the exclusion in writing and monitor leakage. A holdout that is accidentally exposed is not a holdout, and a retroactive report cannot repair the lost comparison.
3. Incremental economics
Join credits, tokens, sessions, and implementation cost to the uplift measured against the holdout. Report cost per incremental outcome rather than cost per conversation. Keep the definition stable enough that finance can compare periods.
4. Independent behavioral review
Sample decisions outside the platform's roll-up. Review whether the agent used the right context, followed the intended policy, and took an action that a human owner can explain. This is complementary to the drift review, not a replacement for it.
Make the stack useful to more than marketing. Legal needs the decision record, finance needs the calculation, data teams need stable identifiers, and the workflow owner needs a way to stop or correct the agent. A measurement system that only produces a quarterly slide will not help when the agent fails at 2 a.m.
Run the review as a repeatable operating loop: export the prior period, sample decisions, check policy and context, compare the holdout, reconcile consumption, then record what changed. That sequence makes the dashboard one input into a decision, not the decision itself.
None of this is anti-Salesforce. The same controls apply to an internal agent or another vendor. The practical boundary is ownership: let the vendor make the machine observable, but do not outsource the definition of value.
Agentforce's command center solves the operational half of the black-box complaint. The economic half remains a buyer responsibility because it depends on your outcomes, your controls, and a counterfactual the platform cannot manufacture for you.

The dashboard is the floor, not the answer
The honest conclusion is not that Salesforce failed to build visibility. It is that visibility has a boundary. Agentforce's Command Center can become a dependable operational instrument if teams preserve raw events, test the export path, and use the same definitions across periods. That is a meaningful improvement over an agent that cannot be inspected at all.
The buyer still owns the harder sentence in the board meeting: “This change produced this incremental result under this test.” That sentence needs an untouched comparison, a decision record, a stable outcome definition, and a cost calculation that includes the work required to run the system. It cannot be inferred from sessions or actions, even when those numbers are impressive.
Use the platform's telemetry to find the decisions worth investigating. Use your own measurement layer to decide whether those decisions changed the business. If the two disagree, investigate the disagreement instead of choosing the prettier chart.
Agentforce is now large enough that this distinction is not academic. Its growth proves a market for autonomous work. It does not prove every autonomous action paid for itself. The Command Center makes the machine easier to watch. The ROI claim becomes credible only when the buyer can show what would have happened without it.
What to ask before the next renewal
Before accepting an ROI slide, ask for one trace from the agent event to the claimed outcome. Start with the campaign or workflow ID and inspect the context the agent received. Identify the exact decision, the permission that allowed it, the action that followed, the cost incurred, and the outcome record used in the report. Then inspect the control. If there is no comparable untouched group or pre-registered comparison, label the result as influenced activity rather than incremental return.
Repeat the test with a failure, not only a successful example. A trustworthy operating layer should let the team explain an escalation, a blocked action, a bad recommendation, and a human override. These cases reveal whether the telemetry is a usable record or a collection of aggregates. They also expose which parts of the workflow still depend on a person remembering what happened.
The most useful vendor conversation is concrete: show us the event, show us the export, show us the comparison, show us the cost, and show us the stop control. If those five requests produce five different owners, the system is not yet ready to carry an unqualified ROI claim.

