Task-sold agents (quotes, approvals, verified runs) need cost per verified
outcome, not seat funnels and login health scores, and CapEcon models that economics so you can turn
it into weekly ship / throttle / kill decisions on accounts and capabilities, ranked by cost of
leaving them live.
This is an open-source teaching repo with a synthetic warehouse plus optional Data
Connect overlays, so you can clone it and run locally.
For task agents, the real unit is a verified outcome per job
(Outcome Definition), and login alone does not count.
Health
ChurnZero-style usage proxies
Instead of usage proxies, look at TTFV, tourist accounts, and whether
seats keep a weekly delegation habit.
Margin
MRR vs usage in another tab
Margin shows up as CPSO, floor vs cap, and the retry or HITL tax on
multi-step runs, not MRR sitting next to a token bill in another tab.
Metering shows the token bill, but it does not tell you if the task was worth shipping.
One honest decision object
A GrowthDecisionRecord is
a priceable row, not an ops alert, so Radar can sort on headline dollars while the card still shows unit economics
and honesty about the cost basis.
primary_metric_usd Headline sort: cost of leaving live (window total)
floor_usd Unit economics: cost per verified outcome where applicable
Pricing dialect follows the preset (internal budget, product SKU, or marketplace take), and
the schema plus contracts live in
ontology/
and
docs/contracts.md.
What CPSO, WTP, and margin leakage mean
CPSO (cost per successful outcome)
Total token, tool, and infra cost attributed to a workflow, divided by the number of runs that reach a
verified successful end state (not merely “the trace finished”). Rising CPSO or high variance
usually means retry loops, context bloat, or expensive model routing. ChartMogul will not show this; LangSmith
alone will not join it to ARPU.
Floor vs headline $floor_usd is the unit cost of producing one
verified outcome when you can measure it.
primary_metric_usd is the Radar headline
(cost of leaving live over the window). They stay separate on purpose so sort order and unit economics do not
get mashed together.
WTP (willingness to pay)
An upper bound or cap next to the floor: what you can sensibly charge or budget for that outcome, often compared
to seat ARPU or SKU value. CapEcon does not invent survey WTP from billing; in rigorous mode a data-derived cap
can come from subscriptions or outcomes (associational), and Packaging Lab / Fit demand can estimate
surplus-optimal list when the panel is strong enough.
Margin leakage
Power users (often the top few percent by token volume) whose revenue minus fully loaded serve cost is near zero
or negative under flat or loosely capped pricing. The board may still see expansion and NRR while ops should
throttle or reallocate. In the taxonomy this is
margin_leakage.
Retry amplification
Average model and tool invocations per logical user intent. Values much greater than 1× quantify the hidden
multi-step cost explosion that CPSO catches when “success” still burns margin.
Where this lives in the app
Run Economics (Price cluster) puts CPSO next to WTP-style caps and power-user margin leakage. Radar and Version
Gate surface the same numbers on decision cards when the workspace has data.
What your agent review could look like
Each row is a GrowthDecisionRecord
that says what's wrong, what it may cost to leave live, and what to do, with policy provenance from
semantics.yaml.
ChartMogul will not show which capability is hurting you, and LangSmith will not show
the dollar cost of leaving it live.
The call
What to ship, throttle, or kill
Ranked GrowthDecisionRecords: headline $, floor per outcome, ops vs commercial.
assistant_heavyproduct_skuseed 428 acct GDRs12 cap GDRs
62Yellow
Agentic health
Health score
62 (yellow)
CPSO
$0.42
TTFV
4.2d
Unattributed
18%
ACC · acme_west
$84,200
ACC · globex_ops
$61,500
ACC · initech_ai
$44,800
ACC · umbra_corp
$31,200
ACC · wayne_labs
$18,900
Cost of leaving live (USD)
Account · Leaking
acme_west
3 exceptions · activation_leak, habit_collapse
Activation or habit problems: value unfinished for many seats.
Paying seats never hit weekly delegation, and habit collapse correlates with churn on
this seed.
Commercial: reallocate (routing tax before raise_list)
Synthetic demoThis is a teaching warehouse you clone and run locally (authored priors, not
production telemetry).
Leaking
Activation or habit unfinished for many seats.
Destructive
Correlates with churn; throttle or rollback.
Uneconomic
Margin-negative even if UX looks fine.
Healthy
Safe to keep shipping.
Two dashboards, one missing meeting
Keep LangSmith for debugging, ChartMogul for board NRR. However, neither emits a ranked
ship / throttle / kill object on the join Account → Run → Outcome → Subscription.
Misses whether the paying account churned, and whether a successful run was a verified outcome.
CapEcon
Join Account, Run, Outcome, Subscription.
Record verdict, cost of leaving live, YAML rule, override.
Call ship, throttle, kill.
Revenue / CS
Tools: ChartMogul, ChurnZero, Stripe.
Sees MRR, NRR, health scores, cancel cohorts.
Misses CPSO, retry loops, HITL takeover, which capability to throttle.
Prefer OTel-compatible exports so ingest stays framework-agnostic.
Paid ChartMogul account with 92% LangSmith success and no verified outcome in 14 days.
Call: tourist / activation_failure
NRR is up while one capability spikes tokens.
Call: margin_leakage / run_cost_blowout, throttle
A Slack trace link plus a red health score, with no shared object.
Call: one GDR, the YAML cited, flywheel outcome 14 days later
Challenge
Traces
Revenue / CS
CapEcon
Cost / margin
Cost per span; no ARPU join
MRR fine while serving cost explodes
CPSO vs ARPU; throttle uneconomic capabilities
Visibility
“What did this trace cost?”
“What is NRR?”
Unattributed spend; cost heatmap by step × cohort
Activation
Run success ≠ first win
Paying customer still in “trial”
TTFV; paying-but-dormant; tourist GDRs
Switching costs
Tool-call graph, not rip-out risk
Churn reason unknown / competitor
Connector blast radius
Trust / reliability
Evals on traces
Health score drops after the fact
HITL trend; trust_break → rollback
Opaque success
Completed trace ≠ trusted SOP
Login/usage health
Verified outcome rate; flywheel write-back
Synthetic teaching modelThe warehouse scaffolds the join, and economics use synthetic runs and seat priors
today. When you have files, Data Connect overlays OTel GenAI JSONL, Langfuse exports, or CSVs into the same
Workspace schema (metadata only: prompt bodies are scrubbed), with no live LangSmith or ChartMogul connector
in this demo.
Common questions
Can’t I join this in Looker/dbt? You can SQL the join, but CapEcon’s IP is the
exception taxonomy, YAML policy, and auditable decision record.
Is this observability? No: keep LangSmith or your OTel pipeline for traces,
while CapEcon consumes trace-shaped facts and emits decisions.
Do I need OpenTelemetry? Prefer OTel GenAI attributes when you can (tokens,
model, agent id), and Langfuse export also works; full nested OTLP protobuf JSON is not the path today.
Do you replace my stack? No: ingest later via Data Connect, and decide now on
the teaching model.
Setup → The call → Price → Learn
The warehouse is synthetic by default, and Data Connect shows the path to real exports
(OTel GenAI JSONL, Langfuse, CSV, Vision) without claiming live connectors.
Setup
Product Profile (presets switch ontology and priors)
Data Connect: OTel GenAI JSONL (prompts scrubbed), Langfuse, CSV, Vision
Outcome Definition (verified success, not login)
The call
Version Gate, Decision Inbox, Radar
Activation, Trust, Connectors, Subgraph, Version Compare
Marketplace Radar and Clinical Radar on matching presets
Price
Run Economics: CPSO, WTP caps, power-user margin leakage (defs)
Math Lab · Packaging (demand curve teaching)
Learn
Executive Summary, Experiments, Agentic Flags
Outcome Flywheel, Math Labs (binomial, power, CLV, drift)
Config
Integrations
Semantics Console, Taxonomy Browser, Record Inspector
Legacy analytics archive (reference only)
Machine-readable policy instead of hardcoded heuristics
Taxonomy → semantics → schema → record.
Change the verdict rules in YAML and the Radar reclassifies without a code deploy.
The ontology/ package defines exception categories, vertical-specific
semantics.yaml action maps, and JSON Schema contracts for
GrowthDecisionRecord. See
ontology/
and
docs/contracts.md.