Capecon

Task-sold agents (quotes, approvals, verified runs) need cost per verified outcome, not seat funnels and login health scores, and CapEcon models that economics so you can turn it into weekly ship / throttle / kill decisions on accounts and capabilities, ranked by cost of leaving them live.

This is an open-source teaching repo with a synthetic warehouse plus optional Data Connect overlays, so you can clone it and run locally.

View on GitHub

What breaks when you run agents like 2016 SaaS

Unit of value
Seat, login, feature adoption

For task agents, the real unit is a verified outcome per job (Outcome Definition), and login alone does not count.

Health
ChurnZero-style usage proxies

Instead of usage proxies, look at TTFV, tourist accounts, and whether seats keep a weekly delegation habit.

Margin
MRR vs usage in another tab

Margin shows up as CPSO, floor vs cap, and the retry or HITL tax on multi-step runs, not MRR sitting next to a token bill in another tab.

Metering shows the token bill, but it does not tell you if the task was worth shipping.

One honest decision object

A GrowthDecisionRecord is a priceable row, not an ops alert, so Radar can sort on headline dollars while the card still shows unit economics and honesty about the cost basis.

  • primary_metric_usd Headline sort: cost of leaving live (window total)
  • floor_usd Unit economics: cost per verified outcome where applicable
  • cost_basis Weakest link: simulated → estimated → metered → invoiced
  • recommended_action Runtime: ship, throttle, rollback, hold
  • commercial_action Packaging: reallocate, raise_list, split_tier (reviewed separately)

Pricing dialect follows the preset (internal budget, product SKU, or marketplace take), and the schema plus contracts live in ontology/ and docs/contracts.md.

What CPSO, WTP, and margin leakage mean
CPSO (cost per successful outcome) Total token, tool, and infra cost attributed to a workflow, divided by the number of runs that reach a verified successful end state (not merely “the trace finished”). Rising CPSO or high variance usually means retry loops, context bloat, or expensive model routing. ChartMogul will not show this; LangSmith alone will not join it to ARPU.
Floor vs headline $ floor_usd is the unit cost of producing one verified outcome when you can measure it. primary_metric_usd is the Radar headline (cost of leaving live over the window). They stay separate on purpose so sort order and unit economics do not get mashed together.
WTP (willingness to pay) An upper bound or cap next to the floor: what you can sensibly charge or budget for that outcome, often compared to seat ARPU or SKU value. CapEcon does not invent survey WTP from billing; in rigorous mode a data-derived cap can come from subscriptions or outcomes (associational), and Packaging Lab / Fit demand can estimate surplus-optimal list when the panel is strong enough.
Margin leakage Power users (often the top few percent by token volume) whose revenue minus fully loaded serve cost is near zero or negative under flat or loosely capped pricing. The board may still see expansion and NRR while ops should throttle or reallocate. In the taxonomy this is margin_leakage.
Retry amplification Average model and tool invocations per logical user intent. Values much greater than 1× quantify the hidden multi-step cost explosion that CPSO catches when “success” still burns margin.
Where this lives in the app Run Economics (Price cluster) puts CPSO next to WTP-style caps and power-user margin leakage. Radar and Version Gate surface the same numbers on decision cards when the workspace has data.

What your agent review could look like

Each row is a GrowthDecisionRecord that says what's wrong, what it may cost to leave live, and what to do, with policy provenance from semantics.yaml.

ChartMogul will not show which capability is hurting you, and LangSmith will not show the dollar cost of leaving it live.

The call
What to ship, throttle, or kill
Ranked GrowthDecisionRecords: headline $, floor per outcome, ops vs commercial.
Synthetic demo This is a teaching warehouse you clone and run locally (authored priors, not production telemetry).
Leaking Activation or habit unfinished for many seats.
Destructive Correlates with churn; throttle or rollback.
Uneconomic Margin-negative even if UX looks fine.
Healthy Safe to keep shipping.

Two dashboards, one missing meeting

Keep LangSmith for debugging, ChartMogul for board NRR. However, neither emits a ranked ship / throttle / kill object on the join Account → Run → Outcome → Subscription.

Traces

Tools: LangSmith, Langfuse, Braintrust, OTel GenAI JSONL (prompts scrubbed).

Sees spans, tokens, evals, latency, retries.

Misses whether the paying account churned, and whether a successful run was a verified outcome.

CapEcon

Join Account, Run, Outcome, Subscription.

Record verdict, cost of leaving live, YAML rule, override.

Call ship, throttle, kill.

Revenue / CS

Tools: ChartMogul, ChurnZero, Stripe.

Sees MRR, NRR, health scores, cancel cohorts.

Misses CPSO, retry loops, HITL takeover, which capability to throttle.

Prefer OTel-compatible exports so ingest stays framework-agnostic.

Paid ChartMogul account with 92% LangSmith success and no verified outcome in 14 days.

Call: tourist / activation_failure

NRR is up while one capability spikes tokens.

Call: margin_leakage / run_cost_blowout, throttle

A Slack trace link plus a red health score, with no shared object.

Call: one GDR, the YAML cited, flywheel outcome 14 days later

Challenge Traces Revenue / CS CapEcon
Cost / margin Cost per span; no ARPU join MRR fine while serving cost explodes CPSO vs ARPU; throttle uneconomic capabilities
Visibility “What did this trace cost?” “What is NRR?” Unattributed spend; cost heatmap by step × cohort
Activation Run success ≠ first win Paying customer still in “trial” TTFV; paying-but-dormant; tourist GDRs
Switching costs Tool-call graph, not rip-out risk Churn reason unknown / competitor Connector blast radius
Trust / reliability Evals on traces Health score drops after the fact HITL trend; trust_break → rollback
Opaque success Completed trace ≠ trusted SOP Login/usage health Verified outcome rate; flywheel write-back
Synthetic teaching model The warehouse scaffolds the join, and economics use synthetic runs and seat priors today. When you have files, Data Connect overlays OTel GenAI JSONL, Langfuse exports, or CSVs into the same Workspace schema (metadata only: prompt bodies are scrubbed), with no live LangSmith or ChartMogul connector in this demo.
Common questions

Can’t I join this in Looker/dbt? You can SQL the join, but CapEcon’s IP is the exception taxonomy, YAML policy, and auditable decision record.

Is this observability? No: keep LangSmith or your OTel pipeline for traces, while CapEcon consumes trace-shaped facts and emits decisions.

Do I need OpenTelemetry? Prefer OTel GenAI attributes when you can (tokens, model, agent id), and Langfuse export also works; full nested OTLP protobuf JSON is not the path today.

Do you replace my stack? No: ingest later via Data Connect, and decide now on the teaching model.

Setup → The call → Price → Learn

The warehouse is synthetic by default, and Data Connect shows the path to real exports (OTel GenAI JSONL, Langfuse, CSV, Vision) without claiming live connectors.

Setup
  • Product Profile (presets switch ontology and priors)
  • Data Connect: OTel GenAI JSONL (prompts scrubbed), Langfuse, CSV, Vision
  • Outcome Definition (verified success, not login)
The call
  • Version Gate, Decision Inbox, Radar
  • Activation, Trust, Connectors, Subgraph, Version Compare
  • Marketplace Radar and Clinical Radar on matching presets
Price
  • Run Economics: CPSO, WTP caps, power-user margin leakage (defs)
  • Math Lab · Packaging (demand curve teaching)
Learn
  • Executive Summary, Experiments, Agentic Flags
  • Outcome Flywheel, Math Labs (binomial, power, CLV, drift)
Config
  • Integrations
  • Semantics Console, Taxonomy Browser, Record Inspector
  • Legacy analytics archive (reference only)

Machine-readable policy instead of hardcoded heuristics

Taxonomy → semantics → schema → record.
Change the verdict rules in YAML and the Radar reclassifies without a code deploy.

The ontology/ package defines exception categories, vertical-specific semantics.yaml action maps, and JSON Schema contracts for GrowthDecisionRecord. See ontology/ and docs/contracts.md.