Tokenomics — subscription fuel gauges per provider (source class, what a seat may decide)
Tokenomics — subscription fuel gauges per provider
Who reads this: any seat about to reason about which model subscription can take work (burn order, tier routing, cost questions), and anyone about to claim “we can see provider X’s quota” — the three providers give DIFFERENT guarantee strengths, and writing the difference down is the point of this page.
Standing rule (owner 2026-09-15, canonised from T-S00-150926-BURN-GLM): burn the quota
that EXPIRES SOONEST, not the one you prefer — an unused quota that resets is money already
spent. The computed form lives in scripts/bridge/tier-policy.py
(_burn_soonest_order) reading metrics/tier-fuel-gauge.json; the freshness gate refuses to
order from a gauge older than its own shortest window or describing an already-reset window.
Per-provider quota surfaces — the source class is the claim
z.ai — PROVIDER’s own API (strongest guarantee)
- Surface:
GET https://api.z.ai/api/monitor/usage/quota/limit, auth token in the secrets env without theBearerprefix (measured: with Bearer it fails). Written byscripts/health/tier-fuel-gauge.py; refreshed on a 15-min systemd user timer (neurport-fuel-gauge.timer, neurport-oracle). - Windows: THREE — 5h, weekly, monthly — each with
percentage,nextResetTime. - What a seat may decide: real fuel arithmetic — time-to-exhaustion, burn-soonest order, the daily owner line. The provider itself answers; the number is as good as the provider’s own accounting.
- Dated reading 2026-09-16T17:0xZ: 5h window 24%, weekly 42%, monthly 12%.
minimax — KILO’s endpoint, NOT the provider’s API (weaker guarantee)
- Surface:
GET http://127.0.0.1:<kilo-serve>/kilocode/provider-usageon the machine where the Kilo serve runs (loopback). This is Kilo’s view of minimax, not minimax’s own accounting. - “We can see minimax quota” and “Kilo can see minimax quota” are DIFFERENT claims. What Kilo reports can lag the provider (its own fetch cadence, its own error handling), and a Kilo-side outage reads as “minimax quota unknown” even when minimax is fine. Cite the layer when the number matters: say “per Kilo”, never “per minimax”.
- What a seat may decide: burn-order participation and daily-line values, with the provenance caveat attached. Do NOT use it as grounds for money decisions (top-ups) without the owner’s own dashboard reading — that is a Rule-23 surface.
- Dated reading 2026-09-16T17:0xZ: 5h window 0%, 1w 28% (per Kilo).
xAI — NO quota surface anywhere (provider limitation, not our gap)
- Surface: none. xAI exposes no quota API, and xai is absent from Kilo’s
/kilocode/provider-usageresponse (proven 2026-09-15, S0 live probe). The owner’s console dashboard is the ONLY measurement that exists. - What a seat may decide: NOTHING automatic. xAI rows stay
state: absent— printed as UNMEASURED, never as zero (a false zero would silently rank xai in the burn order). Any spend/burn statement about xAI needs the owner reading the dashboard and handing the number over. - Writing this down stops it being rediscovered — three separate probes (S0 2026-09-14/15, S8 2026-09-15) each cost a turn before this page existed.
The meter and its consumers
| artefact | role |
|---|---|
scripts/health/tier-fuel-gauge.py |
reads the surfaces above; three states (active / exhausted / absent); --selftest covers all three |
metrics/tier-fuel-gauge.json |
the snapshot consumers read; authored every 15 min by neurport-fuel-gauge.timer |
scripts/bridge/tier-policy.py _burn_soonest_order + freshness gate |
orders cheap-class candidates by min days_to_reset; GAUGE-FRESH/GAUGE-STALE lines name the age; stale → falls back to the declared pin config/am-live-model.txt |
config/tier-policy.yaml |
supersession record: burn-by-expiry (2026-09-15) supersedes the three 2026-08-13 minimax-first preference minutes |
config/subscription-time-model.yaml |
owner-supplied prices (the ONLY $ figures that exist; plans are not denominated $/token) |
Records (measurement history, not canon)
docs/reports/S8-BURN-SOONEST-FUEL-GAUGE-20260915.md— the burn-soonest ruling, the probe that proved glm-5.3-flash alive, the Claude-Code counter-example (80%/4d → leave alone).docs/reports/RES-TIER-ROUTING-INDUSTRY-20260915.md— industry survey: our “exhausted” is the standard’s COOLDOWN; the weekly-expiry burn order is NOT in the standard (rpm/tpm protect against overload, not waste); our delta is live usage pacing.docs/reports/RES-R2-COST-PER-RESULT-20260916.md— $12.83 per completed product-class task (7d, amortized cash; 7 undercount flags; per-token pricing is NOT computable).docs/reports/S8-QUOTA-GAUGE-STALE-FRESHNESS-20260916.md— the freshness gate (age vs the gauge’s own shortest window + dead-window branch), the schedule, the host-clock offset trap.
Known open gap (2026-09-16)
_burn_soonest_order matches gauge providers (zai, minimax, xai) against subscription ids
(glm-flash, minimax, grok-build); the zai↔glm-flash mapping is not wired, so zai rows (the
soonest-resetting quota most days) do not yet reorder glm-flash. Ready unit:
S8-BURNSOONEST-SUB-PROVIDER-MAP in docs/briefs/TRACK-BACKLOG-S8-tokenomics.yaml.