White Paper
AI Unit Economics: Why Token Factories Need More Than Metering
This white paper argues that token metering explains customer billing but not whether an AI service is profitable. GPU rental margins can be only 14–16%, while utilization may sit between 13% and 33%, leaving expensive capacity idle. Tokens also hide major cost differences caused by model size, input versus output generation, context length, batching, KV-cache use, and resource occupancy. A true economics dashboard should combine GPU-seconds, memory, power, MIG allocation, utilization, cache hits, latency, throughput, queue depth, tokens, tool calls, tenant spend, and cost per successful outcome. The paper describes four billing generations: GPU-hour, token, outcome, and policy-driven economics. PaletteAI adds multitenancy, quotas, MIG, dynamic allocation, scheduling, utilization and cost
