How companies can cut AI token costs without reducing staff

AI token spending is becoming a major operating cost for companies at scale Cut costs with smarter AI engineering, not layoffs, and protect performance

Nvidia chief executive Jensen Huang said engineering teams should be measured partly by how much AI token usage they generate, arguing that token spending is now a meaningful operating cost for companies using AI at scale. The article says large technology firms are spending heavily on AI infrastructure while also using AIrelated cuts to fund investment. The piece argues that reducing staff has not reliably delivered better returns. It cites research from Gartner showing that many companies that cut headcount after adopting AI or automation did not see a clear improvement in performance. It also points to examples such as Uber, where heavy use of AI coding tools quickly exhausted the company’s AI budget, and Meta, where layoffs were presented as part of broader cost reallocation. Instead of focusing on layoffs, the article says companies can lower token expenses through engineering changes such as prompt caching, model routing, batch processing, retrievalaugmented generation, prompt compression, and selective use of openweight models. It also notes that some firms are using AI to support workers rather than replace them, citing Klarna’s shift back toward a blended model after customer service quality declined. The article concludes that businesses may gain more from optimizing AI usage and preserving workforce experience than from cutting jobs to cover AI spending.