Spotify's method for Claude Code cuts token usage by 90% but raises costs in some cases
A Spotify engineer adapted the company's approach to reduce Claude Code's token consumption by 90%, but the strategy led to higher expenses in certain scenarios. The method involves routing large files to cheaper models, though it may introduce latency or complexity.

Spotify's engineering team implemented a system to minimize token usage in Claude Code by adapting strategies originally used for internal tools. The approach leverages a shunt mechanism that redirects large file processing to cheaper models, reducing token consumption by up to 90% in some cases. This method was inspired by Spotify's own practices for managing AI workloads efficiently.
The system works by detecting when Claude Code attempts to process files exceeding 350 lines. At that point, a hook intervenes and reroutes the task to a bulk-reader skill. This skill sends the file to a less expensive model, which generates a summary and returns it to Claude. This process significantly reduces the number of tokens consumed during large-scale file analysis.
According to internal reports, the shunt mechanism achieved a mean token reduction of 90% when tested against a 162,000-line Java monorepo. The savings ranged from 82% to 94% across different scenarios, demonstrating the effectiveness of the approach. However, the system's reliance on additional models and routing logic can complicate implementation and increase infrastructure costs in some cases.
The strategy introduces potential trade-offs, including increased latency and the need for additional infrastructure to support the routing logic. While the token savings are substantial, the complexity of managing multiple models and ensuring seamless integration can lead to higher operational costs. These factors must be weighed against the benefits of reduced token consumption, particularly in environments with strict budget constraints.
The approach highlights the ongoing challenge of balancing cost efficiency with system complexity in AI workflows. While Spotify's method successfully reduces token usage, it underscores the need for careful evaluation of trade-offs such as infrastructure costs and latency. As AI tooling evolves, similar strategies may become more common, requiring organizations to adapt their workflows accordingly.