How One Startup Is Bringing Wall Street–Style Contracts for AI Tokens

The Information | 12-08-2026 06:06pm |

If the number of times I called someone this week and heard their cellphone make the familiar ring of a European phone is any indication, many of you are on vacation. Still, there are few breaks these days in the world of financing AI. With trillions of dollars due to pour into chips and data centers and companies spending heavily on their own use of the technology, nearly everyone involved is looking for ways to make their spending and returns more predictable. That’s already spurred the creation of marketplaces and contracts for graphics processing units that let buyers and sellers lock in compute rental pricing. And on Tuesday, CME Group announced an October launch date and details for two exchange-traded GPU futures contracts, pending regulatory review.  Now those efforts are turning to tokens, the units that companies actually pay for when they use AI models. Compute Exchange, a startup that matches chip owners with short-term renters, is moving into a similar matchmaking role for tokens, The Information has learned. It will offer token forwards, privately negotiated contracts that give enterprises a way to lock in their token prices over a timeframe of as long as six months, hedging against future swings in AI costs. The token forwards offer a different pricing model than GPU contracts, which typically lock in the hourly cost of computing capacity used to run AI models. Token contracts would instead provide consistent pricing over a much longer time period, and for an item that enterprises actually use in their day-to-day.  Customers can select from a variety of mostly open-weight models offered by one of the six inference providers that have signed up with Compute Exchange. Forwards aren’t quite the same as standardized, exchange-traded futures like the ones CME is planning, but they mark another step in bringing financial market-style hedging to more kinds of AI costs. Tokens are the snippets of words or phrases that get sent to AI models and come back as its response in a process called inference. It is this action rather than the training of new models that accounts for the majority of the world’s interaction with AI. Inference is set to explode as agents begin to handle more and more of the world’s AI traffic. Deloitte estimates inference will account for about two-thirds of AI workloads by the end of 2026, up from one-third in 2023.The early days of AI experimentation saw companies offer employees unfettered access to the technology. Meta Platforms and others defined the era of tokenmaxxing, setting up internal competitive league tables that encouraged employees to experiment with AI early and often, consuming massive quantities of tokens. Then companies including Uber and ServiceNow disclosed that they had blown through their annual AI budgets in a matter of months, kicking off new restrictions on employee spending. Any opportunity to lock in the price of tokens should help companies, startups and others get a better handle on their spending, according to Val Bercovici, chief AI officer for WEKA, an AI data and memory infrastructure company. WEKA helps AI systems use GPU memory more efficiently. “Everyone is scrambling to understand their cloud bills,” Bercovici said. “Tokenmaxxing quickly turned from a badge of honor [to be] at the top of the leaderboard to a badge of shame.”Recent developments including larger and more complex models, larger context windows and agents that can run for days and weeks on their own, as well as emerging cybersecurity threats, have caused an explosion in token consumption, he said. That makes it difficult for companies to get a handle on everything they need to know to make informed choices about how to best manage costs.“This market is far too complex and too opaque, and you need clearinghouses to add transparency,” he said. In parallel with launching token forwards, Compute Exchange is creating something called standardized token units, a methodology that benchmarks pricing proposals and model performance. According to Carmen Li, CEO of Compute Exchange, the company’s pitch is that token forwards offer an alternative to the predominant method of consuming AI today, which is accessing models through broad-based application programming interfaces. Relying on APIs can make it difficult for users to manage costs or measure performance because the model companies charge based on token usage, while also retaining the ability to change how quickly tokens are processed—for instance, they can change a user’s place in the line or otherwise throttle usage without disclosing it, she said. Users can be relegated to paying on-demand or spot prices rather than a locked-in price, which Bercovici described as the reserve price. Bercovici said spot prices for tokens have become increasingly volatile in recent months, something that’s apparent on OpenRouter, where prices for new models used to be stable over hours, days and even weeks. Now those prices can change hourly, he said. 

Stay Updated with the Latest News!

Don't miss out on breaking stories and in-depth articles.