In short
The U.S. Army rapidly exhausted its AI token allocation after pushing employees to use Ask Sage for generative AI work. The episode highlights the hidden costs and practical limits of “unlimited” AI access inside the Pentagon.
- The Army quickly used up its AI token pool after encouraging broad use of Ask Sage.
- Officials had to restore limits after saying access would be unlimited just weeks earlier.
- The case shows how token-based AI pricing can create sudden budget and usage pressure.
- The Pentagon continues to expand AI use even as concerns rise about reliability and oversight.
The U.S. Army has run through its AI token allowance far faster than expected, forcing it to reinstate limits on generative tools just weeks after the Defense Department touted widespread adoption. The episode underscores a growing problem for government and corporate AI rollouts: access may look “unlimited” on paper, but usage caps and budget pressure quickly turn enthusiasm into rationing.
In an internal email sent in mid-June to members of the Army’s Combat Capabilities Development Command, officials said the Army CIO token pool had already been depleted and that usage would have to be constrained again. The message came after the Army had told workers earlier in the spring that they could use the system without worrying about limits.
What happened to the Army’s AI token supply?
The Army exhausted its allocated token pool after encouraging broad use of a generative AI platform called Ask Sage, then reimposed limits once the pool ran dry. According to the email viewed by WIRED, the Army had announced unlimited token use in May 2026, but by mid-June the central token pool was already empty.
The service did not say whether the pool would be renewed after October 1, leaving employees uncertain about how much AI access they will have in the next budget cycle. The Army and the Department of Defense did not respond to requests for comment, and Ask Sage also declined to comment.
| Key fact | Details |
|---|---|
| Organization | U.S. Army, including DEVCOM |
| AI platform | Ask Sage, a multimodal generative AI workspace |
| Initial access | At least 200,000 tokens per month for some employees |
| Enterprise allotment | 100,000,000 tokens in an annual subscription package |
| Issue | Army CIO token pool was exhausted by mid-June |
| Open question | Whether the pool will be renewed after Oct. 1 |
Why the Army turned to Ask Sage
The Army has been using Ask Sage as an enterprise generative AI workspace designed to help personnel work with large language models in a controlled environment. The platform supports multiple models, including Google’s Gemini, Meta’s Llama and OpenAI’s ChatGPT, allowing users to switch between systems inside one interface.
Ask Sage is also accredited for handling Controlled Unclassified Information, a designation that makes it usable for a range of sensitive but unclassified government work. That accreditation helps explain why the Army and other Defense Department offices have adopted it for administrative and acquisition tasks rather than experimenting only with consumer-grade AI products.
What kinds of work was it used for?
According to the Army’s own materials, one of the tool’s uses has been reclassifying personnel descriptions by helping define job duties, experience requirements and background qualifications. That is the sort of repetitive bureaucratic task defense officials often point to when making the case that AI can save time and reduce administrative load.
The Army employee who spoke anonymously to WIRED said the broader push to use generative AI was not subtle. Workers were nudged to take advantage of the tools, and people who were signed up but not using their tokens regularly received reminders encouraging them to do so. In other words, the system was not simply available; it was being actively promoted.
How did the token system work?
The Army’s access model relied on tokens, the basic unit used to measure AI output or input in many large language model systems. For Ask Sage, documents reviewed by WIRED indicated that one token corresponds to about 3.7 characters.
Employees were initially allocated at least 200,000 tokens per month, according to emails seen by WIRED. If they used up their first allotment, additional tokens could be assigned automatically. Separate from that, the Army had access to an annual 100 million-token subscription as part of its enterprise pack.
That math helps explain how quickly usage can compound. Tokens disappear not just when people write lengthy prompts, but also when a model produces responses, drafts, summaries, translations, charts or images. A large number on a procurement sheet can therefore evaporate surprisingly fast once dozens or hundreds of users start experimenting.
Why do token limits matter?
Token limits matter because they are the clearest sign that AI use is never truly free. Even when agencies or companies market “unlimited” access, someone is still paying for compute, model calls and platform licensing. The Army’s experience shows that generous access policies often collide with real-world demand much sooner than managers expect.
The defense stakes are higher than in many private-sector settings. Federal agencies handle public records, sensitive operational information and mission-critical workflows. If a tool becomes popular before officials understand how much it is being used, the organization may face either runaway costs or abrupt constraints that frustrate employees and slow adoption.
“Apparently the whole Army burned through the whole year of tokens for just one service,” an Army employee told WIRED anonymously, describing the speed of consumption as surprising even inside the department.
How much AI is the Pentagon using?
The Army’s token crunch sits inside a much larger Department of Defense push to make generative AI a routine part of work. In May 2026, the Pentagon said nearly half of its 3.5 million employees were using AI on the job, a figure that reflected rapid internal enthusiasm for the technology across offices and commands.
That same push has created tension between speed and caution. On one hand, defense leaders want the efficiency gains that AI promises. On the other, critics worry that such systems may be unreliable, opaque or used too aggressively in decisions that should involve human judgment, especially when they touch personnel or operational assessments.
Recent reporting has also suggested that the Pentagon is continuing to integrate AI while trimming other parts of the bureaucracy. One example cited in the source material is the Civilian Protection Center of Excellence, which traditionally supported efforts to reduce civilian harm in conflict zones. According to that reporting, the department has reduced staffing there while developing AI tools to speed up some of the center’s work.
Who is asking for more AI use—and who is pushing back?
Inside the federal government, the pressure to expand AI use appears to be coming from leadership as well as enthusiastic users. The Army employee quoted by WIRED said workers were actively encouraged to use their tokens, suggesting that adoption was being treated as a performance and modernization issue, not merely an optional pilot program.
At the same time, some users are skeptical that generative AI is ready for the tasks being assigned to it. The same employee said the tools had not proven especially helpful for their work and that they had encountered reliability problems, including one model that claimed to have completed a task it had not actually finished.
The employee said they could see AI helping with some bureaucratic tasks in government, but warned that using the tools without judgment would not produce a trustworthy or efficient rollout.
That concern is not limited to the military. Similar stories are emerging across major technology companies, where early enthusiasm for AI has collided with cost control and productivity questions.
What other companies are running into the same problem?
The Army is not alone in discovering that AI usage can balloon quickly once employees are given broad access. Meta reportedly encouraged workers to “tokenmaxx,” then quietly removed its internal leaderboard and began trying to slow consumption. Adam Mosseri, who leads Instagram at Meta, has also raised the possibility of capping token use for engineers.
Uber is another example cited in the source material. According to reporting referenced there, some of the company’s engineers used a year’s worth of generative AI tokens in only four months. That kind of surge suggests that once a tool becomes convenient and integrated into workflows, usage can expand much faster than procurement teams anticipate.
The pattern is important because it illustrates a hidden reality of the generative AI economy: apparent abundance can mask hard limits. Models may be easy to access, but they are still constrained by computing costs, licensing agreements and budgeted quotas. When those limits are reached, the conversation shifts from innovation to management.
Why the Army’s experience matters beyond one service
The Army’s token shortage is about more than one subscription or one command. It offers a concrete case study in how government agencies adopt new AI systems under pressure to modernize quickly while still controlling costs, security risks and oversight obligations.
Defense agencies are especially likely to become heavy users because they have large bureaucracies, structured workflows and strong incentives to automate repetitive documentation. But they also have some of the toughest constraints, including security classification rules, public accountability and the need to avoid errors in areas tied to personnel, logistics and operations.
That combination makes the Army a useful test case for the broader public sector. If the military cannot sustain “unlimited” access without hitting a wall, smaller agencies may face the same challenge once AI usage spreads across offices and mission sets.
What are the risks of pushing AI too fast?
The biggest risk is that institutions confuse adoption with success. A tool can be used heavily and still produce low-quality results, especially if employees are nudged to use it even when it does not fit the task. In government, that can lead to wasted money, bad outputs, and misplaced confidence in automated assistance.
There is also the risk of uneven access. If one group of users depletes a shared pool, others may lose availability at critical moments. And if agencies are not clear about whether different classes of data draw from separate pools, employees may be left guessing about what kind of information is safe to put into a system and how costs are being allocated.
Timeline: how the Army’s AI token story unfolded
The sequence below shows how quickly the issue developed.
| Date | Event | Why it matters |
|---|---|---|
| May 2026 | Army CIO announced unlimited token use | Signaled broad confidence in generative AI adoption |
| Mid-June 2026 | Internal email said the token pool had been exhausted | Limits had to be reintroduced almost immediately |
| July 21, 2026 | WIRED published its report | Brought public attention to the Army’s rapid consumption |
| Oct. 1, 2026 | Renewal question remains unresolved | Agency users still do not know if access will continue |
What happens next?
The immediate question is whether the Army CIO pool will be renewed and, if so, under what terms. If the service restores access without changing the pricing or usage model, the same pattern could repeat. If it tightens limits, employees may be forced to ration AI use or abandon it for some tasks altogether.
More broadly, the episode is likely to intensify debate over how federal agencies should measure the real value of generative AI. Raw usage numbers do not tell leaders whether the tools are improving outcomes, saving time or producing more reliable work. They only show that people are trying them.
That distinction may matter most inside the Defense Department, where the push for automation is happening alongside high-stakes decisions and a still-developing understanding of what these tools can and cannot do. The Army’s token crunch suggests that the next phase of AI adoption will be defined less by hype than by accounting, oversight and hard trade-offs.
For now, the message to Army workers is clear: the AI experiment is continuing, but it is no longer operating under the assumption that usage can grow forever without consequences.
Frequently asked questions
Why did the Army run out of AI tokens?
The Army ran out of AI tokens because employees used Ask Sage heavily after the service encouraged broad adoption. Internal emails said the central token pool was exhausted by mid-June, forcing the Army to reimpose limits after previously saying access would be unlimited.
What is Ask Sage?
Ask Sage is a multimodal generative AI platform used by the Army and other Defense Department offices. It lets users access multiple large language models in one workspace and is accredited for Controlled Unclassified Information, making it suitable for some sensitive government tasks.
How many tokens did Army employees get?
Army employees were given at least 200,000 tokens per month in some cases, according to emails reviewed by WIRED. The Army also had access to a 100 million-token annual enterprise subscription, but usage apparently moved fast enough to drain the shared pool.
Did the Army say whether the token pool will be renewed?
The Army did not say for certain. The internal email noted that the current usage level would continue, but it was unclear whether the Army CIO pool would be renewed after Oct. 1, leaving future access unresolved.
Are other companies limiting generative AI use too?
Yes. Meta has reportedly moved to curb internal token use after encouraging heavy adoption, and reporting has said Uber engineers used a year’s worth of generative AI tokens in only four months. The Army’s experience fits a broader industry pattern.









