AT&T is curbing its Anthropic and OpenAI spending by shifting employee AI queries to cheaper open-source models and smart routers, aiming to keep costs flat amid surging usage.
The telecom giant processes about 45 billion AI tokens daily, up from roughly 8 billion earlier. Open-source and open-weight models now handle 40% of employee queries. AT&T plans to raise that share to 60–70% in coming years. It relies on Nvidia’s Nemotron, Meta’s Llama, Google’s Gemma, and its own telecom-tuned OTel models based on Gemma. Chinese options like DeepSeek remain under review but unused for now.
For coding and advanced tasks, LiteLLM routers assess complexity and direct simpler work to lower-cost models. This cut coding AI costs by up to 56%, with only a 2% drop in quality. Open models lag frontier systems by 6–10 months but already match or exceed older proprietary versions for many jobs.
The strategy keeps Anthropic and OpenAI bills steady while usage grows. AT&T has long used cache-aware gateways that route tasks for 80–90% savings in some cases, calling it a defense against the “token apocalypse.”
The move mirrors a broader enterprise trend: firms are imposing AI spending controls and favoring open alternatives as bills climb, pressuring OpenAI and Anthropic to offer better cost tools and pricing.
See What’s Next in Tech With the Fast Forward Newsletter
Tweets From @varindiamag
Nothing to see here - yet
When they Tweet, their Tweets will show up here.




