When you measure AI usage, you get usage
Duolingo, Amazon and Meta all tried to measure AI adoption in 2026, and all three backed off. What they learned is the oldest lesson in product metrics: measure what the usage was supposed to change.
Three companies tried to measure AI adoption this year. All three changed their mind.
None of them are small, and none of them are skeptical about AI. That is what makes the pattern worth a closer look.
Three companies, one reversal
Duolingo, April 2026. A year earlier, CEO Luis von Ahn had announced that the company would be AI-first and that AI use would count in performance reviews. In April he told Fortune that they backtracked. His reasoning, in his words: "the most important thing in your performance is that you are doing whatever your job is as well as possible."
Amazon, May 2026. A group of Amazon employees had built KiroRank, an internal leaderboard of AI token usage. Amazon shut it down, and a senior vice president's message to staff was blunt, as Business Insider reported: "Please don't use AI just for the sake of using AI. Use AI to help you solve customer problems, to help you solve business problems, to innovate."
Meta, September 2026. Meta removed AI usage from engineers' performance reviews after people burned through tokens to climb internal leaderboards, a habit that got its own name: tokenmaxxing. According to The Decoder, reviews now look at the quality, speed and complexity of the work instead.
Three different companies, three different mechanisms: a review criterion, a leaderboard, a token count. Same ending.
When you measure usage, you get usage
None of this is surprising. Put a number in front of people, tie it to their review or their rank, and they will move the number. That is not cheating. That is doing exactly what the system asked for.
The problem is what the system asked for. Token counts, prompts per day and the share of people using a tool are all output metrics. They tell you that something happened. They tell you nothing about whether it helped.
I see the same mistake in product teams that have nothing to do with AI. A team gets measured on features shipped, and it ships features. Whether customers changed their behavior is somebody else's question. AI did not create this problem. It made it cheaper to produce output, so the gap between output and impact got wider, faster. Marty Cagan calls the bigger version of this the AI productivity paradox, and I wrote up his talk in faster is not better.
Measure what the usage was supposed to change
Here is the question I would ask instead. Which product metric moved, and why?
Not "how many engineers use the assistant", but "did the thing we wanted the assistant for get better". If AI was supposed to help the team learn faster, look at how quickly ideas get tested and dropped. If it was supposed to improve the product, look at the product outcome the team owns, the change in user behavior that leads to the business result. I explained that difference in product outcomes vs. business outcomes, and it applies to AI tools exactly as it applies to features.
The "and why" matters as much as the "which". A number that moved without an explanation is a coincidence you have not checked yet. That is the same reason I argue for being data-informed, not data-driven: the metric starts the conversation, it does not end it.
Basically, what is the impact? If nobody can answer that, more usage will not fix it.
What does your company measure?
If your leadership team has an AI adoption dashboard, look at what is on it. If every number on it counts activity, you are one leaderboard away from your own version of tokenmaxxing.
Getting AI to change how teams decide, not just how much they produce, is what the AI-First workshop and sprint is built around. It starts from the outcomes a team owns and works backwards to where AI actually helps.