If AI Can’t Count Its Own Words, Should We Trust It With Our Numbers?
Anyone who regularly uses artificial intelligence for writing eventually encounters a strangely revealing problem. Ask a sophisticated AI model to produce an article of exactly 750 words, and it may confidently deliver 683. Ask for 5,000 words, and it might produce 5,700 while insisting it followed the instruction. Ask it to count the words in what it just wrote, and the answer itself may be wrong.
This seems almost comical until we consider what else we are preparing to entrust to artificial intelligence.
If a machine struggles to count the words in its own essay, why should Americans trust it with financial projections, engineering calculations, medical statistics, economic forecasts, military logistics or trillion-dollar government budgets?
The answer requires understanding an important distinction that is too often lost in the excitement surrounding AI. Large language models such as ChatGPT are not, fundamentally, calculators. They are language-prediction systems. They process text through units called tokens and generate responses by predicting what should come next based on enormously complex mathematical relationships learned during training.
That architecture can produce extraordinary results. An AI can explain quantum mechanics, analyze a contract, summarize legislation, write computer code and discuss Shakespeare. Yet asking the same system to count every word in a long passage can expose an almost embarrassing weakness.
The contradiction is only apparent. A calculator and a language model are designed to do fundamentally different things.
A pocket calculator does not “understand” mathematics in the way we ordinarily use that word, but ask it to multiply 7,413 by 928 and it performs a deterministic operation. A language model, operating by itself, can instead approach arithmetic as another language problem. It may generate the number most likely to be correct rather than mechanically calculate the answer.
That distinction should shape how America deploys AI.
The danger is not that artificial intelligence cannot perform large-scale calculations. AI systems can be connected to calculators, databases, statistical software, computer code and specialized mathematical engines that perform precise operations extraordinarily well. The danger comes when users assume that because an AI sounds intelligent, every number appearing in its answer must have been calculated rather than generated.
Confidence is not verification.
That lesson becomes increasingly important as corporations and governments race to automate decision-making. Imagine an AI system analyzing a pension fund, estimating the cost of a federal program or calculating structural requirements for a bridge. Being “approximately correct” is not necessarily acceptable. A one-percent error in a household budget may be inconvenient. A one-percent error applied to hundreds of billions of dollars can represent billions.
The appropriate response is not to abandon AI. It is to stop pretending AI is magic.
Conservatives should find something familiar in that conclusion. Institutions work best when power is divided, claims are independently verified and no single authority is assumed to be infallible. The same philosophy should govern artificial intelligence.
An AI model should generate analysis. A deterministic mathematical system should perform calculations. Software should verify the output. Human beings should review consequential decisions. Where enormous sums of money, public safety or individual rights are involved, audit trails should show exactly where important numbers originated and how they were calculated.
In other words, trust should be replaced by verification.
There is also an important difference between AI making a calculation and AI managing a calculation. The latter may ultimately prove far more revolutionary. A sufficiently capable model does not need to perform millions of arithmetic operations internally. It can formulate the problem, select appropriate tools, write code, instruct specialized software to execute the mathematics, examine the results and present them in understandable language.
That resembles how competent humans already work.
An engineer does not prove his competence by refusing to use a calculator. An accountant does not demonstrate intelligence by calculating an entire corporate balance sheet mentally. Professionals use specialized tools because reliability matters more than theatrical demonstrations of mental arithmetic.
AI should be judged similarly.
The real concern begins when nobody knows whether the machine used a reliable tool or merely generated a plausible-looking answer. As AI becomes embedded in banking, government, medicine, defense and infrastructure, that distinction cannot remain hidden behind a friendly conversational interface.
Organizations using AI for consequential numerical work should therefore be required—or compelled by liability and common sense—to distinguish between generated estimates and verified calculations. Important results should be reproducible. Assumptions should be visible. Independent systems should be capable of checking the arithmetic.
And humans must remain accountable.
One of the most dangerous phrases of the coming decade may be, “The AI calculated it.” That statement could become a convenient way for bureaucracies and corporations to evade responsibility for mistakes nobody bothered to verify.
The humble word-count problem therefore teaches us something important.
Artificial intelligence can appear brilliant while failing at something elementary. That does not make AI useless or inherently untrustworthy. It means intelligence and reliability are different qualities.
We should use AI enthusiastically where it expands human capability. But when the numbers matter, Americans should insist upon something older than artificial intelligence and considerably less fashionable:
Show your work.Anyone who regularly uses artificial intelligence for writing eventually encounters a strangely revealing problem. Ask a sophisticated AI model to produce an article of exactly 750 words, and it may confidently deliver 683. Ask for 5,000 words, and it might produce 5,700 while insisting it followed the instruction. Ask it to count the words in what it just wrote, and the answer itself may be wrong.
This seems almost comical until we consider what else we are preparing to entrust to artificial intelligence.
If a machine struggles to count the words in its own essay, why should Americans trust it with financial projections, engineering calculations, medical statistics, economic forecasts, military logistics or trillion-dollar government budgets?
The answer requires understanding an important distinction that is too often lost in the excitement surrounding AI. Large language models such as ChatGPT are not, fundamentally, calculators. They are language-prediction systems. They process text through units called tokens and generate responses by predicting what should come next based on enormously complex mathematical relationships learned during training.
That architecture can produce extraordinary results. An AI can explain quantum mechanics, analyze a contract, summarize legislation, write computer code and discuss Shakespeare. Yet asking the same system to count every word in a long passage can expose an almost embarrassing weakness.
The contradiction is only apparent. A calculator and a language model are designed to do fundamentally different things.
A pocket calculator does not “understand” mathematics in the way we ordinarily use that word, but ask it to multiply 7,413 by 928 and it performs a deterministic operation. A language model, operating by itself, can instead approach arithmetic as another language problem. It may generate the number most likely to be correct rather than mechanically calculate the answer.
That distinction should shape how America deploys AI.
The danger is not that artificial intelligence cannot perform large-scale calculations. AI systems can be connected to calculators, databases, statistical software, computer code and specialized mathematical engines that perform precise operations extraordinarily well. The danger comes when users assume that because an AI sounds intelligent, every number appearing in its answer must have been calculated rather than generated.
Confidence is not verification.
That lesson becomes increasingly important as corporations and governments race to automate decision-making. Imagine an AI system analyzing a pension fund, estimating the cost of a federal program or calculating structural requirements for a bridge. Being “approximately correct” is not necessarily acceptable. A one-percent error in a household budget may be inconvenient. A one-percent error applied to hundreds of billions of dollars can represent billions.
The appropriate response is not to abandon AI. It is to stop pretending AI is magic.
Conservatives should find something familiar in that conclusion. Institutions work best when power is divided, claims are independently verified and no single authority is assumed to be infallible. The same philosophy should govern artificial intelligence.
An AI model should generate analysis. A deterministic mathematical system should perform calculations. Software should verify the output. Human beings should review consequential decisions. Where enormous sums of money, public safety or individual rights are involved, audit trails should show exactly where important numbers originated and how they were calculated.
In other words, trust should be replaced by verification.
There is also an important difference between AI making a calculation and AI managing a calculation. The latter may ultimately prove far more revolutionary. A sufficiently capable model does not need to perform millions of arithmetic operations internally. It can formulate the problem, select appropriate tools, write code, instruct specialized software to execute the mathematics, examine the results and present them in understandable language.
That resembles how competent humans already work.
An engineer does not prove his competence by refusing to use a calculator. An accountant does not demonstrate intelligence by calculating an entire corporate balance sheet mentally. Professionals use specialized tools because reliability matters more than theatrical demonstrations of mental arithmetic.
AI should be judged similarly.
The real concern begins when nobody knows whether the machine used a reliable tool or merely generated a plausible-looking answer. As AI becomes embedded in banking, government, medicine, defense and infrastructure, that distinction cannot remain hidden behind a friendly conversational interface.
Organizations using AI for consequential numerical work should therefore be required—or compelled by liability and common sense—to distinguish between generated estimates and verified calculations. Important results should be reproducible. Assumptions should be visible. Independent systems should be capable of checking the arithmetic.
And humans must remain accountable.
One of the most dangerous phrases of the coming decade may be, “The AI calculated it.” That statement could become a convenient way for bureaucracies and corporations to evade responsibility for mistakes nobody bothered to verify.
The humble word-count problem therefore teaches us something important.
Artificial intelligence can appear brilliant while failing at something elementary. That does not make AI useless or inherently untrustworthy. It means intelligence and reliability are different qualities.
We should use AI enthusiastically where it expands human capability. But when the numbers matter, Americans should insist upon something older than artificial intelligence and considerably less fashionable:
Show your work.

