The Proxy Economy
A number should not become a victory until the people carrying the cost can see the bridge between them
AI Has a Proof Problem
The artificial intelligence industry does not lack measurements. It has token prices, benchmark scores, paid seats, pull requests, capital expenditures, adoption curves, latency figures, and market valuations, all arriving quickly enough to give every decision the appearance of evidence.
The problem begins after the measurement. A lower price becomes proof that a workload is ready for production, access becomes proof that an institution has gained scientific capacity, more pull requests become proof that engineering is more productive, paid seats become proof that work has been transformed, manufacturing origin becomes proof of security, and technical foresight becomes proof of investment judgement.
Each substitution feels reasonable because the proxy is related to the claim. Price affects viability, access enables work, output contributes to productivity, supply chains affect security, and expertise sometimes transfers between adjacent domains, but related evidence is not equivalent evidence.
The distance between the measured fact and the larger conclusion is where institutional failure hides. That distance is rarely crossed through validation because it is crossed through narrative, repetition, and the quiet assumption that somebody else must have proved the conversion.
The proxy economy exists because proxies let institutions recognize success before the outcome arrives. The vendor gets revenue, the executive gets adoption, the regulator gets a rule, the program gets participation, and the founder gets capital while someone farther down the chain inherits the missing proof.
The operator inherits remediation, the engineer inherits review, the researcher inherits reproducibility, the customer inherits outcome risk, the consumer inherits insecure behavior, and the investor inherits leverage. That transfer is not an incidental weakness in the measurement model because it is what makes the measurement useful to the people selecting it.
A proxy becomes dangerous when the actor who benefits from the conclusion also gets to decide what the proxy proves. The fraud is usually not in the number itself but in the uninspected distance between the number and the victory declared in its name.
Price, Adoption, and Premature Credit
OpenAI’s July 30 article Advancing the Price-Performance Frontier with GPT-5.6 announced an 80 percent price reduction for GPT-5.6 Luna, a 20 percent reduction for Terra, and a Sol Fast mode delivering responses up to 2.5 times faster at twice the standard price. OpenAI argued that those changes would make a broader range of tool-using and multi-step applications practical to operate at scale.
That is a material economic change, and pretending otherwise would be its own kind of theater. Workloads that were marginal can cross into viability when the cost of model execution falls by that much, especially classification, document processing, routine implementation, background automation, and agent loops.
The trouble begins when token price is allowed to stand in for operating cost. Integration, evaluation, observability, exception handling, security review, human verification, incident response, data preparation, workflow redesign, and failed automation do not disappear because the model call got cheaper.
The model call may become inexpensive while the complete result remains costly. That difference matters more as systems move from answering questions to changing records, communicating with customers, generating production code, initiating transactions, or triggering other automated processes.
Lower inference cost increases the number of actions an organization can afford to attempt. It does not establish that the organization can afford the error rate, supervision burden, accumulated ambiguity, or repair work those actions create.
The economic claim in Advancing the Price-Performance Frontier with GPT-5.6 is therefore narrower than the adoption narrative organizations may build around it. OpenAI changed the price of model execution, while the adopter still has to prove that the complete outcome is cheaper after validation, remediation, support, and failure are included.
When those costs are omitted, the savings have not vanished. They have been transferred to the people who maintain the process after the procurement case has already declared victory.
Microsoft’s July 29 release Microsoft Cloud and AI Strength Fuels Fourth Quarter Results shows the other side of this conversion. Microsoft reported $90 billion in quarterly revenue, $59.3 billion in Microsoft Cloud revenue, 43 percent Azure growth, and more than 30 million paid Microsoft 365 Copilot seats.
Those figures prove commercial demand at extraordinary scale. Customers are purchasing capacity, licenses, and access, and Microsoft has every right to report that success through the measurements it owns.
They do not prove that each customer transformed its work. A paid seat proves procurement, an active user proves interaction, a prompt proves consumption, and revenue proves that Microsoft captured value, but none of those measures proves that the purchasing organization improved the quality, speed, cost, or resilience of its work.
The conversion failure occurs inside the customer. An executive buys the seats, the rollout team reports adoption, managers are asked to increase usage, and employees discover that the old process still exists except it now includes another interface and a recurring request to explain why the utilization dashboard is not greener.
The vendor has completed its proof because it sold the product. The customer still owes a different proof about what changed after the invoice was paid.
That proof belongs in the work itself. Sales cycles should shorten without increasing error or discounting, engineers should deliver reliable changes without moving effort into review and remediation, support teams should resolve more problems at first contact, and managers should recover time for judgement rather than merely produce more documents.
Those outcomes are harder to count than seats, which is precisely why institutions prefer the seat count. Deployment can be announced immediately, while transformation can remain an aspiration long enough for the next budget cycle to inherit it.
Enterprise software had shelfware long before generative AI. Apparently, civilization was unwilling to leave such a durable form of waste unexplored.
Where the Work Went
Justin Reock’s article AI Productivity Gains Are 10 Percent, Not 10x provides the evidentiary object behind this week’s engineering-productivity argument. The article reports preliminary findings from DX’s longitudinal study of more than 400 engineering organizations between November 2024 and February 2026.
AI tool usage increased by an average of 65 percent while median pull-request throughput increased by 7.76 percent, with most organizations landing between 5 and 15 percent rather than the 2x, 3x, and 10x gains circulating through vendor claims and executive expectations. The article was first published in March and later updated, while LeadDev resurfaced the argument in its July 30 newsletter.
The lazy conclusion is that AI coding tools failed to deliver. The more useful conclusion is that the industry selected a convenient unit of output and treated it as a complete description of engineering.
Code generation is one activity inside a delivery system. Software still has to be understood, reviewed, tested, integrated, secured, deployed, observed, supported, and changed again by people who may not have generated it.
The work did not disappear when generation accelerated. It moved into the parts of the system that remained constrained, which is why a developer can finish a change faster while a reviewer spends longer reconstructing its reasoning and a team can merge more pull requests while making the codebase harder to own.
Reock’s AI Productivity Gains Are 10 Percent, Not 10x is careful about that distinction because it reports pull-request throughput rather than total business value. Organizations become less careful when they promote the throughput figure into productivity and then promote productivity into proof that the investment transformed delivery.
Leadership buys tools against a 10x expectation and observes a gain closer to 10 percent. Instead of questioning the original measurement model, the organization blames adoption, prompt quality, engineering resistance, or the absence of yet another tool.
The proxy protects the promise by blaming the system it failed to describe. Ten percent can be extraordinarily valuable across a large engineering organization, and it does not need mythology to justify itself.
The operational question is where the reclaimed time went. If it became better design, deeper testing, faster incident recovery, lower cognitive load, or shorter lead time from validated need to reliable production use, the organization gained something worth defending.
If the reclaimed time became more generated output waiting for inspection, the improvement was borrowed from downstream capacity. The debt will be collected through review, remediation, debugging, and systems that become harder to explain each time another generated layer is added.
OpenAI’s July 22 article Advancing the Next Era of National Science exposes the same displacement in a domain where the word productivity would be too crude. OpenAI committed $4 million in Codex access for approximately 2,000 researchers participating in the Department of Energy’s Genesis Mission, $3 million in API support for major scientific campaigns, and additional access to specialized capabilities.
The article did not claim that access alone would produce discovery. It tied frontier models to research tools, workflows, domain expertise, supercomputers, simulations, facilities, and scientific standards, while keeping researchers responsible for defining questions, selecting methods, challenging outputs, and validating results.
That qualification is sound because frontier access can expand the supply of plausible work faster than an institution can validate it. A model can generate hypotheses faster than a laboratory can test them, compress literature faster than a researcher can reconstruct the source trail, and produce experimental designs faster than physical infrastructure can determine whether they describe the world.
The bottleneck moves toward evidence rather than disappearing. The institution can still report accounts created, credits distributed, campaigns launched, and researchers enrolled before the first scientific claim survives replication.
The program sponsor receives early credit while the researcher inherits the longer proof. The real measures arrive later through weak hypotheses eliminated sooner, experiments made more discriminating, validated results reached faster, and methods that another team can reproduce using preserved inputs, model versions, prompts, sources, and intermediate transformations.
Without those measures, Advancing the Next Era of National Science proves that OpenAI committed meaningful access and resources. It does not yet prove that scientific capacity increased because that conclusion belongs to the results produced through the program rather than the size of the program itself.
Classification, Expertise, and Transfer
The FCC’s July 28 document FCC Updates Covered List to Include Foreign-Produced Advanced Robotic Devices and Power Inverters announced that new covered models would generally be denied the equipment authorization required for importation, marketing, or sale in the United States. Previously authorized models and previously purchased devices remain unaffected.
The fact sheet says the FCC acted after national-security determinations by an executive-branch interagency body. It identifies supply-chain vulnerability, cybersecurity exposure, surveillance, manipulation of data and physical operation, and remote commandeering among the risks associated with networked robotic systems.
Those risks are real, and the policy may reduce a legitimate category of geopolitical and supply-chain exposure. Manufacturing origin can affect legal obligations, component provenance, update authority, vendor control, and the relationship between a supplier and a foreign government.
Origin is therefore relevant evidence, but it is not a complete security assessment. A domestically produced robot can still use weak authentication, transmit excessive telemetry, retain maps indefinitely, depend on an opaque cloud service, accept insecure updates, or become useless when the vendor abandons it.
A foreign-produced device can implement stronger technical protections while carrying a separate national-security exposure. Those facts can coexist because origin and behavior are not the same property.
The FCC fact sheet answers which new devices may enter the market under the determination the Commission was required to implement. It does not answer how every admitted device authenticates updates, limits telemetry, protects stored data, survives cloud failure, exposes remote access, or permits independent inspection.
Policy receives a clean classification because classifications are legible and enforceable. Manufacturers and consumers inherit the harder verification problem, which is why exclusion must not be reported as proof that the remaining market is secure.
The action may reduce one real category of risk without resolving the rest. A border can block a product, but it cannot inspect the behavior of every product allowed through.
The Wall Street Journal’s July 31 article Citadel Buys Situational Awareness’s Stock Portfolio After Big Losses in AI reports the same promotion of partial evidence in a setting where the correction arrived through margin calls rather than regulation. The article says the AI-focused hedge fund sold the bulk of its stock portfolio to Citadel after deep losses, having amassed well over $20 billion under management through large leveraged bets tied to the AI trade.
The Journal also reports that founder Leopold Aschenbrenner was seen by some investors as an AI oracle and that other investors closely tracked the fund’s movements. That supports a bounded inference that technical proximity and a compelling thesis about AI infrastructure helped create confidence in an adjacent form of judgement.
The inference should not be inflated into a complete account of every investor’s motives. It is enough to observe that technical foresight and portfolio construction are different objects of evidence even when they point toward the same industry.
A person can correctly anticipate that advanced AI will require enormous quantities of compute, memory, networking, energy, and data-center capacity while remaining wrong about which firms will capture the value, when markets will price it, how positions will correlate, and how much leverage the thesis can survive. Technical foresight asks what may happen, while portfolio construction asks what can be owned, at what price, under what downside, with what liquidity, and for how long.
Leverage makes the distinction brutal because it removes time. An unleveraged investor can be early and wait, while a leveraged investor can be directionally correct and still be forced to sell before the thesis matures.
The portfolio can fail operationally while the technological argument remains intellectually defensible. That does not prove the AI thesis wrong because it proves that thesis quality and risk management require different evidence.
Prestige often erases that boundary. A successful founder becomes a public-policy authority, a celebrated engineer becomes an organizational strategist, a scientist becomes a business oracle, and a compelling writer becomes an allocator of billions.
The initial expertise may be real, which makes the unsupported transfer harder to challenge. The institution granting authority receives the comfort of association with brilliance, while investors inherit the difference between insight and discipline when the adjacent judgement finally meets a condition capable of punishing error.
Who Gets the Credit and Who Gets the Cost
The articles do not describe the same technology, but they expose the same incentive structure. Advancing the Price-Performance Frontier with GPT-5.6 allows adopters to describe workloads as viable before they count the operating burden, while Microsoft Cloud and AI Strength Fuels Fourth Quarter Results allows customers to describe procurement and usage as transformation before they prove changed work.
AI Productivity Gains Are 10 Percent, Not 10x shows how generated output can be recognized before downstream effort is counted. Advancing the Next Era of National Science shows how access can be reported before results survive validation.
FCC Updates Covered List to Include Foreign-Produced Advanced Robotic Devices and Power Inverters shows how exclusion can be counted before admitted devices prove their behavior. Citadel Buys Situational Awareness’s Stock Portfolio After Big Losses in AI shows how technical foresight can be rewarded before adjacent judgement survives leverage.
In every case, the proxy allows the actor nearest the decision to recognize success early. The cost of being wrong moves outward or downward toward the people responsible for maintaining, validating, repairing, or financing what the proxy did not prove.
This is why better measurement alone will not solve the problem. The proxy was not selected only because the institution lacked imagination because it was selected because it aligns with what the institution can count, what its leaders can report, and when they want credit.
Vendors can count sales, executives can count deployment, regulators can count exclusions, programs can count participants, and funds can count assets. Operators are left counting what happened afterward, which is the hidden accounting system beneath the proxy economy.
Complex institutions cannot abolish proxies because leaders need indicators before final outcomes arrive, researchers need provisional measures, engineers need telemetry, markets need forecasts, and regulators need classifications. The corrective is to stop allowing the beneficiary of the claim to hide the conversion between indicator and outcome.
Every promoted proxy should face three questions in the same review. The institution should state what was actually measured, what larger conclusion is being claimed, and who inherits the cost if the connection between them fails.
The third question is the one most measurement frameworks omit because it exposes the transfer. When token cost falls, the adopter must show the complete operating economics and name who absorbs remediation when the workflow fails.
When access expands, the sponsor must show validated research outcomes and preserve the method needed to reproduce them. When engineering output rises, leadership must show system-level value and count the work transferred into review, integration, security, and support.
When seats are purchased, the customer must show changed work rather than a completed rollout. When origin determines admission, the regulator and buyer must still show the security properties of what remains, and when expertise is promoted into adjacent authority, the institution must show that the relevant judgement has survived conditions that can actually punish error.
Without that bridge, the institution has not proved its claim. It has selected a nearby fact, collected the credit, and assigned the remaining distance to someone with less authority.
AI did not invent this behavior, but it has industrialized it. The numbers arrive faster, the claims grow larger, the decisions become more expensive, and the people choosing the proxy are increasingly separated from the people who pay when it fails.
That is the proxy economy, and its central shortage is not data. It is disciplined inference joined to consequence because a number should not become a victory until the people carrying the cost can see the bridge between them.
Source Articles
The reporting basis includes OpenAI’s Advancing the Price-Performance Frontier with GPT-5.6 and Advancing the Next Era of National Science. It also includes Microsoft’s Microsoft Cloud and AI Strength Fuels Fourth Quarter Results.
The engineering section uses Justin Reock’s AI Productivity Gains Are 10 Percent, Not 10x. LeadDev resurfaced the argument in its July 30 newsletter, while the canonical DX article supplies the direct public link and underlying evidence.
The final section uses the FCC document FCC Updates Covered List to Include Foreign-Produced Advanced Robotic Devices and Power Inverters and The Wall Street Journal’s Citadel Buys Situational Awareness’s Stock Portfolio After Big Losses in AI. The FCC document supplies the regulatory action and stated risk basis, while the Journal supplies the hedge-fund reporting and the basis for the bounded inference about expertise transfer.
Artifacts are cheap, judgement is scarce.
Per ignem, veritas.



