Wednesday, August 26, 2026
Forge News Breakdown
A new paper examining agentic AI in military command and control reviewed 240 documented testing and evaluation practices across eight evaluation dimensions and three lifecycle stages. The researchers found eight assumptions underneath established testing methods, grouped around system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight. [1]
The important finding is narrower, and harder, than “agents are difficult to test.” The researchers argue that the evidence itself can remain valid while the argument connecting that evidence to future fielded behavior becomes insufficient. A system can satisfy the testing process without the result supporting the level of confidence attached to deployment. Persistent memory changes state. Agents select tools and information at runtime. Subagents can enter after certification. Variation compounds across trajectories. Assemblies produce behavior that component tests may never expose. [1]
The timing is not academic. In June, the U.S. Department of War launched Agent Network, an AI-enabled system intended to continuously scan intelligence and operational systems and present commanders with options within seconds. The department says Agent Network will not autonomously select or strike targets, will keep human judgement at the center of consequential decisions, and will undergo rigorous testing, operational evaluation, and oversight throughout development and fielding. [2]
Forged Analysis
That commitment sounds responsible. The new research identifies why fulfilling it is more difficult than writing it.
Traditional assurance depends on correspondence. We test a particular system under particular conditions and use those observations to justify confidence in the system we later operate. The test never captures everything, but the relationship between the test article and the fielded article has to remain stable enough for the inference to hold.
Agentic systems can weaken that correspondence without a conventional release event. Memory changes. Available tools change. Retrieved information changes. Delegation changes the assembly. Each intermediate action becomes context for the next one. The model version can remain identical while the effective operating system drifts away from the thing that generated the original evidence. [1]
In The Illegibility Crisis, I use the same distinction at the organizational layer. A visible signal can remain intact after the inference it once supported stops being safe. [3] The assurance problem here has the same shape. The test did not become fake. What changed is how much the organization is justified in believing because the test passed.
That changes what “continuous assurance” has to mean. Monitoring whether an agent is running is insufficient. The operating system needs to know whether its existing evidence still describes the system carrying authority now. The paper points toward bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, measured run-to-run variance, re-baselining, staged fielding, and explicit expiry conditions for evidence. Where uncertainty cannot be eliminated, somebody has to own it. [1]
Counter-pressure
This is a preprint and an analytical synthesis, not experimental proof that every current military AI testing regime is invalid. The authors explicitly argue that narrower assurance claims remain recoverable, and tightly constrained agents with fixed tools, limited state, and narrow mission envelopes present a very different problem from persistent multi-agent systems operating against adversarial information. [1]
That distinction keeps the argument useful. The answer is not to declare agentic systems untestable. It is to stop allowing a broad assurance claim to outrun the evidence underneath it.
Operating Takeaway
Do not ask only whether the agent passed.
Ask what was actually tested, which operating conditions the evidence covers, what can change without triggering re-evaluation, how behavioral drift becomes visible, when the evidence expires, and who has authority to narrow or revoke the system’s operating scope when the assurance case no longer holds.
A test result is evidence. Authority is a separate decision.
Artifacts are cheap, judgement is scarce.
Per ignem, veritas.
Sources
[1] U. Richard, H. Frase, S. Cao, D. Cooke, S. Kwon, and A. Tan, “Testing and Evaluation of Agentic AI Systems In Military Command and Control,” arXiv:2608.20597, Aug. 20, 2026. [Online]. Available: Testing and Evaluation of Agentic AI Systems In Military Command and Control on arXiv. [Accessed: Aug. 26, 2026].
[2] U.S. Department of War, “DOW Unleashes ‘Agent Network’ to Transform AI-Enabled Battle Management and Targeting,” June 25, 2026. [Online]. Available: Department of War release on Agent Network. [Accessed: Aug. 26, 2026].
[3] P. LaPosta, The Illegibility Crisis: Instrumentation for AI-Era Leadership, 1st ed. Forged Culture, 2025. [Online]. Available: Leanpub. [Accessed: Aug. 26, 2026].



