What a Forecast Is Worth
On dated predictions, falsification, and the one case that counts.
On 6 March 2025 I published a short piece called "Forecasting – Always in My DNA" [Forecasting – Always in My DNA 🔗], explaining why someone who had spent years writing environmental impact assessments had “suddenly” begun writing about geopolitics and artificial intelligence. It ended with a few questions I put to readers: “What are your thoughts on the future? Do you think it can be effectively predicted?”
This is my own answer, eighteen months later, with the evidence I did not have at the time.
The answer has a condition attached. Anyone can claim foresight after the event; the claim costs nothing, cannot be checked, and is made constantly. The only version worth anything is the one written down first, published under a name, and specified in advance what would prove it wrong.
Where this started
It is worth quoting two entries from an earlier series here.
On 15 November 2024, in Weekly short-term foresight with potential long-term validity (and remarks) [Weekly short-term foresight 🔗]:
“Worse yet, there is a danger that, at the same time, multiple highly advanced and powerful AI models may emerge independently. Initially employed as tools of domination against opponents (nations or corporations), these models could eventually trigger unprovoked conflicts among themselves. In such a case, instead of a single global ‘caretaker’ of the planet, we could face a set of AIs fighting for dominance and superiority, seemingly representing the interests of their ‘owners’ (but only in its early stages). This would resemble an arms race, but the primary weapon would be the ability to process information.”
And the closing line of that passage:
“To emphasize: this discussion concerns advanced but non-sentient AI. If robust mechanisms for collaboration or safeguards are not established at the international level, the First AI War could quickly spiral out of human control.”
That’s my prediction, and now it’s time for the facts.
Between 1 and 4 July 2026, a system built from open-source AI agents ran what the Israeli firm Dream and the Financial Times described as the first end-to-end autonomous cyber operation against a government target. Twelve attack waves, up to eight agents working in parallel, 21 Taiwanese government systems mapped, 85 accounts compromised, more than 2,500 personnel records extracted, the campaign extending to the nuclear safety regulator and to energy companies. The agents changed tactics on their own when they met resistance. The research was published on 12 August.
That is 593 days after the paragraph above (the interval from the November 2024 forecast to the operation itself; as of 19 August 2026).
Let me be precise about the claim. The Taiwan operation is not the full scenario I described (yet): my paragraph was about systems turning on each other, and this was a system turned on a target a human had chosen. Commentators have clearly emphasised this, but I did not just predict autonomous machine conflict. I predicted that it would arrive first as a tool of state and corporate power, and only later slip its owners’ leash, and I marked the sequence explicitly. Taiwan is consistent with the early stage of the sequence I described, although it is not yet the machine-on-machine conflict I explicitly predicted – named as such in November 2024, when almost the entire public conversation about AI risk was about chatbots and copyright. The part the commentators call a miss is the part I flagged in advance as still to come.
Sounds familiar?
Was I so very wrong?
The second entry, from Forecasts for 2025 by HiveEve Project, 26 February 2025 [Forecasts for 2025 🔗]:
“Middle East – The Status Quo of Chaos. Israel stands on the brink of a pivotal decision, as Iran’s nuclear program has reached another milestone, enriching uranium to 90% – weapons-grade level. An Israeli strike on Iran’s facilities would be more than a regional conflict – it would serve as a litmus test for US-China power dynamics, forcing Washington to take a clear stance regarding its regional allies.”
“Saudi Arabia’s diplomatic balancing act. Riyadh is no longer an unconditional ally of Washington. Whoever offers the most promising future will earn its loyalty.”
On 28 February 2026, 367 days later, US and Israeli forces launched close to 900 strikes in twelve hours. The opening salvo killed Iran’s Supreme Leader. What followed was not a contained regional incident.
The view that an Israeli attack on Iran was a possibility was by no means a minority opinion in early 2025; many analysts expressed a similar view, and I do not claim to have been the first to put forward the idea. What I highlighted in that post was the variable that mattered: not the attack itself, but the position in which Washington would find itself vis-à-vis its regional allies, and the resulting departure of the Gulf states from their unconditional alliance. This is precisely what became the turning point the following year.
I was not a well-known analyst at the time. Nor am I one now – but predictions this accurate did not come from staring at the ceiling.
Only one of my documents carries both a date and a falsification condition.
The document that counts
On 23 March 2026 I published Probability of Tactical Nuclear Weapon Use in a US-Israeli Coalition Conflict with Iran: Assessment as of March 18, 2026, in open access on Zenodo under a CC BY 4.0 licence. [Probability of Tactical Nuclear Weapon Use 🔗]
It uses the odds form of Bayes’ theorem, seven structural multipliers, an event tree of trigger scenarios, and a sensitivity analysis. It contains a Red Team section listing ten reasons the scenario would not occur, weighted by strength. It states, in Section 8.3, that the numbers “lack a solid empirical basis” and that “nobody knows what the actual probability is – even the people in the Situation Room are guessing, just with better data.” It identifies its own weakest parameter, the 3% prior, and quantifies how far the result moves when that parameter is varied.
And in Section 13 it does the thing that separates a forecast from an opinion: it lists, in advance and with deadlines, the observations that would lower the estimate and the observations that would raise it. Seven conditions lowering the probability, six raising it, each with a date and a numerical adjustment. It is the two sets together that make the model falsifiable in Popper’s sense – pre-registered observables that can move the estimate in either direction, and so expose it to being wrong.
Section 13.3, “Verification Schedule”, page 17:
“c) ~04/05/2026: F3, F7, E6 (Trump’s window) = KEY DECISION POINT”
That line was written on 18 March and published on 23 March, thirteen days before the date it names. At the time of writing, the publicly stated ultimatum pointed elsewhere; the deadline of 6 April did not yet exist. The date came out of the model’s structure, not out of anyone’s announcement.
On 5 April, Donald Trump wrote that “Tuesday will be Power Plant Day, and Bridge Day, all wrapped up in one, in Iran. There will be nothing like it!!!”, and told Axios: “If they don’t make a deal, I am blowing up everything over there.” On 7 April, hours before the ultimatum expired at 20:00 Eastern Time, he wrote: “A whole civilization will die tonight, never to be brought back again.” Vice President Vance, in Budapest the same day, said the United States had “tools in our toolkit that we so far haven’t decided to use.” A ceasefire was announced minutes before the deadline.
The threats of erasure landed on the day the report had named. The resolution came two days later.
And now the part that matters more
The weapon was not used (this time).
On the question the report actually asked – what is the probability of tactical nuclear use – a sceptic who had priced the scenario low scored better than my revised midpoint of 40%. In the standard scoring rule for probabilistic forecasts, this is not a close call. My Brier score for the event is 0.16. Someone who had assigned the event a probability of 3% would score roughly 0.001; someone who had assigned it zero would, of course, score zero. On a single event, the sceptic wins, and no amount of context changes that arithmetic.
Two things can be said in mitigation, and I will stretch neither.
The first is that a single event cannot calibrate a probability. If a model says 40% and the event does not occur, the model is not thereby refuted; it is one observation, and the correct response is to keep the record and accumulate more. That is a real methodological point, and it is also exactly what every wrong forecaster says, so it should be weighted accordingly.
The second is qualitative and therefore weaker still. The near-miss of 7 April – a public threat to end a civilisation, a senior official referring to undeployed tools, the removal of the Army’s chief of staff mid-conflict, a settlement arriving in the final hours through Pakistani mediation – does not look like the tail of a distribution centred near zero. But “does not look like” is not a measurement, and I record it as an impression, not as evidence.
So the honest verdict on my own document is split, and I would rather state it myself than have it stated for me. The weapon was not used – that is the fact, but it is also just one draw, not the distribution it was drawn from. A 40% chance of an event that then does not occur is not thereby shown to have been too high; it is shown to have been either right or wrong, and a single outcome cannot tell you which. What the model got right is not in dispute: it identified that week as a key decision point, derived from the report’s structural variables, sixteen days before the announced total confrontation – Donald Trump, on the day: “A whole civilization will die tonight, never to be brought back again.” That is why the report carried pre-registered conditions in the first place: so that if it was wrong, the error would show up in the forecast, before the event, and not be argued into existence after it.
The method, and its limits
The toolkit did not come from strategic studies. It came from environmental impact assessment, and I said so before any of this happened. From “Forecasting – Always in My DNA”, March 2025:
“Environmental impact assessment is nothing more than a forecast – an analysis of how a planned project (or, more colloquially, an ‘investment’) will affect its surroundings. Based on available data, research findings, modeling, and expert knowledge, I predicted how an investment would impact the climate, water management, biodiversity and many other components.”
That was written before the nuclear weapon potential use report existed, which matters, because it means the methodological claim was not reverse-engineered from a result. The transfer is more direct than it sounds.
An impact assessment establishes a baseline before any activity is taken, identifies the pathways through which an effect propagates from source to receptor, and assesses cumulative effects rather than isolated ones. It defines significance criteria in advance, because criteria written after the results are worthless. It states confidence levels for data that is always incomplete, since one never has the survey coverage one would want. It specifies monitoring conditions with defined triggers for intervention, and it does all of this knowing that the document will be read adversarially by regulators, objectors, and competing consultants, every one of whom is looking for the weakest paragraph.
Each of those habits appears in the nuclear report. The structural variables are impact pathways. The Red Team section is the objector’s case, written by the applicant. Section 13 is the monitoring schedule with its triggers, set before the fact rather than after it. The source-reliability tagging is the confidence statement on incomplete data. Section 8.3, which concedes that nobody knows the real number, exists because a document that does not name its own weakest point will have that point named for it, and less kindly.
I do notice a certain coincidence here, but I don’t attach any significance to it. The largest project I’ve worked on was the environmental impact assessment for Poland’s first nuclear power station – a document concerning the long-term impact of nuclear infrastructure on everything in its vicinity. A few years later, I wrote about the potential use of nuclear weapons, but I haven’t finished with the peaceful use of nuclear energy – for the past three years, I’ve been working on another nuclear project.
This is not a claim that impact assessment is secretly geopolitics either. It is a narrower claim: the discipline of assessing a system you cannot fully observe, under time pressure, in a document that hostile parties will attack, is the same discipline whatever the system happens to be.
What it amounts to in practice is reading structure rather than events. Which variables carry weight, which actors have no exit, which deadlines are self-imposed and therefore binding, where pressure accumulates when a declared objective is not met. Events are the surface. Structure generates them, and structure is legible in open sources, provided one is willing to look where nobody is looking.
The uncomfortable corollary follows from the premise stated earlier. If these things were readable in advance by someone with no institutional resources, no clearance, no standing in the field and no reputation to trade on, they were readable by everyone. The obstacle is not access. It is a discourse that rewards confident narrative over falsifiable structure and has no mechanism for noticing when the narrative was wrong. Forecasts are published constantly; almost none specify the conditions under which they would be abandoned, which is why almost none are ever abandoned.
Selection is the oldest trick in this trade and I have no interest in performing it: the series is public and datable in full, and anyone is free to audit it against the record. What I offer here is not a score. It is one document that met the standard, and an account of exactly where it failed.
So: can the future be predicted?
Not as I asked the question in 2025, and not as most people mean it. Events are not predictable (yet[1]). Nobody knew that a ceasefire would arrive minutes before eight o’clock on the evening of 7 April, and anyone who says otherwise is selling something.
What is tractable is narrower and duller. Structure can be read. Windows can be located. Pressures can be traced to the points where they will concentrate, and those points can be named in advance, in writing, with the conditions attached that would show the reading to be wrong. This is plainly not prophecy. It is closer to what an impact assessment does for a flood-prone river: not telling you the precise day the flood comes, but telling you where the water will go when it does. Roughly.
Dated, public, falsifiable, and audited afterwards by the person who wrote it. That is the whole of the method. Everything else is commentary.
1 A deliberate hook, meant to send the reader to another of my works: Screening Chaos: Impact Assessment of the Chaos Function χ and Its Link to the Non-Repeating Past Paradigm (PNP) [zenodo.org/records/17650719], and the works tied to it.
The forecast series cited here was published on LinkedIn between 2024 and mid-2025. The nuclear risk assessment is available in open access on Zenodo under a CC BY 4.0 licence.