On Thursday, September 3, 2026, three of the leading generative artificial intelligence services experienced availability problems during the same day. OpenAI reported elevated errors on ChatGPT and Codex. Anthropic observed errors on several Claude models as well as on Claude.ai, its API, Claude Code, and Claude Cowork. Grok, the xAI service now attached to SpaceX, also experienced an outage. During part of the European afternoon, these separate incidents overlapped.
The image is striking: three AI giants disrupted almost simultaneously. Yet it calls for a precise reading. The available information does not demonstrate a general failure of artificial intelligence or the failure of a common technical provider. The publicly released descriptions differ and do not allow a single cause to be established. What the episode reveals, however, is very concrete: a company that has turned these platforms into production components can lose several solutions it believed to be independent within the same operational window.
OpenAI: a mitigation applied in 34 minutes, then monitored for about 1 hour 38 minutes
According to OpenAI, the incident began on Thursday, September 3 at 7:43 a.m. Pacific Time, i.e. 14:43 UTC and 16:43 in Belgium. A routing error made ChatGPT and Codex unavailable for some users across several platforms. The term is important: OpenAI did not announce a total shutdown of all its services for all its customers, but an increase in errors and unavailability affecting some users.
The first technical phase was relatively short. An OpenAI spokesperson indicated that a fix had been applied around 8:17 a.m. Pacific Time, i.e. 15:17 UTC. Approximately 34 minutes therefore elapsed between the announced start of the routing error and the application of the corrective measure. At that point, however, OpenAI did not immediately close the incident: the company moved into a monitoring phase in order to verify recovery.
The public status page ultimately classified the event as resolved at 18:55 CEST, Central European Summer Time, i.e. 16:55 UTC. If we measure the administrative duration between the opening of the investigation at 14:43 UTC and the closure of the incident at 16:55 UTC, we get approximately 2 hours and 12 minutes. This duration does not mean that every user remained blocked throughout that entire period. It includes the 34 minutes before mitigation, then approximately 1 hour and 38 minutes of recovery and monitoring.
OpenAI also reported a specific consequence: some Codex remote control users might have needed to re-pair their mobile device after the incident. This is a useful example of the side effect an outage can produce, even after the main service has returned. A technical recovery does not always guarantee that all sessions, connections, or automations instantly return to their previous state.
Anthropic: nearly three hours of impact on several models and interfaces
The Anthropic incident began earlier. At 13:26 UTC, the company announced that it was investigating an increase in errors affecting requests sent to several Claude models. At 13:41, Anthropic said it had identified the cause, without publishing the technical details on its status page. A spokesperson later described to The Register an “infrastructure issue” that had caused a partial outage.
At 13:50 UTC, Anthropic expanded the list of affected models: Mythos and Fable 5.1, Mythos and Fable 5, Opus 5, Opus 4.8, and Opus 4.6. The incident did not concern only the conversational interface. The official page states that it affected Claude.ai, the Claude API, Claude Code, and Claude Cowork. For an organization, this distinction is important: when an API and integrated tools are affected at the same time as the website, the potential effects go beyond the inability to start a manual conversation.
Anthropic continued working on the fix during the afternoon. At 15:25 UTC, most models had returned to their usual error rate, but Opus 4.8 and Opus 5 remained affected. A fix was deployed at 16:06 UTC, then the company indicated that the impact had ended at 16:16 UTC. The official closure was published seven minutes later, at 16:23.
Between the first public alert at 13:26 and the declared end of the impact at 16:16, approximately 2 hours and 50 minutes elapsed. Taking the final publication at 16:23, the ticket remained open for 2 hours and 57 minutes. The technical start time of the errors, however, is not specified. The Register refers to an outage of 3 hours and 6 minutes, a figure that does not exactly match the boundaries currently displayed on Anthropic’s status page. This discrepancy reinforces the need to distinguish between technical duration, impact duration, and the administrative duration of the ticket. In any case, recovery was gradual: only two models were still affected at 15:25.
Grok: an incident linked to the Memphis data center
The third incident concerned Grok. xAI’s status page indicated that the company had begun examining the problems around 6:30 a.m. Pacific Time, i.e. 13:30 UTC. The initial message remained general: Grok was experiencing difficulties and the teams were working to restore service.
Later in the day, SpaceX apologized for the problems encountered and attributed the incident to an outage at its Memphis data center. The company also apologized to affected compute partners, then indicated that systems had been restored and were operating normally.
It is not possible, based on the publicly accessible information, to establish a restoration time as precise as for OpenAI or Anthropic. The start of the investigation, around 13:30 UTC, is documented; the exact time at which the impact ended is not sufficiently documented to calculate a reliable duration. Any more precise quantified duration for Grok would therefore be an estimate and not a confirmed fact.
How long did the incidents overlap?
Anthropic had published its alert at 13:26 UTC when xAI began investigating around 13:30. OpenAI then opened its incident at 14:43. Between 14:43 and 15:17 UTC, the three providers therefore simultaneously had publicly open or reported incidents, while OpenAI had not yet applied its corrective measure and Anthropic was continuing its work.
This 34-minute window documents with certainty an overlap in reports. It does not allow us to state with the same precision that every user of the three services experienced 34 minutes of simultaneous unavailability, due to the lack of an exact public restoration time for Grok. Nor should it be confused with the full duration of each incident: Anthropic declared an impact until 16:16 UTC and OpenAI monitored its recovery until 16:55.
Different published causes, no common outage demonstrated
The simultaneity immediately raised the hypothesis of a shared technical dependency. The three companies do indeed use numerous infrastructure, network, and security services. The Register contacted Cloudflare, used to varying degrees by the three players. Cloudflare responded that its services were operating normally and that no significant disruption was underway. The AWS, Google Cloud, and Microsoft Azure status pages showed no relevant incident at the time, while bearing in mind that these public dashboards can publish some problems with a delay.
The published explanations do not allow a common cause to be established: OpenAI cites a routing error, Anthropic an unspecified infrastructure issue, and SpaceX an outage at its Memphis data center. However, they are also not sufficient to demonstrate that the events were entirely independent. Cloudflare and the status pages of the major clouds reduce the plausibility of a visible common outage, without excluding any unpublished shared dependency. Presenting the episode as a coordinated outage or as the collapse of a single cloud provider would therefore be misleading.
The real lesson lies elsewhere. Several reported failures can simultaneously produce the same effect for the end customer, whether they are entirely separate or share a dependency that is still unknown. From the point of view of a company whose teams or applications are waiting for a response from a remote model, the exact origin does not change the immediate consequence: the work can no longer be executed as planned.
The immediate consequences for users
The directly documented consequences are elevated error rates, services unavailable for some users, and degradation of several interfaces and APIs. A human user could see a request fail, wait for a response that never arrived, or have to try again. An application connected to an API could receive an error, exceed its timeout, or place a task in a queue.
The business consequences described below are not all reported as having actually occurred at customers during this incident. They illustrate the risks an organization exposes itself to when its workflows depend directly on a remote API without a degraded mode. A support assistant may stop preparing responses. A document-processing chain may accumulate delays. A development tool may interrupt code generation or review. An automated agent may fail in the middle of a procedure. Poorly configured retry mechanisms may abruptly increase the number of requests at the very moment when the service is weakened. When recovery occurs, queues still have to be emptied, errors reconciled, and incomplete actions checked.
The more AI leaves the stage of an individual tool to become a building block of automation, the more an outage resembles a classic information-system incident: interruption, delay, loss of context, partial operations, and the need for controlled recovery.
Why multiple subscriptions do not necessarily constitute resilience
At first glance, the answer seems simple: if one provider goes down, it is enough to use another. This strategy sometimes works for an employee who manually writes or summarizes a document. It becomes much less obvious for an integrated application.
Providers do not use exactly the same APIs, the same models, the same limits, the same tools, or the same authentication mechanisms. A prompt optimized for one model may behave differently on another. The business context or document index may not be available from the second provider. Security policies, logging, connectors, quality assessments, and contractual guarantees may also differ.
A company may therefore have three accounts and no operational failover. As long as it has not prepared interface compatibility, data availability, selection rules, secrets management, and the recovery procedure, it has commercial alternatives, not necessarily technical alternatives.
This situation corresponds to the risk of concentration in information technologies: a function becomes vulnerable when its continuity depends on a provider that is difficult to replace within the time actually available. The DORA regulation explicitly governs this risk in the European financial sector. The principle is broader: resilience is not measured by the number of logos appearing in a contract, but by the real capacity to substitute a component without stopping the business function.
What this incident says about European sovereignty
The problem is not that these providers are American. International platforms provide high-performing models, considerable capacity, and speed of access that are difficult to reproduce in isolation. Reducing sovereignty to the nationality of a provider would prevent us from understanding the operational risk.
The question is one of control. Can an organization choose its model? Can it change provider without rebuilding its entire application? Do its internal knowledge assets remain accessible in a format it controls? Can it decide that sensitive data does not leave an approved perimeter? Can it maintain a priority service when an external provider becomes unavailable?
European data localization answers part of these questions, but not all of them. Data can be stored in Europe while remaining linked to a service, an API, a model, or a control plane governed elsewhere. Operational sovereignty therefore requires combining localization, governance, portability, reversibility, and continuity. The European Data Act is gradually facilitating switching providers of data processing services, but a contractual right to switch does not by itself create an architecture capable of failing over in a few minutes.
Conclusion: the usefulness of local capacity
The Thursday, September 3 incident finally shows the concrete value of local AI infrastructure: not to systematically replace the cloud or claim to reproduce all the models of major providers, but to preserve an execution path under control when those providers become unavailable. An open-weight model installed on local capacity, combined with private RAG and a controlled API gateway such as OPA Core, could have provided a degraded mode for certain compatible uses, maintained access to sensitive knowledge within the approved infrastructure, and reduced the impact of an external outage. The sources for this incident do not demonstrate that a particular local architecture would actually have ensured this continuity; it depends on sizing, monitoring, backups, and above all tests carried out before the incident. In this context, local becomes a pragmatic component of a hybrid architecture: a means of transforming total dependency into a governed continuity capability.
Sources
OpenAI Status, “Elevated errors across ChatGPT and Codex”, September 3, 2026: https://status.openai.com/incidents/01M1KWEDH417T2CF44YYHZDFCR
Anthropic Status, “Elevated errors for multiple models”, September 3, 2026: https://status.claude.com/incidents/461yvfrzpwtt
xAI Status, Grok incident INC25664c15: https://status.x.ai/grok-com/INC25664c15
The Register, “True AI-pocalypse as ChatGPT, Claude, and Grok all go down at once”, September 3, 2026: https://www.theregister.com/ai-and-ml/2026/09/03/chatgpt-claude-and-grok-all-had-outages-at-the-same-time/5294322
European Commission, COM(2020) 67, “Shaping Europe’s digital future”: https://eur-lex.europa.eu/legal-content/FR/TXT/?uri=CELEX:52020DC0067
DORA Regulation (EU) 2022/2554: https://eur-lex.europa.eu/legal-content/FR/TXT/?uri=CELEX:32022R2554
European Data Act (EU) 2023/2854: https://eur-lex.europa.eu/legal-content/FR/TXT/?uri=CELEX:32023R2854
This analysis concerns operational resilience and does not constitute legal advice.
Tom Cheniaux - rephrased using AI
Let's talk about it