AI · Truth & Epistemics October 5, 2026 11 min read

Malaysia’s AI Hub: Who Controls Recovery During a Crisis?

A hand operates a switch between server racks and branching routes, with Kuala Lumpur in the background.

Malaysia can add data centres, attract cloud investment and build a larger AI ecosystem. But when power fails, international connectivity becomes constrained and replacement hardware is difficult to obtain, who can actually restore the service?

The answer depends on which part of the system has failed. Government holds authority over legislation and coordination. Operators have the access and expertise to perform technical work. Suppliers can determine which options are available when replacement components or capacity are needed.

The resilience of Malaysia’s AI hub should be assessed through the functions that can be restored, by whom, with which resources, and on the evidence of which tests. The number of data centres indicates capacity. Recovery evidence shows how far that capacity remains usable when conditions change.

This distinction matters for AI sovereignty: the ability to make independent decisions becomes more meaningful when those decisions can be carried out.

Investment creates options; testing establishes capability

Malaysia has tangible cloud infrastructure. AWS’s official region table lists Asia Pacific (Malaysia) with three Availability Zones. Azure’s table also lists Malaysia West with three zones. This infrastructure creates options for building applications that can withstand certain failures. Sources: AWS Regions, Azure region list.

An application running in only one zone does not gain resilience across multiple zones simply because its region has three. The services used, configuration, data replication and application design still need to be examined. AWS documentation explains that resilience is a shared responsibility between provider and customer. With EC2, for example, customers manage the resilience configurations they need. Source: AWS shared responsibility for resiliency.

Recovery tools are also becoming more available. On 21 July 2026, AWS announced the availability of Elastic Disaster Recovery in its Malaysia region. That establishes the availability of a service option. Successful recovery of a particular application still requires evidence of that customer’s configuration, exercises and test results. Source: AWS DRS announcement.

Policy also shapes the direction of development. Malaysia’s Ministry of Digital launched the National Cloud Computing Policy on 13 August 2025, with the aim of a sovereign, secure, inclusive and sustainable digital ecosystem. Implementation needs to be assessed separately from the goals announced. Source: Ministry of Digital.

A useful assessment question is: when a dependency becomes unavailable, which functions still work, and how does the organisation demonstrate that?

Three forms of power in Malaysia’s AI hub

To answer that question, we need to distinguish legal authority, operational control and the influence created by dependency. The following distinction is an analytical framework, rather than an official hierarchy covering every agency or company.

Legal authority: who can set rules and coordinate?

Legal authority, or de jure power, comes from legislation and institutional mandates. NACSA’s legal page describes the framework of the Cyber Security Act 2024, including the roles of sector leads and National Critical Information Infrastructure entities, and the management of cyber threats and incidents affecting that infrastructure. Source: NACSA.

NADMA describes its role as the lead and coordinating agency for national disaster management and disaster risk reduction. Source: NADMA’s role.

These mandates help explain the channels for coordination. Determining the actions required in a specific incident still depends on the nature of the disruption, jurisdiction and applicable procedures. The existence of a cyber incident or disaster framework alone does not identify who will carry out every action across the grid, networks and applications.

Operational control: who can change the situation now?

Operational control, or de facto power, rests with those who have the access, skills and resources to act.

GSO states that it is responsible for the real-time operation and management of Peninsular Malaysia’s grid, including generation planning, scheduling and dispatch. This geographical scope matters: the role should not be generalised to every electricity system in Malaysia. Source: Grid System Operator.

At the application level, a customer’s team may need to activate a recovery plan, validate data and change how the service operates. The division of responsibilities depends on the services used, as described in the AWS documentation above.

A plan can look complete on paper yet stall when an emergency account cannot be used, qualified staff are unavailable or replacement capacity has not been secured. This is why named action owners and evidence of exercises need to accompany the technical design.

Dependency power: who determines the replacement options?

Structural power emerges when an organisation needs something that is difficult to replace: particular compute capacity, spare parts, model services or cable repairs. A supplier’s influence increases when the customer has few usable alternatives.

For submarine cables, ICPC describes an established global maintenance framework involving repair vessels, marine engineering expertise and maintenance agreements. It also emphasises efficient permitting for repair work. International cooperation can be a source of resilience here. Source: ICPC explanation.

Dependency becomes harder to manage when alternatives have not been verified, capacity is insufficient or contracts do not explain how to obtain replacements during disruption. These conditions explain limits on operational options. They do not establish that suppliers deliberately intend to constrain their customers or Malaysia.

Power, networks and chips operate on different timescales

Imagine overlapping disruptions: one site loses power, international connections become congested and an order for new hardware is delayed. All three can put pressure on AI services, but through different mechanisms.

Three disruption layers and recovery questions
LayerRecovery question to answer
Power and facilitiesCan backup power and cooling support the selected functions, and who verifies this?
NetworksAre alternative routes genuinely distinct, reachable and backed by sufficient capacity?
Compute and hardwareAre replacement capacity or parts available, compatible and obtainable under the applicable conditions?

This table is a scenario framework. It provides no measured recovery times and makes no guarantee about any facility’s capabilities.

A power disruption may demand an immediate response. Redirecting traffic requires routes that remain usable. Chip supply pressure may emerge during hardware replacement or capacity expansion. In this scenario, the component that most constrains recovery can change as the incident develops.

Trade rules also need to be read precisely. In a statement dated 14 July 2025, MITI announced a strategic trade permit requirement for the export, transshipment and transit of high-performance AI chips originating in the United States. The statement concerns specific transaction types. It does not say that all chip imports are prohibited or that chips already in operation will be switched off. Source: MITI statement.

The statement is used here as historical evidence of the trade dimension of chip supply chains. Requirements for current transactions need to be checked against the applicable rules and authorities. This article does not assess whether a particular transaction is eligible.

Three assumptions that can mislead

“Data is stored in Malaysia, so all inference happens in Malaysia”

Storage location and processing location need to be checked separately. Microsoft Foundry documentation states that, for Global deployments, prompts and responses may be processed in any geography where the relevant model is deployed. DataZone uses the data zone boundaries defined by Microsoft. Data stored at rest follows the geography designated by the customer. Source: Foundry privacy and processing locations.

A location label alone therefore does not resolve the question of data sovereignty. Organisations need to know which services and deployment types they actually use. This explanation is specific to the cited Foundry documentation; it does not claim that every cloud service follows the same rules.

“Quota is approved, so capacity must be available”

Azure documentation distinguishes quota from capacity. Quota is a subscription’s permission to use resources; capacity is the infrastructure available in a particular region or zone. Azure also distinguishes Reserved Instances, which provide billing discounts, from on-demand capacity reservations, which provide capacity within the relevant configuration and SLA terms. Source: Azure capacity reservations.

These product distinctions matter in recovery planning. An accepted virtual machine capacity reservation is not a general guarantee that every GPU, AI model or inference service will be available whenever a customer wants it.

“We have two providers, so we are independent of any single failure”

Two providers can broaden the options. Assessment still needs to examine shared dependencies: do both require the same route, identity system or access to the same data source?

These are design questions to answer through operator information and application tests. The number of contracts alone does not show whether a user’s function can be restored.

Start with the user’s function, then trace its dependencies

Consider a constructed example: an organisation uses an AI assistant to search internal documents. When the model provider cannot be reached, it still wants employees to read documents they are authorised to access and check the original sources.

The selected minimum function might be keyword search and document viewing, with AI summaries temporarily suspended. That option is useful only if the search index, storage and login can operate without the disrupted component.

Likewise, switching models may not help if the team cannot open the data, use decryption keys or obtain administrative access. The recovery chain needs to be examined from the user’s perspective: can they log in, obtain valid information and complete the defined task?

This example is not the result of testing a real system. It shows how to turn a broad discussion about “keeping AI running” into a specific function that can be checked. If output accuracy cannot be verified, the plan needs a way to suspend that feature and direct users to appropriate sources or staff.

Six questions before declaring a system resilient

Technology teams, service owners and policymakers can use these questions as a starting point:

  • Which minimum function must remain available? Define the user’s task, the data required and the acceptable accuracy limits.
  • Who can activate recovery? Name the action owner, their backup, the access required and the escalation channel.
  • Which replacements are actually available? Distinguish catalogue options, contractual commitments, allocated capacity and tested capabilities.
  • Where do storage and processing occur? Check the actual configuration, including storage, inference and supporting services.
  • What did the last test show? Compare the recovery time objective, or RTO, with the time recorded; compare the recovery point objective, or RPO, with the state of the data after recovery.
  • What happens when recovery fails? Provide limited functionality, human referral or suspension of features whose results cannot be trusted.

The review record should state the configuration, date, test conditions and items that remain unchecked. Success in one scenario does not establish readiness for every form of overlapping disruption.

A phased plan for gathering evidence

The plan can move through early, middle and final phases according to system scale, access to evidence and team readiness. Progress is determined by each phase’s outputs and the gaps that remain open.

Early phase — Understand functions and dependencies. Select the core functions, map their dependencies and name the action owners. Distinguish what is controlled internally, what the provider controls and what requires other parties. Identify the evidence needed for each resilience claim. The intended outputs are a function map, clear responsibilities and a list of evidence still required.

Middle phase — Check the recovery options. Review recovery access, data configuration, operator support and the terms governing alternative capacity. Distinguish proposed replacements from options already confirmed as available. Record what remains unknown, along with an action owner and a follow-up step. Before testing, clarify which options will be tested, their prerequisites and the gaps that could affect the results.

Final phase — Test, assess and address gaps. Run scenario exercises and recovery tests in an approved, isolated environment. Record actual times, the state of the data, functions restored and reasons for failure. Use the results to improve the plan and determine what needs to be retested. Any test that could affect a live service requires the relevant operator’s procedures and approval.

The phases can overlap or repeat. If testing reveals an unmapped dependency or an unusable replacement, return to the relevant phase. The final phase produces evidence for the scenarios tested and a list of open gaps; it does not mean that every risk has been resolved.

Local capability can be assessed through demonstrated actions: who can diagnose a problem, obtain authorised access and restore the defined function? This offers a more specific measure than training totals or investment size alone.

Assess AI sovereignty function by function

Foreign investment can expand technology and recovery options. Dependencies on it need to be understood through contracts, configurations and usable replacements. Owning every component internally does not, by itself, establish that every function can be restored.

All From AI’s earlier article, Malaysia’s Deep-Tech Leap: Operator to Innovator, discusses local capability and technology ownership. Recovery assessment adds an operational question: when a core function fails, who has the mandate, access, skills and resources to restore it?

The public sources reviewed here establish certain mandates, infrastructure options and product rules. They are insufficient to deliver a comprehensive verdict on the ability of all Malaysian AI hubs to withstand a crisis. This review did not test facilities, inspect confidential contracts or measure actual application recovery. A lack of public evidence also does not establish that plans or capabilities are absent.

For one service you manage or use, try completing this sentence: “If this dependency fails, this person can restore this function, using this replacement; the evidence is this test.” The parts you cannot yet complete indicate where the next review should begin.

Frequently asked questions

Do three Availability Zones guarantee that an application keeps running?

The number of zones indicates infrastructure options. Application resilience still depends on services, configuration, replication and testing. A design also needs to be assessed against the types of disruption it is intended to handle.

Does a chip supply disruption immediately shut down installed GPUs?

A delay in new supplies does not, by itself, shut down hardware already operating. The effect on a particular system depends on spare-part requirements, hardware failures, alternative capacity and available support.

Can this article be used as a national readiness audit?

This article provides an analytical framework and links to selected official sources. An operational audit requires additional evidence about actual systems, contracts, access, capacity and test results.

Method note: this article adapts the supplied research material and checks selected claims against official sources reviewed on 4 October 2026. The power framework, disruption scenario and phased plan are the author’s analysis and proposals. No operational tests were performed.

Continue the conversation

Get the next dispatch.

Occasional essays, tool notes and fiction updates. Confirm by email and unsubscribe at any time.