Amaze Contact →
Sovereign AI

Data sovereignty and AI: keeping data onshore

Data sovereignty isn't just where data is stored. It's who controls it, under what jurisdiction, and who can compel disclosure for AI workloads.

7 min read
Authorised request path

Key takeaways

  • Data sovereignty is not just about where data is stored. It concerns who controls it, under what legal jurisdiction, and who can compel disclosure.
  • AI workloads extend the sovereignty perimeter: training datasets, inference prompts, and model outputs all carry sensitivity and must be treated as protected data.
  • The Privacy Act 1988 and sector-specific rules in finance, health, and government create real, enforceable obligations for where AI data is held and how it is processed.
  • The US CLOUD Act exposes data processed by US-owned platforms to compelled foreign disclosure regardless of where the physical servers are located.
  • Practical AI data sovereignty requires infrastructure due diligence, contractual controls, and a vendor operating entirely under Australian legal jurisdiction.

What is data sovereignty in an AI context

Data sovereignty refers to the principle that data is subject to the laws and governance structures of the country in which it is collected, stored, or processed.

In traditional IT environments, that concept was relatively contained. Data sat on servers in a known location, managed by a known operator. Sovereignty questions were largely settled by identifying where the data centre was and whether the operating entity was Australian.

AI changes the equation. The data pipeline is broader and less linear. Training a model involves ingesting large datasets, often containing sensitive or regulated information. Running inference means processing real user inputs: queries, documents, images, voice, in real time. The model itself retains statistical patterns derived from training data, patterns that can sometimes reconstruct or reveal source records. Outputs may include information traceable to individuals.

Each of these stages involves data in motion and data at rest. Each creates a potential sovereignty gap when infrastructure, the operating entity, or the control plane sits outside Australian jurisdiction.

Data sovereignty in an AI context means knowing where training data is processed, where inference runs, who controls the model, and what legal framework governs every stage of that pipeline.

Why AI workloads raise new sovereignty questions

Traditional workloads carried sovereignty risk, but the exposure was relatively bounded. A database had defined records. An application had defined data access paths. The scope of a breach or a compelled disclosure was traceable.

AI workloads are broader. Training data may include customer records, health datasets, financial histories, or government documents. A poorly governed training pipeline can expose regulated data to infrastructure or operators that fall outside the required legal framework. This can happen without a deliberate decision, simply by using a SaaS AI tool that routes data through offshore infrastructure.

Inference carries its own exposure. Users querying an AI system input real information: medical questions, legal documents, sensitive business data. If inference happens on offshore infrastructure, prompt content travels outside Australia. Depending on the operator’s jurisdiction, that content may be subject to foreign government access requests under laws such as the US CLOUD Act.

Model outputs are a further consideration. Large language models can reproduce fragments of training data, sometimes verbatim. If training included regulated personal information and the model is accessible outside the original compliance boundary, the result may constitute a data handling breach under Australian law.

None of these are theoretical risks. The Office of the Australian Information Commissioner has published guidance on AI and privacy. APRA has flagged AI risk in its supervisory letters to regulated entities. Australian businesses deploying AI carry responsibility for managing these exposures, not outsourcing that responsibility to a vendor.

Regulatory landscape: Privacy Act and sector-specific rules

Australia’s Privacy Act 1988 is the primary framework for personal information. Its Australian Privacy Principles (APPs) set requirements for how personal information is collected, held, used, and disclosed. APP 8 specifically addresses cross-border disclosure. Sending data overseas requires either that the recipient country has comparable privacy protections, or that the Australian entity accepts ongoing accountability for how that data is handled offshore.

For AI workloads, legal accountability does not transfer because a vendor’s terms of service include a privacy clause. If personal information is processed on offshore infrastructure, the Australian business remains accountable under the Privacy Act.

Sector-specific rules add further obligations.

Financial services. APRA’s CPS 234 requires APRA-regulated entities to maintain information security controls commensurate with the criticality and sensitivity of their data assets. Outsourcing to a third party does not reduce that obligation. Offshore AI processing of customer financial data sits within scope of CPS 234 and requires the regulated entity to conduct and document third-party risk assessments.

Health. The My Health Records Act and associated privacy guidelines govern the handling of health information. The Australian Digital Health Agency has published guidance noting that AI tools processing health records must meet the same compliance standards as any other health data handling system. The sensitivity classification does not change because the processing is automated.

Government and defence. Agencies handling data classified under the Protective Security Policy Framework (PSPF) and the Information Security Manual (ISM) must use infrastructure that meets defined sovereignty and access control requirements. Many classification levels cannot be processed on hyperscaler infrastructure with foreign ownership structures or offshore control planes.

Critical infrastructure. The Security of Critical Infrastructure Act 2018 (SOCI Act) creates obligations for operators of critical infrastructure assets across defined sectors, including data storage and processing. AI workloads that form part of critical infrastructure operations may be within scope of SOCI obligations.

Practical steps to ensure AI data sovereignty

Achieving data sovereignty for AI workloads is an operational and contractual task, not only a policy statement.

Data mapping. Before deploying any AI tool or building any AI pipeline, map the data that will flow through it. Identify what is regulated, what is sensitive, and where it originates. That mapping drives the sovereignty requirements the deployment must satisfy at every layer.

Vendor due diligence. Assess every vendor in the AI stack: infrastructure providers, model providers, API gateways, fine-tuning platforms, and any SaaS tools that process the data. For each, ask: is the vendor an Australian entity? What jurisdiction governs their operations and legal obligations? Do they have CLOUD Act exposure through their corporate structure? What happens to data after processing, during model training updates, and at contract end?

Contractual controls. Sovereignty cannot be satisfied by a vendor’s marketing materials or general privacy statements. Require contractual commitments on data residency, processing location, data retention and deletion timelines, your audit rights, and the vendor’s obligations when they receive a law enforcement request. If a vendor cannot provide those commitments in writing, they cannot meet Australian sovereignty requirements.

Operational monitoring. Sovereignty is not a set-and-forget configuration. Vendor contracts change. Data flows change as applications evolve. Build regular review processes to confirm that AI workloads continue to meet their sovereignty requirements as the technology stack changes over time.

How Amaze supports sovereign AI data handling

Amaze operates AI infrastructure as an Australian company, under Australian law, with no foreign parent entity and no CLOUD Act exposure. Training, inference, and data storage remain in Australia, in facilities located in Sydney (ap-syd-2) and Melbourne (ap-mel-1).

Amaze is ISO 27001 certified, providing an independently verified baseline for information security governance. The infrastructure is designed for regulated workloads: certified facilities, Australian-staffed operations, and contractual commitments that support the audit trail requirements of Privacy Act, APRA, and ISM-scoped workloads.

The model library, including NVIDIA-based inference for Meta Llama, Mistral, DeepSeek, and Qwen, is hosted and operated entirely within Australian jurisdiction. Data used for fine-tuning or inference does not leave the country. Pricing is AUD-denominated, with no egress surprises.

Data sovereignty for AI is the default position, not a premium tier.

Related reading: Data residency and compliance in the age of AI and How sovereign data centres support AI compliance.

Frequently asked questions

Is data sovereignty the same as data residency? No. Data residency is about where data is physically stored. Data sovereignty is broader: it covers who controls the data, under what legal jurisdiction, and who can compel its disclosure, which matters even when the storage location is correct.

Do AI prompts and outputs count as sovereign data, or just the training data? All of it. Training data, inference prompts, and model outputs each carry sensitivity and fall within the sovereignty perimeter. A model that runs on offshore infrastructure exposes prompt content to foreign jurisdiction even if the original dataset never left Australia.

Does the Privacy Act 1988 apply differently to AI systems than to other software? The obligations are the same, but the exposure is larger. AI systems ingest and process sensitive data at greater scale and speed, and outputs can sometimes reproduce fragments of training data, which increases the chance of an inadvertent breach under the Australian Privacy Principles.

What is the first practical step toward AI data sovereignty? Data mapping: identifying what data will flow through an AI system, what is regulated or sensitive, and where it originates, before deployment. That mapping determines the sovereignty requirements every vendor and processing layer must satisfy.

Tagged data sovereigntyAI compliancePrivacy ActCLOUD Actregulated workloads

Build on sovereign Australian infrastructure.

Talk to a solution architect about deploying your workload on Amaze.