Amaze Contact →
Sovereign Cloud

How an edge rollout rewrites storage strategy

Edge does not shrink what you store. It changes where data lands and what your five-year storage budget has to model.

9 min read

Key takeaways

  • Gartner projected that by 2025 around 75 per cent of enterprise-generated data would be created and processed outside a traditional centralised data centre or cloud, up from under 10 per cent in 2019.
  • Edge changes the direction of traffic. Your storage strategy stops being a capacity question and becomes a movement and retention question.
  • Public cloud egress is priced to punish the return trip: after 100 GB free per month, AWS lists US$0.09/GB for the first 10 TB, and Asia Pacific regions such as Singapore start at US$0.12/GB.
  • Hyperscalers stretched server useful life from three or four years to six, and in 2025 Amazon pulled a subset back to five. Your own refresh assumptions deserve the same scrutiny.
  • Model five years as three separate lines: capacity growth, replication and movement, and refresh timing at each tier.

An edge rollout usually starts as a latency project. Put inference next to the camera, the sensor, the point of sale or the plant floor, cut the round trip, and get a decision back in tens of milliseconds instead of hundreds. That part works.

What it quietly does is invalidate the assumption underneath your centralised storage strategy: that data is created near the core, stored near the core, and read from near the core. Gartner projected that by 2025 roughly 75 per cent of enterprise-generated data would be created and processed outside a traditional centralised data centre or cloud, up from less than 10 per cent in 2019. Once creation moves out, everything downstream of creation moves with it.

Edge does not reduce what you store

The first budget mistake is treating edge as a filter. The pitch is that you process locally and only send the interesting parts upstream, so central capacity growth flattens.

In practice it rarely flattens, for three reasons.

You keep the raw data anyway. Model retraining needs it. Incident investigation needs it. In regulated settings, evidentiary and audit obligations need it, and a derived summary is not the record. So the raw stream is retained somewhere, and “somewhere” is now dozens or hundreds of locations rather than one.

You add new data. Edge inference produces its own output: predictions, confidence scores, model versions, feature snapshots. That output is small per event and large per year, and it is the data your auditors will ask about, so it gets a longer retention class than the raw input.

You duplicate for availability. An edge node with no local resilience is an outage waiting for a truck roll. So each site gets a local copy plus a central copy, and often a third copy for backup. Your total stored bytes go up even when your central array does not.

Budget the aggregate, not the core. If your capacity model only counts the racks you can see, it will be wrong from month one.

Data gravity now pulls in the wrong direction

Data gravity is the observation that data attracts the compute and applications that use it, and that moving large volumes is slow and expensive enough that the workload usually moves instead. Centralised architecture used that in your favour. Everything landed in one place, so everything ran in one place.

Edge inverts it. Gravity now exists at every site. The consequence is that “just replicate it all to the centre nightly” stops being a background task and becomes the dominant cost and the dominant failure mode.

The shift shows up in five places a storage plan already tracks.

DimensionCentralisedEdge-distributed
Where data landsCreated, stored and read near the coreDozens or hundreds of sites, plus a central copy and often a backup copy
Direction of movementInbound to one placeOutbound from every site to the centre
Egress exposureOne consolidation pointPer-site replication metered per GB at each hop
Refresh cycleCentral arrays on a procurement cycleDesynchronised, a rolling percentage of the fleet with truck-roll cost
Budget line most affectedCapacity growthMovement, scaling with site count and per-site sensor density together

The design question is no longer how much storage you need. It is which class of data has to move, how fast, and in which direction:

  • Must move, in near real time. Alerts, exceptions, aggregated telemetry, anything feeding a central dashboard or a regulatory report.
  • Must move, eventually. Training corpora, long-term archives, anything supporting a model refresh cycle rather than an operational one.
  • Should never move. High-volume raw capture whose value expires in days, and anything whose movement creates a compliance exposure you would rather not carry.

Sorting your data into those three buckets does more for a five-year budget than any procurement negotiation, because it directly sets the line item most people underestimate.

The three cost lines that move

Egress. Public cloud pricing is built around the assumption that data comes in free and leaves at a price. AWS publishes 100 GB per month free, then US$0.09 per GB for the first 10 TB, US$0.085 for the next 40 TB, US$0.07 for the next 100 TB, and US$0.05 above 150 TB, at standard US and European rates. Asia Pacific regions cost more, with Singapore starting at US$0.12 per GB. On those published rates, 10 TB a month of internet egress works out to roughly US$913, and 100 TB a month to roughly US$7,980, before the meters that sit alongside it: NAT gateway processing at US$0.045 per GB, cross-availability-zone traffic at US$0.01 per GB in each direction, and cross-region transfer at US$0.01 to US$0.02 per GB.

An edge fleet generates exactly the traffic pattern those meters are priced for. Run the arithmetic on your own daily replication volume across the whole fleet, multiply by 60 months, and put the number in the business case before someone else finds it in year three.

Regulators have noticed the lock-in effect. Under the EU Data Act, switching charges between data processing services are prohibited outright from 12 January 2027. That covers switching rather than day-to-day traffic, and it is European law, not Australian, but it signals where scrutiny of exit costs is heading.

Replication. Each additional copy multiplies capacity, network and management cost together. Decide the copy count per data class, not per system. Three copies of a raw video stream is a policy failure, not a resilience posture.

Physical siting. Central capacity is no longer freely available. CBRE reported Asia Pacific data centre availability down 43 per cent year on year in the first quarter of 2026, with Sydney vacancy at 4.5 per cent and asking rents around US$188 per kW per month. If your plan assumes you can take another 200 kW in Sydney in year three at today’s rate, price the risk that you cannot.

What the five-year model actually needs

Split it into three lines rather than one growth curve.

Capacity by tier and by location. Aggregate edge, regional and central separately, each with its own growth rate. Edge capacity tends to grow in steps as sites are added. Central capacity grows smoothly with retention. Combining them into one number hides both.

Movement. Gigabytes per day per site, times site count, times the tariff at each hop. This is the line that behaves nonlinearly, because it scales with sites and with per-site sensor density at the same time.

Refresh. This is where most five-year models are quietly wrong. The hyperscalers extended assumed server useful life from three or four years out to six, which changed their reported economics considerably. Then in 2025 Amazon shortened the useful life of a subset of its servers from six years back to five, citing the increased pace of technology development. If the operators with the best fleet data in the world are revising their assumptions in both directions, a fixed five-year refresh line in your spreadsheet is an assumption, not a plan.

Edge makes refresh harder in a specific way: it desynchronises it. Central arrays get refreshed on a procurement cycle. Edge nodes get refreshed when a site is upgraded, when a sensor generation changes, or when a node fails in a place that is expensive to reach. Model edge refresh as a rolling percentage of the fleet per year with a truck-roll cost attached, not as a single event in year five.

Density helps here. Capacity per rack unit has moved fast enough that a refresh can be a consolidation rather than an expansion. A single 122 TB QLC drive rated at 25 W maximum and under 5 W idle changes what a small edge chassis can hold, and what a central tier costs to power. Recheck your assumed terabytes per rack unit each budget cycle rather than carrying last cycle’s figure forward.

The sovereignty variable

Every replication decision is also a jurisdiction decision. When you decide that edge sites in five states replicate to a central tier, you have decided where the consolidated record lives and which law reaches it. When a managed observability tool ships logs off the node by default, that decision was made for you.

Amaze operates Australian regions in Sydney (ap-syd-2) and Melbourne (ap-mel-1) as an Australian company under Australian law, with no foreign parent entity and no CLOUD Act exposure, ISO 27001 certified, and AUD-denominated pricing. For an edge architecture with many sites and one consolidation point, the practical effect is that the consolidation point has one jurisdiction and one currency, and the cost of moving data between your own sites is something you can model rather than discover.

Related reading: Data residency and compliance in the age of AI and Cloud governance for AI: guardrails for responsible use.

Frequently asked questions

Does edge computing reduce total storage cost? Usually not on its own. It reduces the volume that has to travel and the latency of a decision, but total stored bytes generally rise because raw data is retained at the site, inference output is added, and each site needs local resilience. The saving comes from deciding what never moves, not from storing less.

How do I estimate egress before the edge fleet is built? Instrument one representative site for a month and measure gigabytes leaving it per day by data class. Multiply by planned site count and by the published per-GB rate at each hop, then add the adjacent meters such as NAT processing and cross-zone traffic. A single measured site beats a modelled estimate for the whole fleet.

What should a five-year infrastructure budget assume for refresh? Assume different lives for different tiers, and revisit them annually. Central storage on a longer cycle, edge nodes on a rolling replacement percentage with field service cost attached. The industry has revised its useful-life assumptions in both directions in recent years, so treat any fixed number as a variable under review.

Where does data gravity hurt most in an edge architecture? At the consolidation point. Once a central tier holds the joined, cleaned, historical dataset, retraining, analytics and reporting all gravitate to it, and moving that tier later becomes the most expensive change in the architecture. Choose its location and jurisdiction deliberately at the start.

Tagged storage strategydata gravityegressbudgetingedge computing

Build on sovereign Australian infrastructure.

Talk to a solution architect about deploying your workload on Amaze.