Seven practices, one engineering team. We design, build, benchmark and operate the compute layer beneath your hardest problems — on-premise, hybrid or sovereign cloud.
Not sure where your workload fits? Our engineers will benchmark it before you commit.
Schedule a discovery call
Industries
Bespoke architecture, sector by sector.
A genomics pipeline and a trading engine are not the same machine. We start from your regulatory perimeter, your data gravity and your deadlines — then design backwards.
We build the compute layer Australia keeps assuming it already has.
Cloud Natives designs, builds and operates sovereign AI, HPC, storage and
low-latency infrastructure for Australian research, government, defence and
industry. Every claim on this page is meant to be checkable. Where a figure is
still unverified we have said so in the markup rather than quietly rounding it.
It usually arrives as a procurement clause — data stays onshore, the supplier is
local — and then nobody specifies the fabric, the filesystem, the support path or
the key custody that would make the clause true.
The constraint we design around first
Where the data already sits, and what it costs to move it. Nearly every estate we
are asked to build is really a data-gravity problem wearing a compute badge. A
telescope archive, a genomics cohort, a decade of tick data or a classified
imagery holding cannot be shifted twice for convenience. The machine goes to the
data, the residency boundary follows the data, and the network bill is decided
long before anyone picks a GPU.
Australia funds ambitious science and ambitious defence programmes, then rents the
layer underneath them by the hour inside someone else's control plane. That is a
reasonable trade for bursty work. It stops being reasonable when the workload is
classified, when the dataset is too large to egress, or when a three-year training
run meets a capital budget.
So we build that layer in-country: the cluster, the interconnect, the parallel
filesystem, the scheduler policy, the key hierarchy, the change process and the
people who can be phoned at three in the morning. None of it is exotic. It is
unglamorous, it is chronically under-resourced, and it is the part that decides
whether a research programme lands or slips a year.
We are deliberately small. That caps how many estates we can hold at once, and it
is the reason the engineer who chose your blocking ratio is still the one who owns
it two years later. When we are at capacity we say so rather than hiring a bench
to absorb the work.
A data residency clause is not a design. Someone still has to decide where the
metadata lives.
The position we open most sovereignty conversations with. Residency, jurisdiction
and key custody are three separate decisions, and a tender response that treats
them as one has not been engineered yet.
Where sovereignty actually leaks. Backup targets in another region.
A KMS root held by the vendor. Telemetry and crash dumps egressing to a support
portal. A vendor support engineer with remote hands and a foreign passport. Object
metadata replicated for durability. Each one is fixable, and none of them are fixed
by the word "sovereign" in a contract.
Three things we are not
Not a reseller
We sell engineering time and an operated outcome. Hardware passes through because
it has to, not because the margin on it funds the business. Where a component
carries a partner margin we name it in the bill of materials.
Not a consultancy
We do not write a strategy and leave. Every engagement ends with something that
boots, passes an acceptance test written in advance, and has a named owner on an
escalation roster.
Not a hyperscaler
We cannot match on-demand elasticity and we do not pretend to. For spiky,
short-lived or globally distributed workloads the public cloud is usually the
right answer, and we will model it honestly alongside our own.
Operating principles
Five commitments you can hold us to.
Not values on a wall. Each one is checkable during an engagement, each one has a
cost we absorb, and we have lost work over every single one of them.
01
We benchmark before we quote
Send us the workload, not the requirement. An input deck, a packet capture, a
model checkpoint, a day of market data, a Nextflow run — whatever the real thing
is. We profile it, we find the term that dominates elapsed time, and the quote
arrives with the measurement attached. If the workload cannot leave your
environment we quote a benchmarking engagement first and nothing else, and we will
run it on your hardware under your supervision.
02
We publish the method with the number
Every figure we state carries how it was produced: the tool and version, build
flags, node and rank count, dataset shape, the run-to-run spread, and the runs we
discarded and why. Where a result is a p50 we say p50. Where it is a best case we
say best case. We also write up the engagements where our own design lost to a
cheaper one, because a vendor who has never been beaten is a vendor who has not
measured.
03
No reference architecture sold as bespoke
Vendor reference designs are useful and we read all of them. We will not hand one
back with a cover page and a design fee. If your configuration genuinely is a
standard build — and often it should be, because standard builds have firmware
that has been tested — we tell you that and price it as integration work. The
bespoke engineering goes where your constraint is unusual, which is rarely the
node and frequently the data path.
04
The engineer who built it answers the phone
Design, build and operations sit with the same small team. The person who chose
your oversubscription ratio is named on the escalation roster for it, and that name
is in the contract. The cost of this is capacity: we cannot take unlimited
engagements, and we would rather close intake for a quarter than break the link
between the design and the on-call roster.
05
We will tell you when we are the wrong team
Some workloads belong on a hyperscaler. Some belong on a national research
facility where you already hold an allocation and are not using it. Some need a
software engineer for six weeks, not a cluster — a single-threaded I/O loop is
cheaper to fix than to out-build. When the measurement points somewhere other than
us we write that down with the numbers, hand it over, and bill only the
benchmarking effort we agreed in advance.
How to test this in a procurement. Ask for the method behind any number
in our response. Ask which engagement we lost on price in the last year and why. Ask for
the name that will be on the escalation roster. If a supplier cannot answer all three in
writing, discount their figures accordingly — including ours.
Team
The people who own the build.
Six practice owners. Each one is accountable for a layer of the stack end to end,
from the specification through acceptance testing to the on-call rotation for the
systems they designed.
Placeholder team — client content required. The six profiles below are
invented to establish layout, role coverage and tone of voice. Before launch we need,
for each person: legal name, role title, a two-sentence biography you are happy to be
held to, a portrait at 800 × 1000 px or larger, and a LinkedIn URL. Tell us which
individuals hold a security clearance and at what level, because defence and government
buyers will ask and that detail is not safe to guess.
Spent eleven years specifying research and trading infrastructure for other people
before starting Cloud Natives to do the measurement work properly. Still writes the
acceptance test plan personally for every engagement above a set contract value.
Owns cluster design and Slurm policy, from fair-share weighting and QoS tiers through
to topology-aware placement and backfill behaviour. Has spent more of his career on
queue wait than on FLOP rates, which is roughly the correct ratio.
Designs GPU fabric and the inference path: vLLM serving topology, KV-cache
behaviour, quantisation trade-offs and GPUDirect data movement. Argues that most
training clusters are storage projects wearing an accelerator badge, and is usually
right.
Leads IRAP engagement, Essential Eight uplift and the key-custody design that turns a
residency claim into something demonstrable. Writes the system security plan before
the rack elevations, not after the assessor asks for it.
Designs the tiering from NVMe scratch and NVMe-oF through Lustre or BeeGFS to object
and tape, plus the metadata layer that decides whether any of it is usable.
Benchmarks small-file and metadata operations first, because that is where estates
actually fail.
Runs the in-country operations roster, the on-call rotation and the post-incident
review process. Holds the rule that no change ships without a documented rollback, a
named owner and a stated blast radius.
Partnerships
We hold vendor accreditations. We do not let them pick the design.
A partner tier buys firmware access, an engineering escalation path and spares
logistics in-country. It should never be the reason a component appears in your bill
of materials.
What a partnership is worth
Real engineering value, and it is specific.
Firmware and errata ahead of general release, with the known-issue list
Joint escalation into the vendor's own engineers, not a reseller queue
Evaluation hardware early enough to benchmark before you commit
Spares held onshore, with a stated replacement window
What it must never buy
The specification. Rebate structures move on a quarterly cycle and estates live for
five to seven years, so a design tuned to this quarter's incentive is a design that
ages badly. We keep three rules.
Where a component carries partner margin for us, it is named in the BOM
Single-source choices get a written justification you can challenge
No accreditation target is allowed to set a purchase volume
How we keep candidates comparable
One harness, one dataset, one set of build flags, run across every candidate. The
comparison goes to you with the raw output, including the configurations we could
not make work and what we think that means. Vendor-supplied numbers are recorded as
vendor-supplied and are never mixed into our own results.
NVIDIA
AMD
Intel
Supermicro
Dell Technologies
Lenovo
DDN
WEKA
VAST Data
Arista
Juniper
Red Hat
No vendor logo ships without a partner agreement. The wall above is a
layout placeholder. Send us the confirmed partner list, the tier held with each vendor,
the approved monochrome SVG and any mandatory attribution wording from their brand
guidelines. Anything we cannot evidence gets removed from the wall rather than softened
in the caption.
Certifications & governance
What we hold, what is in progress, what we will not claim.
An accreditation without a certificate number, a scope statement and an expiry date
is a logo, not a control. Treat every row below as unproven until the evidence
column is filled in.
Accreditation register — placeholder status, pending verification by Cloud Natives
Accreditation
Scope as stated
Status
Evidence we must supply
IRAP assessment
Sovereign managed platforms and the NOC that operates them, to PROTECTED
Information security management across engineering delivery and operations
To confirm
Certificate number, certifying body, scope statement, current Statement of Applicability
ISO 9001
Quality management for design, build and acceptance testing
To confirm
Certificate number, last surveillance audit date, open non-conformances
Defence Industry Security Program
Membership category and the facilities and personnel it covers
To confirm
Membership confirmation, entry level for each of the four categories, sponsor
Essential Eight
Our own corporate and management environments, not the client estate
To confirm
Target maturity level, whether the assessment is self-assessed or independent, date
Personnel clearances
Which roles hold Baseline, NV1 or NV2, and which engagements require them
To confirm
Count by level and sponsoring entity. Individual names are never published
Insurance
Professional indemnity, public liability and cyber liability
To confirm
Insurer, policy numbers, limits and expiry — buyers request these at tender stage
Procurement and reporting posture
Modern slavery
Hardware supply chains run through contract manufacturing and mineral extraction,
so this is a live risk for us rather than a paperwork exercise. We intend to
publish a statement under the Modern Slavery Act 2018 (Cth) once consolidated
revenue passes the reporting threshold, and to operate a supplier code and a
component-level questionnaire regardless of whether the threshold applies.
Read the current statement.
Environmental
We report facility PUE per engagement, and WUE and grid carbon intensity where the
data centre provides them. We will quote a design that is worse on power if it is
better on time-to-result, and we will say which trade you are making. We do not
present purchased offsets as an efficiency improvement.
Indigenous procurement
Work delivered under Commonwealth contracts is aligned to the Indigenous
Procurement Policy. We maintain a register of Indigenous-owned suppliers for
logistics, fit-out, cabling and facilities work, and report spend against contract
targets where a contract sets them.
Accessibility
This site targets WCAG 2.2 level AA, and the dashboards and operator consoles we
build are held to the same standard. If something here fails for you, tell us and
we will fix it. Accessibility statement.
Questions procurement asks us first
Ownership and control are in the fact list at the top of this page and must be
confirmed against the current company register before you rely on them. For
sovereign engagements we accept a change-of-control notification clause with a
defined notice period and a termination right, because a change in beneficial
ownership is exactly the event a residency requirement exists to survive.
We subcontract physical work — cabling, fit-out, freight and some facilities
trades — and we name those parties before award. Engineering, operations and
incident response stay with our own staff in Australia. Vendor support is the
honest exception: a firmware defect can require a vendor engineer offshore, and
we design the access path for that case in advance, with session recording,
time-boxed credentials and a nominated Australian escort on the call.
Everything we build is documented to the standard another team could take over:
as-built topology, configuration in version control you hold, runbooks, the
acceptance test suite and the monitoring definitions. We use open orchestration
and open filesystems by default — Slurm, Kubernetes, Lustre, BeeGFS, Ceph — so
the exit path does not depend on us existing. Ask for the handover pack during
evaluation rather than at termination.
A single workstation or a two-node build is worth a conversation and is often the
right first step for a research group or an early-stage team. What we decline is
work where nobody will own the outcome operationally, or where we are being asked
to supply a quote to make someone else's tender look competitive. We would rather
say no in week one than deliver something with no owner.
Careers
We hire slowly, and we are direct about the trade.
There is no graduate programme and no bench. If you want breadth across silicon,
fabric, filesystem and scheduler, this is a good place to work. If you want to
specialise very narrowly at enormous scale, a hyperscaler will serve you better.
What the work actually is
Profiling somebody else's code before you are allowed to recommend hardware
Writing the acceptance test, then being held to it in front of the client
Data centre time: racking, cabling, labelling, and the long boring bring-up
On-call for the estates you designed, compensated and rostered, not volunteered
Writing up the failures publicly, including your own
What it is not
Not a pre-sales role dressed as engineering — the measurement decides the quote
Not fully remote. Fabrics and filesystems need hands, and clients need faces
Not a place to avoid documentation. If it is not written down it is not finished
Not always cleared work, but some of it is, and that constrains who can staff it
Open roles and how to apply. We post roles only when the work is signed,
so the list is short and sometimes empty. Where an engagement requires a security
clearance we sponsor it for staff who are eligible, and we will tell you at first
interview whether a role depends on one. Send a short note about a system you built and
what it taught you, to
hello@cloudnatives.example.
A curriculum vitae is welcome but the note matters more.
Ask us something specific
Ask for the certificate number.
Then ask for the method.
We would rather answer a hard due-diligence question early than win a tender we
cannot evidence. Bring the security questionnaire, the residency constraint or the
awkward architecture review — whichever one is currently blocking you.