Seven practices, one engineering team. We design, build, benchmark and operate the compute layer beneath your hardest problems — on-premise, hybrid or sovereign cloud.
Not sure where your workload fits? Our engineers will benchmark it before you commit.
Schedule a discovery call
Industries
Bespoke architecture, sector by sector.
A genomics pipeline and a trading engine are not the same machine. We start from your regulatory perimeter, your data gravity and your deadlines — then design backwards.
Buy the decision you can defend in eighteen months.
There are two expensive mistakes available to you. Buying hardware before you know
your load, and renting a workload you have understood for two years because nobody
did the arithmetic. Both cost roughly the same. We help you pick the cheaper one,
with the model written down so the next round of the argument starts from numbers.
The constraint we design around first
Optionality has a price. So does keeping it forever.
Early on you pay a premium to stay flexible, and that premium is worth it because
you do not yet know what you are building. The failure is not noticing when the
uncertainty resolves and you are still paying for it.
The two failure modes, in the order they usually arrive
Both are common enough to describe without naming anyone. They look different
from the inside and identical on a cap table.
01
Premature capex
A seed-stage team buys a GPU server because the on-demand bill looked
alarming. Six months later the model architecture has changed, the card is
the wrong shape, and the capital is in a rack rather than in runway. The
hardware was not the mistake. Committing capital against an unknown load was.
02
The Series B rebuild
A platform assembled at speed is now carrying real revenue. Cost per customer
is rising rather than falling, nobody can say which workload consumes what,
and the architecture assumes a single managed service per function. The
rebuild takes two quarters of engineering attention at exactly the point
where that attention was promised to product.
03
The version nobody talks about
Over-engineering for portability from day one. Two of every abstraction, a
service mesh for four services, and a platform team of one who is now the
only person who can deploy. The optionality is real and the cost is the
product you did not ship.
The useful work is deciding which parts of your load are now predictable, and moving
only those. Everything else stays rented, deliberately, with a review date.
Chart — cost per unit of work over time for rented, reserved and owned capacity, with the crossover band marked, 1200×900
What we produce first. A decision document: your workloads
classified by predictability, a cost model with the assumptions visible and
editable, the crossover point for each workload, and a recommendation that may well
be to change nothing. It is useful to your board whether or not you engage us
further.
Cloud versus owned
Six variables decide this. The list price is not one of them.
Comparisons that put an instance-hour rate beside a server quote are theatre. The
honest model includes the people you would have to hire, the capital you would tie
up, and the fact that half your load is genuinely unpredictable.
01
Utilisation, measured rather than assumed
Owned hardware is priced per month regardless of use, so its economics are
entirely a function of how busy it is. Below roughly half utilisation, owning is
usually indefensible. Above sustained high utilisation on a workload whose shape
is stable, renting is usually a premium you no longer need to pay. The interesting
cases sit in between, which is why we insist on measurement instead of a
spreadsheet estimate.
02
Data egress and gravity
Egress charges are a small line until your product starts serving media, exports
or model weights. More importantly, data accumulates somewhere and then pulls
compute towards it. If your training corpus lives in one provider's object store
and grows weekly, the cost of moving later rises every month you defer the
decision. We quantify that as a number, not a warning.
03
Steady state versus bursty
Almost every startup has both. Inference serving a live product is steady and
predictable within a band. Training runs, batch reprocessing and customer
onboarding jobs are spiky by nature. The correct answer is nearly always a split:
own or reserve the steady floor, rent the spikes. Sizing owned capacity for your
peak is how facilities end up at 20% utilisation.
04
Cost of capital, and whose capital it is
Venture equity is the most expensive money you will ever spend on a depreciating
asset. Before recommending purchase we look at whether equipment finance, an
operating lease or a vendor arrangement changes the answer, because converting
capex to a monthly obligation often preserves the runway that matters more than
the total.
05
The team you would need to hire
Owning infrastructure means someone carries a pager, patches hypervisors, handles
a failed drive at 2 a.m. and keeps the capacity model current. That is a real
salary, plus a second one so the first can take leave. If you do not intend to
hire it, the honest options are a managed arrangement where we carry it, or
staying in cloud. Pretending it is free is the most common error in these models.
06
Time to capacity
An instance is available in a minute. Hardware has a lead time, a rack, a power
allocation and a commissioning window. If your growth curve means you need double
the capacity inside a quarter, that lead time is a business risk and belongs in
the comparison alongside the price.
The most common outcome of this analysis is: stay where you are, fix three
things. Right-size what you run, take a commitment on the predictable portion,
and instrument cost per unit of work so the question can be reopened with evidence in
two quarters. We charge for the analysis, not for an outcome that favours us.
Proof of concept to production
Four stages, each with an exit test.
The point of staging is that each step is cheap to abandon. If a stage cannot be
abandoned without losing the work, it was scoped wrong. Every transition has a
written test, so nobody has to argue about whether you are ready.
Proof of concept — weeks, rented entirely
On-demand everything, including the expensive accelerators. You are buying
information about your own workload, not capacity. The exit test is narrow: can we
state the resource profile of one unit of work, and does the approach do what the
product needs. Do not optimise cost here. Optimise learning rate.
Pilot — first real users, first real telemetry
Production-shaped but small: a managed Kubernetes cluster or a handful of
instances, one Postgres you did not build yourself, object storage, and
OpenTelemetry instrumentation from the first commit. The exit test is a measured
cost per customer or per thousand requests, with a stated growth assumption. If
you cannot produce that number, you are not ready for the next stage regardless of
how the product is performing.
Production — commit to the predictable part
Now the split becomes worth engineering. Reserved or committed capacity for the
steady floor, autoscaling for the rest, and a clear separation between the
workloads you understand and the ones you do not. This is also where availability
targets stop being aspirational and get written down with the consequences of
missing them.
Scale — own the floor, if the numbers say so
Colocated or owned capacity for the workload whose shape has been stable for two
or more quarters, sized to the floor rather than the peak, with burst retained in
cloud. The test for entering this stage is that you can predict next quarter's
base load within about 20% and you have the operational cover, ours or yours, to
run it.
2 quartersMinimum stable load history before we recommend owning anything
±20%Forecast accuracy on base load we treat as the threshold for purchase
Floor onlyWhat owned capacity is ever sized against — the peak stays rented
GPU access for AI startups
Four ways to get accelerators. Each one expires.
The interesting question is not which is cheapest. It is when each one stops being
the right answer, because the signal that you have outgrown an arrangement usually
arrives a quarter after the cost did.
Indicative guidance on GPU access modes — replace with a current model for your workload and market
Mode
Suits
Relative cost per GPU hour
When it stops making sense
On-demand cloud
Experimentation, unknown architecture, short training runs, bursts
Highest
When the same job shape has run weekly for a quarter, or when you start
queueing for availability in your region
Reserved or committed
Steady inference, recurring training cadence, predictable base load
Materially below on-demand for the committed portion
When the commitment term outlasts your confidence in the hardware
generation, or when you are consistently exceeding the reservation
Colocated, hardware owned or financed
Sustained high utilisation, data residency requirements, latency to your
own data
Lowest per hour once utilisation is high, plus power, space and cross-connect
When utilisation falls below the level that justified it, or when a new
accelerator generation changes your cost per token faster than you can
depreciate
Owned, on your premises
Development and lookdev-style interactive work, sensitive data, air-gapped
requirements
Low per hour, high in facilities and attention
Almost immediately, at density. Office power and cooling become the binding
constraint well before the compute does
Depreciation is the argument people skip
Accelerator generations arrive faster than a comfortable depreciation schedule.
Buying a card outright is a bet that your cost per unit of work will still be
competitive in three years against hardware that does not exist yet. Sometimes
that bet is fine, particularly where data cannot leave the country or the
building. It should be made explicitly rather than inherited from a spreadsheet
that assumed a five-year life.
Serving efficiency usually beats buying more
Before adding accelerators we look at what you are getting from the ones you
have. Continuous batching and paged attention in a server such as vLLM,
quantisation where evaluation shows quality holds, KV-cache reuse, and routing
easy requests to a smaller model. On inference workloads these changes commonly
move cost per million tokens further than a hardware upgrade would, and they cost
engineering weeks rather than capital.
Diagram — GPU access modes plotted against utilisation and commitment length, with the practical switching points, 1200×900
Portability and instrumentation
Keep the exit cheap. Do not build the exit.
Full portability is an expensive insurance policy against an event that may never
occur. The practical target is that leaving would be a project, not a rewrite, and
that you know today what that project would cost.
Where we accept coupling, and where we do not
Managed services are usually worth the lock-in when they replace work you would
otherwise have to staff. The place to hold the line is anywhere that would force
a rewrite of your own code rather than a change of configuration.
Accept
Managed relational databases, managed queues, identity, secrets management, and observability back ends. Replacing these later is a migration with a known shape.
Accept, with care
Provider-specific serverless runtimes for edges of the system. Keep the business logic in ordinary libraries so the trigger is the only thing that is provider-shaped.
Hold the line
Storage access patterns. An S3-compatible interface means a move is a credential and endpoint change rather than a rewrite, and Ceph or MinIO give you the same API on your own hardware.
Hold the line
Container packaging and orchestration. If it runs under Kubernetes with declarative manifests, it will run somewhere else. If it only runs through a console click, it will not.
Hold the line
Infrastructure as code and reproducible builds. Terraform or equivalent, in version control, with no undocumented manual state. This is also the artefact diligence will ask to see.
Do not bother
Abstraction layers that let you swap providers in theory. They cost real velocity and are almost never exercised. Write the migration plan instead, and price it once a year.
Instrument now, so the next decision has data
Every decision on this page depends on measurements that are cheap to collect
early and awkward to reconstruct later. Six things, set up once.
Cost allocation tags from the first resource, so spend can be attributed to a product, a customer tier and an environment without archaeology
A defined unit of work — a request, an inference, a document processed — and its cost tracked as a first-class metric
Utilisation history at fine enough resolution to see the shape of your peaks, not just monthly averages
Egress and inter-zone traffic broken out separately, because these are the lines that surprise people at scale
Data growth rate per dataset, which is what determines when gravity starts making decisions for you
OpenTelemetry traces on the request path, so a latency regression is a query rather than a week of guessing
The cheapest version of this is a day of work. Tags, one dashboard,
one recurring export. Teams that do it at pilot stage can answer an investor's cost
question in an afternoon. Teams that do not are usually rebuilding history from
invoices while under time pressure.
Diligence and enterprise readiness
Your first enterprise customer is a security audit with revenue attached.
Investors ask about unit economics and concentration risk. Enterprise buyers ask
about data handling, access control and what happens when you are acquired. Both
conversations go faster when the answers already exist as documents.
Security posture proportionate to your stage
The Essential Eight is a sensible frame for an Australian company and is cheap to
adopt early: patching, multi-factor authentication, application control, admin
privilege restriction and tested backups. Full ISO/IEC 27001 or SOC 2 certification
is a deliberate commercial decision, usually driven by a specific deal. We will tell
you when you do not need it yet.
Data handling and residency, written down
Where customer data lives, which processors touch it, how long it is retained, how a
deletion request is executed, and what crosses a border. Australian Privacy
Principles apply to most of this. Government and health buyers will ask about
residency specifically, and an answer that consists of a shrug ends the
conversation.
Unit economics that survive a hostile read
Gross margin per customer with infrastructure attributed honestly, including the
free tier, the support overhead and the workloads that run whether or not anyone is
using the product. Investors discount numbers they cannot trace. A cost model with
visible assumptions is worth more than a flattering one.
Key-person and single-point risk
A diligence process will find the service only one engineer understands, the
hand-configured host that is not in code, and the credential in someone's personal
vault. These are findings that cost negotiating position. They are also
straightforward to clear if you start before the term sheet.
Essential Eight uplift as a starting pointAustralian data residency by defaultInfrastructure as code, handed overCost model you own and can edit
When not to engage us
Four situations where we are the wrong purchase.
We would rather lose the engagement than take money for infrastructure a company
does not need yet. It is also a better commercial position: the same founders come
back when the load is real.
01
You have not found product-market fit
If the shape of the product is still moving, infrastructure spend is a distraction
and hardware is a liability. Use on-demand, accept the premium, and put the
capital into finding out what you are building. Come back when a workload has
repeated often enough to have a shape.
02
Your monthly infrastructure bill is small
Below a few thousand dollars a month, the achievable saving does not cover the
cost of the analysis, let alone a migration. Right-size what you have, take the
obvious commitment discount, and spend the attention on your product. We will say
this on the first call rather than the third.
03
Nobody will own the platform after we leave
Owned infrastructure needs a person, or a managed arrangement where that person is
ours. A team with no platform capacity and no budget for operational cover should
not take on hardware. That is not a sales objection to overcome; it is a
prediction about an outage in eight months.
04
The real problem is application efficiency
Sometimes the bill is high because of an N+1 query, an unbounded retry, a nightly
full table scan or a model that is twice the size the task requires. Buying
capacity to cover that is the most expensive available fix. We will point at it,
and you may not need us further.
What we do charge for. The analysis, the architecture, the build, and
the operational cover afterwards. Not for an outcome. If the model says stay in cloud,
that is the deliverable, and it is the same price.
Let's Talk
Bring your last three invoices and your growth assumption.
We will build the model with you.
Three months of billing detail, a utilisation export and an honest forecast are
enough to find the crossover point for each of your workloads. You keep the model.
If it says change nothing, that is a legitimate result and we will tell you plainly.