Start with the shape of the bill, not the model
Before touching a single instance type, split your ML spend into four buckets. In almost every account we review the split looks roughly the same and the surprise is always in the last bucket.
| Bucket | Typical services | What usually goes wrong |
|---|---|---|
| Training | SageMaker Training, EC2 P/G/Trn instances, EKS GPU nodes | On-Demand GPUs for jobs that could checkpoint; oversized instance families; failed jobs re-run from scratch. |
| Inference | SageMaker Endpoints, EKS, Lambda, Bedrock | Endpoints sized for peak and never scaled down; one model per endpoint; no autoscaling in non-production. |
| Data & storage | S3, EBS, FSx for Lustre, EFS, data transfer | Duplicate datasets, hot storage for cold data, FSx file systems left running after training, cross-region egress. |
| Everything else | Notebooks, Studio, dev clusters, logs, experiment tracking | Notebook instances running 24/7, GPU dev boxes “for later”, CloudWatch ingestion of debug logs. |
Tag ruthlessly: ml:project, ml:stage (train / infer / dev) and ml:owner on every resource. Without the stage tag you cannot tell whether your GPU spend is producing models or producing heat.
Training: pay for compute only while it computes
Use Spot with checkpointing, always
Training is the ideal Spot workload because it is interruptible by design if you checkpoint. SageMaker Managed Spot Training handles the interruption and restart for you and regularly delivers savings of well over half versus On-Demand. On EKS or raw EC2, write checkpoints to S3 every N steps and let the job resume from the last one. The rule is simple: if a training job cannot survive an interruption, it is a reliability bug, not a reason to pay On-Demand.
Pick the smallest accelerator that keeps the GPU busy
Teams default to the biggest P-family instance because it feels safe. Look at GPU utilisation first. If a P4d or P5 spends most of its time waiting on the data loader, a G5 or G6 instance with faster storage would finish in similar wall-clock time at a fraction of the price. Profile before you provision.
Evaluate AWS Trainium for recurring jobs
For workloads you retrain on a schedule, Trn1 instances are frequently the cheapest way to get a given number of training tokens through the door. The porting effort with the Neuron SDK is real but it is paid once and the discount is paid every run. It is a strong PoC candidate: we typically prove it on one model in two weeks before touching the rest of the fleet.
Fix the boring efficiency losses
- Mixed precision (bf16 / fp16) is often a near free throughput win on modern GPUs.
- Data loading from S3 through FSx for Lustre or the S3 Express One Zone class keeps the accelerator fed. Then delete the FSx file system when the job ends; it is billed per provisioned GB, not per use.
- Distributed training should be justified by a scaling test. Eight GPUs at 40% efficiency cost more than two GPUs at 90%.
- Hyper-parameter sweeps deserve early stopping. Kill runs that are clearly losing after the first evaluation.
Inference: the silent majority of most ML bills
Training is spiky and visible. Inference is flat, permanent and easy to forget. Over a year it is usually the larger number.
Match the endpoint type to the traffic shape
| Traffic | Best fit | Why |
|---|---|---|
| Steady, latency-sensitive | Real-time endpoint with autoscaling, Savings Plans | Predictable base load is what commitments are for. |
| Spiky or low volume | SageMaker Serverless Inference | You pay per invocation and compute time, with zero cost at idle. |
| Large payloads, tolerant of seconds | Asynchronous Inference | Queues requests and can scale the instance count to zero. |
| Offline scoring | Batch Transform on Spot-friendly schedules | No endpoint needs to exist between runs. |
| Many small models | Multi-model endpoints | Hundreds of models share one fleet instead of one instance each. |
Right-size the accelerator, then remove it
A large share of production models do not need a GPU at all once quantised and compiled. Try CPU inference on Graviton with ONNX Runtime, or AWS Inferentia (Inf2) for transformer workloads, before defaulting to G5. Measure cost per thousand inferences at your latency target. That single metric usually settles the argument.
Scale to zero wherever the business allows
Staging and development endpoints should not exist outside working hours. Use scheduled scaling or delete and recreate from infrastructure as code. A GPU endpoint left running over a weekend is the single most common line item we remove in a first review.
Generative AI: manage tokens like you manage instances
Amazon Bedrock moves the cost unit from hours to tokens, and most teams have never budgeted in tokens before. The levers are different but the discipline is the same.
- Model tiering. Route by task difficulty. A fast, small model for classification, extraction and routing; a frontier model only for the steps that need reasoning. A simple router with an evaluation set behind it routinely cuts spend dramatically with no measurable quality loss.
- Prompt caching. Long system prompts, tool definitions and retrieved documents that repeat across calls should be cached. Cached input tokens are billed at a steep discount.
- Batch inference. Anything that does not need a synchronous answer, such as nightly summarisation, enrichment or evaluation runs, belongs in Bedrock batch mode at roughly half the on-demand token price.
- Provisioned Throughput only with proof. It is a commitment. Buy it when your steady-state token volume is measured, not forecast.
- RAG hygiene. Tighter chunking, fewer retrieved passages and re-ranking reduce input tokens on every request. Retrieval quality and cost improve together.
- Output limits. Set
max_tokensdeliberately. Verbose models are expensive models.
Track cost per successful task, not cost per token. A cheaper model that needs three retries is not cheaper.
Data and storage: the bill that never sleeps
- Move datasets to S3 Intelligent-Tiering and let lifecycle rules archive old training snapshots.
- Find the duplicate copies. Every team keeps “their” copy of the same raw data. One canonical bucket with versioning beats four.
- Convert EBS volumes on training and notebook instances to gp3 and stop paying for io2 IOPS nobody uses.
- Keep training data in the same region as the compute. Cross-region reads during training are a quiet, recurring tax.
- Add VPC endpoints for S3 and ECR so data and container pulls never traverse a NAT Gateway.
The “everything else” bucket
- Notebooks: enforce auto-shutdown with SageMaker Studio lifecycle configurations or idle-timeout scripts. Nobody trains on a notebook instance at 3 a.m.
- GPU dev boxes: replace with on-demand SageMaker Studio spaces or ephemeral EKS pods that disappear at night.
- Logs and metrics: debug-level training logs shipped to CloudWatch add up. Sample or ship to S3 instead.
- Experiments: set retention on experiment artefacts and model registry versions that never got promoted.
Make it stick: the FinOps layer
Optimisation done once decays in a quarter. Optimisation done as a practice compounds.
- Publish three unit metrics: cost per training run, cost per thousand inferences, cost per thousand tokens. Review them weekly with the team that can change them.
- Put AWS Budgets and Cost Anomaly Detection on the ML accounts with alerts routed to engineers, not only finance.
- Cover the measured steady state with SageMaker Savings Plans and Compute Savings Plans, and leave the spiky part to Spot and serverless.
- Give every experiment a budget and a kill date at creation time.
First-week checklist
- Tag every ML resource with project, stage and owner.
- Move all training jobs to Spot with checkpointing.
- List every endpoint and delete or schedule the non-production ones.
- Measure GPU utilisation on training and inference; downsize anything under 50%.
- Turn on prompt caching and route easy GenAI tasks to a smaller model.
- Auto-shutdown notebooks; delete idle FSx file systems and duplicate datasets.
- Define cost per run, per inference and per token, and put them on a dashboard.
None of this slows a good team down. Done well, it speeds them up, because the budget freed from idle GPUs is the budget that funds the next experiment.