On this page
There is a moment in every ML project where the notebook says 94 percent accuracy and everyone celebrates. Then someone asks how users will actually use it, and the room goes quiet.
I built the serving layer for an in house image model that processes user uploaded medical images inside a live product workflow. The model was the easy 20 percent. This post is the other 80: the contracts, queues, versioning and failure design that turn a promising checkpoint into a feature people rely on.
The serving layer is a contract, not a wrapper
The notebook assumed clean inputs. Users upload sideways HEIC photos from an iPhone, 40MB scans, screenshots of photos, and occasionally a PDF renamed to .jpg. Your serving layer owns that gap:
- Validate early. File type, dimensions and size, rejected with a human error message before wasting GPU time. "Image too small to analyse, please retake closer" beats a stack trace every day.
- Normalize deterministically. The exact resize, color and orientation pipeline the model was trained on, implemented once and versioned with the model. Skew between training preprocessing and serving preprocessing is the most common silent accuracy killer I have seen, and it never announces itself.
- Type the output. A schema with scores, labels, confidence and model version. Not a raw tensor dump the frontend has to guess about.
Never run inference inside the request
Inference takes seconds. Requests should take milliseconds. Our flow:
- The upload endpoint validates, stores the file, creates an analysis row with status queued, and returns instantly.
- A queue worker picks it up, runs preprocessing and inference, writes results.
- The client follows status over polling or a websocket, and the product flow continues when the analysis completes.
This buys three things a synchronous endpoint can never give you: retries when a worker dies mid job, backpressure when a clinic uploads a hundred images at once, and independent scaling of GPU workers versus web servers. GPU boxes are expensive; you want exactly as many as the queue depth justifies, not as many as your web traffic implies.
Batching: the cheapest performance win in ML serving
GPUs love company. Processing images one by one wastes most of the capacity you are paying for. Even a tiny batching window, collecting requests for 50 to 100 milliseconds and running them as one batch, can multiply throughput. You do not need fancy infrastructure: a worker that drains up to N items from the queue per cycle captures most of the benefit. On our workload that one change halved the number of GPU instances we needed, which the finance side of the project noticed before the engineering side did.
Version everything, trust nothing
Every result row records the model version, the preprocessing version and the thresholds used. When the data science team ships v2, two capabilities matter more than any accuracy claim:
- Shadow mode. Run v2 alongside v1 on real traffic, store both results, compare offline. No user impact while the new model proves itself on data nobody curated.
- Instant rollback. Model binaries are artifacts in object storage with version tags, and switching back is a config change, not a deploy.
There will be a first time a "better" model turns out worse on real world data. Ours was v2 stumbling on a phone camera type the training set barely contained. Shadow mode caught it in a week of quiet comparison instead of a week of clinical complaints.
Fail like a product, not like a script
Some images are blurry, cropped wrong, or genuinely unprocessable. Decide the product behavior deliberately: ask for a retake with specific guidance, degrade to a partial result, or route to human review. A model that fails gracefully builds trust with every failure. A spinner that ends in "error" spends trust it never earns back.
The checklist I hand to teams
Input validation with human messages. Preprocessing versioned with the model. Async inference behind a queue. Batching. Shadow mode for new versions. One command rollback. Confidence thresholds tuned with the product owner, not defaulted. Failure states designed in the UI, not discovered by users. None of it is glamorous, which is exactly why it separates shipped ML from demoed ML.
The stack that carried it
FastAPI for the inference service, Celery and RabbitMQ for orchestration, S3 compatible storage for artifacts, Postgres for results, Docker for packaging, deployed across AWS and Azure. Nothing exotic. Production ML is mostly excellent plumbing, and plumbing is a compliment.
Have a model that needs to become a product feature? That is my favorite kind of project. Let's talk.
Need this built?
These services can turn the ideas in this article into production software.

