Serving a Custom ML Model in Production: The Plumbing Nobody Talks About

The data scientists hand you a model that works in a notebook. Making it survive real traffic, uploads, GPUs, queues and versioning is a different job entirely.

Ijaz KhanSeptember 24, 2026 4 min read
On this page

There is a moment in every ML project where the notebook says 94 percent accuracy and everyone celebrates. Then someone asks how users will actually use it, and the room goes quiet.

I built the serving layer for an in house image model that processes user uploaded medical images inside a live product workflow. The model was the easy 20 percent. This post is the other 80: the contracts, queues, versioning and failure design that turn a promising checkpoint into a feature people rely on.

The serving layer is a contract, not a wrapper

The notebook assumed clean inputs. Users upload sideways HEIC photos from an iPhone, 40MB scans, screenshots of photos, and occasionally a PDF renamed to .jpg. Your serving layer owns that gap:

  • Validate early. File type, dimensions and size, rejected with a human error message before wasting GPU time. "Image too small to analyse, please retake closer" beats a stack trace every day.
  • Normalize deterministically. The exact resize, color and orientation pipeline the model was trained on, implemented once and versioned with the model. Skew between training preprocessing and serving preprocessing is the most common silent accuracy killer I have seen, and it never announces itself.
  • Type the output. A schema with scores, labels, confidence and model version. Not a raw tensor dump the frontend has to guess about.

The server racks where notebook models meet production traffic

Never run inference inside the request

Inference takes seconds. Requests should take milliseconds. Our flow:

  1. The upload endpoint validates, stores the file, creates an analysis row with status queued, and returns instantly.
  2. A queue worker picks it up, runs preprocessing and inference, writes results.
  3. The client follows status over polling or a websocket, and the product flow continues when the analysis completes.

This buys three things a synchronous endpoint can never give you: retries when a worker dies mid job, backpressure when a clinic uploads a hundred images at once, and independent scaling of GPU workers versus web servers. GPU boxes are expensive; you want exactly as many as the queue depth justifies, not as many as your web traffic implies.

Batching: the cheapest performance win in ML serving

GPUs love company. Processing images one by one wastes most of the capacity you are paying for. Even a tiny batching window, collecting requests for 50 to 100 milliseconds and running them as one batch, can multiply throughput. You do not need fancy infrastructure: a worker that drains up to N items from the queue per cycle captures most of the benefit. On our workload that one change halved the number of GPU instances we needed, which the finance side of the project noticed before the engineering side did.

Version everything, trust nothing

Every result row records the model version, the preprocessing version and the thresholds used. When the data science team ships v2, two capabilities matter more than any accuracy claim:

  • Shadow mode. Run v2 alongside v1 on real traffic, store both results, compare offline. No user impact while the new model proves itself on data nobody curated.
  • Instant rollback. Model binaries are artifacts in object storage with version tags, and switching back is a config change, not a deploy.

There will be a first time a "better" model turns out worse on real world data. Ours was v2 stumbling on a phone camera type the training set barely contained. Shadow mode caught it in a week of quiet comparison instead of a week of clinical complaints.

Fail like a product, not like a script

Some images are blurry, cropped wrong, or genuinely unprocessable. Decide the product behavior deliberately: ask for a retake with specific guidance, degrade to a partial result, or route to human review. A model that fails gracefully builds trust with every failure. A spinner that ends in "error" spends trust it never earns back.

The checklist I hand to teams

Input validation with human messages. Preprocessing versioned with the model. Async inference behind a queue. Batching. Shadow mode for new versions. One command rollback. Confidence thresholds tuned with the product owner, not defaulted. Failure states designed in the UI, not discovered by users. None of it is glamorous, which is exactly why it separates shipped ML from demoed ML.

The stack that carried it

FastAPI for the inference service, Celery and RabbitMQ for orchestration, S3 compatible storage for artifacts, Postgres for results, Docker for packaging, deployed across AWS and Azure. Nothing exotic. Production ML is mostly excellent plumbing, and plumbing is a compliment.

Have a model that needs to become a product feature? That is my favorite kind of project. Let's talk.

#ml#model-serving#fastapi#ai#mlops

Need this built?

These services can turn the ideas in this article into production software.

Keep reading

All articles
For Founders6 min read

Adding AI to Your SaaS: Scope, Risks and Launch Checklist

A founder’s guide to adding AI to an existing SaaS: choose one workflow, define permissions and acceptance criteria, then plan a controlled rollout.

Read article
Backend6 min read

Django vs FastAPI: How I Choose for Client Projects

I have shipped production systems with both. The honest answer to which one is better is a decision tree, not a fan war. Here is mine.

Read article