Yes. I have built the serving layer for an in-house trained vision model processing user-uploaded medical images in production: queued inference through background workers, structured outputs with confidence levels and separate scaling for the inference tier. That pattern applies to most custom model deployments.
Have a question specific to your product?
Share the context and I will reply with a practical next step.