Senior MLOps / AI Platform Engineer
Aivar Innovations
Overview
Aivar Innovations is looking for a Senior MLOps / AI Platform Engineer to help design and build an enterprise-grade MLOps and AIOps software platform running on Kubernetes. You will be responsible for developing the infrastructure and platform capabilities required to deploy, manage, scale, and observe machine learning, deep learning, and generative AI workloads across cloud and on-premises environments. You will work extensively with Kubernetes, GPUs, model-serving frameworks, distributed systems, and cloud-native technologies. You will collaborate closely with Product, Engineering, AI/ML, DevOps, and Customer Delivery teams to transform complex AI infrastructure requirements into reliable, secure, and easy-to-use platform capabilities. This role requires strong hands-on engineering experience and a practical understanding of how machine learning models move from experimentation into reliable production environments.
Requirements
5–8 years of experience in software engineering, platform engineering, DevOps, SRE, MLOps, or related infrastructure roles. Strong hands-on experience with Kubernetes, including writing Kubernetes Operators and Custom Resource Definitions using frameworks such as Kubebuilder, Operator SDK, or equivalent. Experience designing and operating cloud-native infrastructure on AWS, particularly Amazon EKS, EC2, S3, ECR, IAM, VPC, and CloudWatch. Experience deploying and operating machine learning, deep learning, or generative AI models in production. Experience running and troubleshooting GPU-accelerated workloads on Kubernetes, with an understanding of GPU scheduling, utilization, memory constraints, and performance. Familiarity with model-serving frameworks such as KServe, NVIDIA Triton Inference Server, vLLM, Ray Serve, TorchServe, or equivalent technologies. Strong programming experience in Go or Python, with experience building production-grade APIs, controllers, or distributed backend services. Experience with containers, Helm, CI/CD, infrastructure as code, and observability tools such as Prometheus, OpenTelemetry, and Grafana. Strong understanding of Linux, networking, storage, security, and distributedsystem fundamentals. Strong debugging, problem-solving, communication, and cross-functional collaboration skills.
Preferred Qualifications
Experience building an MLOps platform, AI infrastructure platform, internal developer platform, or Kubernetes-based enterprise product. Experience writing GPU kernels or performance-critical code using CUDA C/C++ or Triton is a plus. Experience with large language model serving, distributed inference, batching, quantization, or inference-performance optimization. Experience with NVIDIA GPU Operator, MIG, GPU time-slicing, Dynamic Resource Allocation, or similar GPU-management technologies. Familiarity with AWS Inferentia, Trainium, SageMaker, or Amazon Bedrock. Experience operating AI platforms across hybrid-cloud, on-premises, airgapped, or multi-tenant environments. Why You’ll Love Working at Aivar Build a Core AI Platform: Help create a Kubernetes-native platform that enables enterprises to deploy and operate AI workloads at scale. Solve Challenging Infrastructure Problems: Work on Kubernetes, AWS, GPUs, distributed systems, model serving, and enterprise AI operations. Influence
Product Direction
Work closely with Product, Engineering, and Leadership to shape the platform architecture and roadmap. Modern Technology Stack: Work with cloud-native infrastructure, GPU technologies, generative AI systems, observability platforms, and modern engineering practices. High Ownership: Lead major technical initiatives and take platform capabilities from architecture through production deployment. Accelerated Growth: Build your career in a fast-growing AI startup where your technical decisions will have visible and lasting impact.