Senior Systems Engineer - Data DevOps/MLOps
EPAM Systems
The successful applicant should have thorough understanding of data engineering, automated data pipelines, and deployment of machine learning models in production. This position requires a collaborative individual capable of architecting, implementing, and overseeing large-scale data and ML pipelines that support company goals.
Responsibilities
- Build, launch, and oversee CI/CD pipelines supporting data integration and ML model rollout
- Establish and maintain cloud-based infrastructure for data processing and model training
- Streamline data validation, transformation, and workflow orchestration through automation
- Partner with data scientists, software engineers, and product teams to ensure seamless ML model integration into production environments
- Improve model serving and monitoring capabilities to increase performance and reliability
- Oversee data versioning, lineage tracking, and reproducibility of ML experiments
- Continuously identify opportunities to improve deployment workflows, scalability, and infrastructure resilience
- Enforce robust security measures to protect data integrity and ensure regulatory compliance
- Diagnose and resolve problems across the entire data and ML pipeline lifecycle
- Bachelor's or Master's degree in Computer Science, Data Engineering, or related discipline
- Minimum 5 years of experience in Data DevOps, MLOps, or comparable positions
- Skilled in cloud platforms such as Azure, AWS, or GCP
- Experienced with Infrastructure as Code tools like Terraform, CloudFormation, or Ansible
- Strong knowledge of containerization and orchestration tools, including Docker and Kubernetes
- Practical experience with data processing frameworks such as Apache Spark and Databricks
- Skilled in programming languages like Python, with familiarity in data manipulation and ML libraries such as Pandas, TensorFlow, and PyTorch
- Knowledgeable in CI/CD tools such as Jenkins, GitLab CI/CD, and GitHub Actions
- Experienced with version control systems and MLOps platforms including Git, MLflow, and Kubeflow
- Solid grasp of monitoring, logging, and alerting tools such as Prometheus and Grafana
- Strong problem-solving skills with the ability to perform well both independently and collaboratively
- Excellent communication and documentation abilities
- Experience with DataOps principles and tools like Airflow and dbt
- Understanding of data governance platforms such as Collibra
- Exposure to Big Data technologies including Hadoop and Hive
- Cloud or data engineering certifications