Lead Systems Engineer - Data DevOps/MLOps
EPAM Systems
Responsibilities
- Build, launch, and oversee CI/CD pipelines that enable smooth data integration and ML model rollout
- Create a strong infrastructure foundation for training, processing, and serving machine learning models through cloud-based platforms
- Streamline operations by automating essential workflows like data transformation, validation, and orchestration
- Partner with cross-functional teams such as data engineers and scientists to bring ML solutions into live production environments
- Enhance reliability, performance monitoring, and model serving within production systems
- Maintain reproducibility, lineage tracking, and data versioning throughout ML workflows and experiments
- Spot and act on opportunities to boost the infrastructure's resilience, efficiency, and scalability
- Apply strict security protocols to protect data and maintain regulatory compliance
- Troubleshoot and fix technical problems within ML deployment workflows and data pipelines
- A Bachelor's or Master's degree in Data Engineering, Computer Science, or a related area
- Over 8 years working in MLOps, Data DevOps, or similar fields
- Strong command of cloud platforms including GCP, AWS, or Azure
- Proficiency with Infrastructure as Code tools such as Ansible, CloudFormation, or Terraform
- Capability in orchestration and containerization technologies like Kubernetes and Docker
- Practical experience using data processing frameworks such as Databricks and Apache Spark
- Strong Python skills along with familiarity with libraries like PyTorch, TensorFlow, and Pandas
- Familiarity with CI/CD tools including GitHub Actions, GitLab CI/CD, and Jenkins
- Working knowledge of MLOps platforms and version control systems such as Kubeflow, MLflow, and Git
- Awareness of alerting and monitoring tools such as Grafana and Prometheus
- Solid ability to solve problems and make decisions independently
- Strong skills in technical documentation and communication
- Experience with DataOps tools and methodologies such as dbt or Airflow
- Familiarity with data governance platforms such as Collibra
- Exposure to Big Data technologies like Hive or Hadoop
- Certifications demonstrating expertise in data engineering tools or cloud platforms