6 + YoE - Data Engineer – Big Data / PySpark - any UST Location - Immediate Joiner

JobCrexa

Candidates ready to join immediately can share their details via email for quick processing.

CCTC | ECTC | Notice Period | Location Preference

Act fast for immediate attention!

Must-Have Skills

  • 6+ years of overall experience in Data Engineering / Big Data.
  • Strong understanding of Big Data concepts and architecture .
  • Strong hands-on experience with Apache Spark .
  • Expertise in:
  • Spark Performance Tuning
  • Spark Optimization
  • Query/Job Performance Improvement
  • Troubleshooting Spark workloads
  • Strong hands-on experience with PySpark and Spark .
  • Strong programming experience in Python .
  • Good experience working with MySQL / SQL .
  • Strong experience in designing and developing Data Pipelines .
  • Hands-on experience with Apache Airflow for data pipeline orchestration and scheduling.
  • Experience working with at least one Cloud Platform .
  • GCP experience is preferred.
  • Good understanding of CI/CD and DevOps concepts .
  • Experience integrating data engineering workloads with CI/CD pipelines.
  • Strong debugging, troubleshooting, and problem-solving skills.

Preferred Skills

  • Hands-on exposure to relevant GCP data services .
  • Experience handling large-scale and high-volume datasets.
  • Understanding of distributed data processing and data architecture.
  • Experience improving the scalability, reliability, and performance of data pipelines.
  • Exposure to Agile development and DevOps practices.

Key Responsibilities

  • Design, develop, and maintain scalable Big Data and Data Engineering solutions .
  • Develop data processing applications using Python, PySpark, and Apache Spark .
  • Perform Spark performance tuning and optimization for large-scale workloads.
  • Build, maintain, and monitor robust ETL/ELT data pipelines .
  • Develop and manage workflow orchestration using Apache Airflow .
  • Work with MySQL/SQL for data extraction, transformation, and validation.
  • Deploy and support data engineering solutions in cloud environments, preferably GCP .
  • Work with DevOps teams to implement and maintain CI/CD pipelines .
  • Troubleshoot production issues and optimize data processing performance.
  • Collaborate with engineering and business teams to deliver reliable and scalable data solutions.

Primary Skill Combination

Big Data + Apache Spark + PySpark + Python + Airflow + SQL/MySQL + Cloud (GCP Preferred) + CI/CD

Mandatory Focus: Strong hands-on Apache Spark performance tuning and optimization experience.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.