Contractor - PySpark Engineer

Vivantify Technology Solutions India Pvt. Ltd.

Location: Pan India

Experience: 5–7 Years

About The Role

We are looking for an experienced PySpark Engineer with strong expertise in Apache Spark, PySpark, Spark SQL, Python, and Spark Streaming. The role focuses on designing and developing enterprise-grade Spark applications and scalable ETL data pipelines, with hands-on experience in Delta Lake, cloud-based Big Data platforms, and Spark cluster management.

Job Description

The role involves developing and maintaining Spark applications using PySpark and building enterprise-grade data pipelines. You will work with Delta Lake, Spark Streaming, Structured Streaming, SQL, and cloud-based Big Data platforms, with a focus on application performance and optimization.

The position also involves setting up CI/CD pipelines, testing Spark applications, managing Spark clusters, and supporting highly available data environments. You will collaborate with Agile development teams to implement data strategies, build data flows, and support conceptual data models.

The role additionally includes environment planning, resource management, and the ability to manage a small team of technical specialists while supporting sprint-based development requirements.

Key Responsibilities

  • Design and develop Spark applications using PySpark
  • Build and maintain enterprise-grade ETL data pipelines
  • Implement solutions using Delta Lake
  • Develop and support Spark Streaming and Structured Streaming applications
  • Optimize Spark application performance and manage cluster resources
  • Write and optimize SQL/Spark SQL scripts
  • Build and maintain CI/CD pipelines for PySpark applications
  • Configure, monitor, and manage Spark clusters across Standalone, YARN, and Kubernetes environments


Required Skills

  • Apache Spark
  • PySpark
  • Python
  • Python for Data
  • Spark SQL
  • Spark Streaming and Structured Streaming
  • Delta Lake
  • ETL data pipeline development
  • Spark application optimization and performance tuning
  • Spark cluster sizing and resource management
  • Spark Standalone, YARN, and Kubernetes
  • GCP Big Data services, including Dataproc and GCS
  • Kubernetes
  • CI/CD and DevOps pipeline setup
  • Spark application testing and performance testing
  • Highly available Spark cluster operations and monitoring
  • Agile development methodology


Good to Have

  • Java or Scala
  • AWS
  • Azure


Mandatory Skills

  • Apache Spark
  • Python
  • Python for DATA
  • SparkSQL
  • Spark/PySpark
  • Spark Streaming
  • SQL
  • Spark clustering
  • DevOps pipeline setup
  • Kubernetes
  • GCP Cloud
  • Big Data platforms


Candidate Profile

The ideal candidate should have strong hands-on experience designing Spark applications and building enterprise-grade ETL data pipelines using PySpark, Python, Spark SQL, and Spark Streaming. The candidate should also have experience with cloud Big Data platforms, Spark cluster management, CI/CD, performance optimization, and Agile delivery, along with the ability to manage a small team of technical specialists.

Benefits Package

  • Competitive salary and benefits package
  • Opportunities for professional growth and development
  • A collaborative and inclusive work environment
  • The opportunity to work on exciting and challenging projects with leading clients

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.