Senior DevOps Engineer_Req 117
3Pillar
Location: [Remote / Hybrid — specify]
Team: Platform Engineering / Innovation
Employment type: Full-time
Work Hours: EST 9am–5pm (EST availability required)
About The Role
We are hiring a Senior DevOps Engineer to own and evolve the deployment of our AI-powered application modernization platform (NEXUS) into client production environments. This role is a unique hybrid of Platform Engineering and Technical Advisory. You will not only design the underlying multi-cloud infrastructure (AWS, Azure, GCP) but also partner directly with client IT groups to establish SysOps frameworks. You will leverage our AI-driven automation tools to provision infrastructure, integrate AIOps for anomaly detection and remediation, and ensure clients receive seamless product updates via our Helm-based delivery pipeline.
This is a senior individual contributor role that requires a "product mindset"—you are designing for scale, security, and long-term operational sustainability in client environments.
What You Will Do
Client Collaboration & SysOps Partnership
- Serve as the primary technical partner to client IT groups, guiding them on best practices for operating bespoke NEXUS implementations.
- Work with client-aligned delivery teams to manage the lifecycle of the platform, ensuring timely and reliable product updates via Helm charts.
- Collaborate with client stakeholders to tailor infrastructure and operational workflows, ensuring NEXUS integrates seamlessly with their existing enterprise monitoring and ticketing systems.
- Build & Leverage AI for Provisioning: Architect and implement AI-assisted workflows to automate the provisioning and configuration of infrastructure across multi-cloud environments.
- AIOps Implementation: Integrate and adopt 3Pillar’s AIOps platform into client ecosystems to enable proactive, automated anomaly detection and incident remediation.
- Knowledge Base Evolution: Maintain and refine the AIOps knowledge base, training agents to perform accurate root-cause analysis (RCA) and automate recurring issue resolution. 1
- Evolve our Helm-based platform layer and Terraform module libraries to support standardized, secure deployment patterns across AWS, Azure, and GCP. 1 2
- Design for "Delegation": Move our automation from simple execution to self-adapting workflows that handle ambiguity and get smarter over time. 3
- Ensure all infrastructure-as-code and automated remediation scripts meet rigorous enterprise security and compliance standards, including PII access logging and least-privilege IAM.
You should be comfortable working with most of the following (or able to ramp quickly):
Area
Technologies
IaC
Terraform (>= 1.5), remote state (S3 + DynamoDB), module versioning via Git tags
Cloud
AWS (primary for client deployments): EKS, VPC, NLB, EFS, ECR, Route 53, KMS, IAM/IRSA
- Azure: AKS, Key Vault, ACR, VNet
- GCP: GKE, Artifact Registry
Production cluster operations, storage classes, ingress, HPA, troubleshooting CrashLoop / init failures
Packaging
Helm (umbrella charts, atomic releases), feature-flag-driven deployments
Platform services
HashiCorp Vault, Keycloak, OAuth2 Proxy, cert-manager, Zalando Postgres Operator, Neo4j, Redis Operator, Strimzi Kafka, Prometheus/Grafana/Loki, Apache Airflow, MinIO
Application platform
Python/FastAPI backend, React frontend, MCP server, background workers, retrieval/RAG services; container images on a private registry (e.g. GHCR)
CI/CD & GitOps
GitHub Actions, Flux (internal environments), image signing (Cosign)
Observability & ops
kubectl, helm, structured logging (Loki), metrics/alerts, operational runbooks
Security
TLS (Let's Encrypt DNS-01), secrets via Vault/VSO, SSO/OIDC, VPN / CIDR-based ingress controls
Required Experience
- Client-Facing Advisory & Enablement: Proven experience in a technical advisory or client-facing engineering role. You are comfortable partnering with client IT teams, guiding them through technical adoption, and translating complex platform requirements into operational success.
- DevOps & Infrastructure Expertise: 5+ years in DevOps, SRE, or Platform Engineering. You possess production-grade Terraform experience—architecting modular, multi-cloud IaC frameworks that enable consistent, repeatable deployments across different environments.
- AIOps & Observability: Hands-on experience integrating or managing AIOps platforms or automated observability stacks. You understand the lifecycle of incident response, anomaly detection, and automated remediation.
- AI-Driven Automation: Demonstrated experience building automation workflows. You don’t just write scripts; you build systems that reduce manual operational toil and empower users through self-service provisioning and automated configuration.
- Platform Engineering at Scale: Deep expertise in Kubernetes and Helm. You have a track record of deploying and operating complex, multi-service platforms and managing their lifecycle (chart development, values layering, atomic releases).
- Foundational Cloud Operations: Proven AWS expertise (EKS, networking, IAM/IRSA) with a strong aptitude for multi-cloud environments (Azure/GCP).
- Identity & Security: Solid experience with secrets management (HashiCorp Vault or similar) and identity/SSO integration (Keycloak/OIDC).
- Communication & Mentorship: Excellent written and verbal communication. You are able to explain complex technical concepts to non-technical stakeholders and mentor teams on best practices.
- Consultative Experience: Prior work as a Solutions Architect, Technical Account Manager, or Implementation Engineer in a SaaS or Managed Services environment.
- AI/LLM Infrastructure: Experience managing infrastructure for AI/ML workloads, including GPU node pools, vector databases (e.g., Neo4j), or RAG (Retrieval-Augmented Generation) pipelines.
- GitOps & Lifecycle: Hands-on experience with GitOps methodologies (Flux, Argo CD) for managing continuous deployment at scale.
- Enterprise Compliance: Experience operating within regulated or enterprise environments (PII logging, audit trails, network isolation, SOC2/compliance-heavy workflows).
- Kafka & Distributed Systems: Experience with event streaming platforms (e.g., Strimzi/Kafka) or specialized operators for stateful services.
- Successful enablement of client IT teams to independently manage and operate the NEXUS platform.
- Quantifiable reduction in manual operational toil for clients through the successful adoption of AIOps-driven anomaly detection and automated remediation.
- Seamless, low-friction delivery of platform updates via automated Helm-based pipelines.
- Deployment of robust, AI-provisioned infrastructure stacks that adapt to client-specific demands while maintaining security and compliance.