Cloud Solution Architect – Enterprise AI Infrastructure & FinOps

Zydus Group

We are looking for a Cloud Solution Architect whose primary responsibility is to design, build, manage, and optimize the enterprise cloud infrastructure that runs our AI ecosystem.

Job Description:

Role: Cloud Solution Architect – Enterprise AI Infrastructure & FinOps

Experience: 12–18 years total, with at least 6 years in cloud architecture and at least 3 years architecting production AI/ML or GenAI platforms

Education: B.E./B.Tech in Computer Science, Information Technology, Information Security, Electronics & Communication, or related Engineering disciplines.

Location: Ahmedabad/Mumbai


Role Summary

The individual hired as Cloud Solution Architect will be responsible for cloud foundation end to end: landing zones, networking, compute, storage, identity, security, Kubernetes platforms, GPU capacity, infrastructure-as-code, monitoring, resilience, and day-to-day infrastructure operations.


The role will ensure that AI and GenAI workloads operate securely, reliably, and at scale while driving cloud governance, operational excellence, and FinOps practices to maintain transparency, control, and optimization of cloud and AI spending.


This is a hands-on infrastructure leadership role. Application and AI engineering teams build agents and workflows; you make sure the platform they run on is secure, stable, scalable, compliant, and cost-efficient.


Key Responsibilities

1. Cloud Infrastructure Architecture

  • Own the enterprise cloud infrastructure architecture on Azure, AWS, and/or GCP, including multi-cloud and hybrid designs.
  • Design and maintain enterprise landing zones: management group and account/subscription structure, policies, guardrails, naming and tagging standards, and environment separation (dev, test, validation, production).
  • Architect network infrastructure: hub-and-spoke or Virtual WAN topologies, VNets/VPCs, subnets, private endpoints, DNS, load balancers, application gateways, WAF, ExpressRoute/Direct Connect, site-to-site VPN, and hybrid connectivity to data centers and manufacturing sites.
  • Design compute platforms across VMs, VM scale sets, containers, Kubernetes (AKS, EKS, GKE), serverless (Functions, Lambda, Cloud Run), and managed PaaS services.
  • Architect storage and database infrastructure: object storage, file shares, managed disks, backup vaults, and managed databases (SQL, PostgreSQL, Cosmos DB, DynamoDB, Redis).
  • Design for high availability, multi-region resilience, disaster recovery, and business continuity with defined RTO and RPO targets.
  • Maintain architecture documentation, reference designs, and architecture decision records (ADRs).

2. Cloud Infrastructure Management and Operations

  • Own the day-to-day health, stability, and performance of cloud infrastructure supporting AI and enterprise workloads.
  • Lead infrastructure provisioning, configuration, patching, upgrades, and lifecycle management for VMs, Kubernetes clusters, networking, and platform services.
  • Manage Kubernetes platform operations: cluster upgrades, node pools (CPU and GPU), ingress, service mesh, autoscaling, namespaces, and multi-tenancy for AI teams.
  • Define and track infrastructure SLOs, SLAs, and capacity plans, and lead major incident response, root cause analysis, and problem management.
  • Implement backup, restore, and DR testing on a regular schedule and maintain runbooks.
  • Establish ITIL-aligned processes for incident, change, problem, and service request management in collaboration with IT operations and managed service providers.
  • Manage cloud vendors and managed service partners, including SLA and performance governance.
  • Drive operational automation to reduce manual effort and human error.

3. Infrastructure-as-Code, Automation and Platform Engineering

  • Define and enforce infrastructure-as-code standards using Terraform, Bicep/ARM, Pulumi, or CloudFormation.
  • Build reusable IaC modules, templates, and "golden paths" so teams can self-serve compliant infrastructure quickly.
  • Implement CI/CD and GitOps for infrastructure (GitHub Actions, Azure DevOps, GitLab, Argo CD, Flux).
  • Apply policy-as-code (Azure Policy, AWS SCPs/Config, OPA/Gatekeeper) to prevent misconfiguration and enforce security and cost guardrails.
  • Automate environment provisioning, scaling, patching, compliance checks, and cleanup.
  • Build an internal developer platform experience for AI, data, and application teams.

4. AI Infrastructure Enablement

  • Design and operate the infrastructure foundation for AI workloads, covering agent runtimes, workflow engines, data pipelines, model serving, and vector and search services.
  • Provision and manage GPU and accelerator capacity (NVIDIA A100/H100/L40S, Inferentia/Trainium, TPUs) for inference, fine-tuning, and batch processing, including quota management and capacity reservations.
  • Host and secure LLM access infrastructure: Azure OpenAI, AWS Bedrock, Vertex AI, and self-hosted open-source models (vLLM, TGI, Triton, KServe, Ray).
  • Deploy and operate a centralized AI/LLM gateway for secure model access, routing, rate limiting, caching, logging, and cost attribution.
  • Provide infrastructure for agent frameworks and orchestration tools (e.g., LangGraph, Semantic Kernel, Azure AI Foundry, Bedrock Agents, Temporal, Airflow, Logic Apps, Step Functions), including MCP servers and secure tool connectors.
  • Support data platforms used by pipelines, such as Databricks, Snowflake, Microsoft Fabric, Synapse, Kafka/Event Hub, and vector stores (Azure AI Search, OpenSearch, pgvector, Pinecone, Qdrant).
  • Ensure network isolation, private connectivity, and secure integration between AI services and enterprise systems (SAP, LIMS, MES, QMS, document repositories, and on-premises data).
  • Partner with AI engineering teams on LLMOps/MLOps infrastructure: model registries, experiment tracking, CI/CD for AI assets, and environment promotion.
  • Architect the agent runtime platform: hosting, lifecycle management, state and memory, tool execution, session handling, and scaling.
  • Define standards for agent frameworks and patterns (e.g., LangGraph, Semantic Kernel, AutoGen, CrewAI, OpenAI Agents SDK, AWS Bedrock Agents, Azure AI Foundry Agent Service).
  • Design multi-agent orchestration patterns: supervisor/worker, hierarchical, planner-executor, and agent-to-agent (A2A) collaboration.
  • Design the orchestration layer for AI-enabled business workflows combining agents, deterministic logic, APIs, and human tasks.
  • Standardize on durable orchestration and workflow engines (e.g., Temporal, Azure Durable Functions, Logic Apps, AWS Step Functions, Airflow, Prefect, n8n, Power Automate) based on use case.
  • Define event-driven architectures using Kafka, Event Hub, EventBridge, or Pub/Sub for real-time, asynchronous triggers.

5. Cloud FinOps and AI Cost Management

  • Establish and run the cloud FinOps practice based on FinOps Foundation principles (Inform, Optimize, Operate) and the FOCUS billing standard.
  • Implement mandatory tagging and cost allocation so every resource and AI workload maps to an owner, cost center, application, agent, workflow, or pipeline.
  • Deliver showback and chargeback reporting to business units and product teams.
  • Track cloud spend across compute, storage, networking, data egress, PaaS services, GPUs, and LLM token consumption.
  • Define unit economics such as cost per agent interaction, cost per workflow run, cost per pipeline execution, and cost per environment.
  • Drive infrastructure cost optimization through:
  • Rightsizing VMs, clusters, and databases
  • Reserved instances, savings plans, and committed use discounts
  • Spot and preemptible capacity for non-critical and batch workloads
  • Autoscaling, scale-to-zero, and scheduled shutdown of non-production environments
  • Storage tiering, lifecycle policies, and orphaned resource cleanup
  • GPU utilization improvement and workload consolidation
  • LLM cost controls such as model routing, caching, batch APIs, and provisioned throughput decisions
  • Set budgets, quotas, and anomaly alerts per subscription, team, and AI workload.
  • Produce forecasts and executive dashboards linking cloud and AI spend to business value.
  • Support Finance and Procurement on enterprise agreements, commitments, and vendor negotiations.

6. Cloud Security and Compliance

  • Design and enforce cloud security architecture based on zero-trust principles.
  • Own identity and access architecture (Entra ID, AWS IAM, GCP IAM), RBAC, privileged access management, managed identities, and workload identity for agents and services.
  • Implement secrets and key management (Key Vault, KMS, HashiCorp Vault), encryption at rest and in transit, and certificate lifecycle management.

7. Monitoring, Observability and Reliability

  • Architect enterprise monitoring and observability for infrastructure and AI workloads using Azure Monitor, CloudWatch, Google Cloud Operations, Datadog, Dynatrace, Prometheus/Grafana, and OpenTelemetry.
  • Monitor infrastructure health, performance, capacity, GPU utilization, network latency, and cost in unified dashboards.
  • Integrate AI-level telemetry (token usage, latency, failures) from tools such as Langfuse or LangSmith with infrastructure observability.
  • Apply SRE practices: error budgets, proactive alerting, chaos testing, and post-incident reviews.

8. Technical Leadership and Governance

  • Act as the design authority for cloud infrastructure, reviewing and approving infrastructure designs from project and product teams.
  • Define cloud standards, policies, and best practices, and drive consistent adoption across teams.
  • Lead and mentor cloud engineers, DevOps/platform engineers, and operations staff.


Required Skills and Experience

  • 12+ years in IT infrastructure, with at least 7 years of hands-on cloud infrastructure architecture and operations on Azure, AWS, or GCP (multi-cloud preferred).
  • Deep expertise in cloud networking, identity, compute, storage, and security services.
  • Strong hands-on experience with Kubernetes (AKS/EKS/GKE) and containers in production, including GPU node pools.
  • Advanced infrastructure-as-code skills (Terraform preferred) and CI/CD/GitOps practices.
  • Proven experience managing production cloud environments at enterprise scale, including incidents, DR, patching, and change management.
  • Demonstrated experience implementing cloud FinOps: tagging, allocation, optimization, and reporting.
  • Practical experience supporting AI/ML or GenAI workloads: LLM services, model serving, GPU infrastructure, vector databases, and data platforms.
  • Strong understanding of cloud security, zero trust, and compliance frameworks.
  • Scripting skills in Python, PowerShell, or Bash.
  • Experience in regulated industries; pharma or life sciences GxP experience is a strong advantage.


Preferred Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
  • Certifications such as Azure Solutions Architect Expert, AWS Solutions Architect Professional, Google Professional Cloud Architect, CKA/CKAD, HashiCorp Terraform Associate, FinOps Certified Practitioner/Engineer, Azure AI Engineer, AWS ML Specialty, or ITIL.
  • Experience with hybrid infrastructure connecting cloud to on-premises data centers and manufacturing/OT environments.
  • Exposure to AI governance and Responsible AI frameworks.


How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.