Senior DevOps Engineer
Asymbl
About Asymbl
Asymbl is the workforce orchestration company bringing together recruiting technology to hire people, digital labor strategy to onboard digital workers, and platform expertise to bring it all together. We help businesses attract, design, manage, and scale a hybrid workforce that drives more meaningful business impact, faster.
We are building a company where human and digital workers operate as one coordinated system of work, and where go-to-market execution is as intentional as the products and services we deliver.
About the Role
We're also Customer Zero for everything we ship. Rosa, Digital Recruiter, and the rest of our digital workforce run inside our own business before they run inside anyone else's. That means release engineering here isn't a back office concern. It's the floor the whole product stands on.
You'll own how our technology gets built, shipped, and run. This role reports to the Engineering Manager and is based in our Jaipur office.
Three systems have to move together for one customer to get one release. There's Salesforce, where our workforce applications live. There's Amazon Web Services (AWS), where our services and the runtimes our digital workers operate in are provisioned per customer through infrastructure as code. And there's a front end on Vercel with its own repository and pipeline. Separate repos, separate pipelines, one customer-facing release. Holding that together is the job.
We're also not a single-cloud shop, and we're not going to pretend the toolchain is settled. Some of what we run sits on Google Cloud. Some customers will insist their whole pipeline stays inside their own AWS boundary, which means building the same release path a second way. Having an opinion about those tradeoffs matters more here than having memorised one vendor's console.
Digital workers are also a different kind of workload from a normal application. Long-running jobs. Bursty load. Per-call costs that can move a lot in a week. Failure modes where nothing is technically down and the work is still wrong. Most release playbooks weren't written for that. You'll help us write ours.
One more thing worth naming plainly, because it's unusual: we're building our runbooks so digital workers can execute them, not just read them. Onboarding a customer, cutting a release, running an upgrade. Each of those is becoming a documented process a digital worker can run with a human approving the checkpoints. Deciding which steps stay human and which move to a digital worker, then building and maintaining the skills behind them, is part of this role.
This is a hands-on senior role on a small team. You'll run the pipeline, hold the pager, and set the standards other engineers follow.
Why Join Us?
At Asymbl, you won't inherit a release function, you'll build one. Small teams, short feedback loops, and a strong bias toward shipping. We use our own products daily, which means engineers hear about problems from colleagues before they hear about them from customers.
We're direct with each other. We write things down. And we'll admit some of this is still being figured out, because building release infrastructure for a workforce of humans and digital workers is new ground for everybody.
What You'll Own
- Salesforce deployments across a complex multi-org environment. Code and manual configuration, documented installation procedures, and post-deployment checks that catch problems before customers do
- Both of our Salesforce release paths: major releases with schema and packaging gates, and fast-track patches. Both need quality assurance (QA) sign-off, and both need to stay predictable under pressure
- Second-generation packaging (2GP) and the version dependencies between our packages, so a customer never lands in a half-upgraded state
- Sandbox lifecycle. Refreshes for new and existing sandboxes, user migration, and enough environment hygiene that developers aren't blocked waiting on an org
- Per-customer AWS provisioning through infrastructure as code. Bootstrapping new customer accounts, standing up the full stack for a tenant, and verifying it end to end before anyone signs off
- Customer upgrade strategy for our AWS services. Nobody owns this today, and we'd rather say that out loud than pretend otherwise. You'll define how existing customers move to new versions, what's automated, what needs a window, and what the rollback looks like
- The serverless runtime our digital workers operate in. AWS Lambda, Step Functions for the long-running orchestration, and Batch for the bulk jobs. Concurrency limits, quota increases, memory and duration tuning, and the retry and idempotency behaviour that decides whether a failed run is recoverable or just gone
- Secrets across the whole estate. AWS Secrets Manager, environment variables scoped per environment in Vercel, and Doppler as the source of truth that syncs into all of it. Rotation, least-privilege access, and a clear answer to “which system holds the real value”
- Connectivity and identity across clouds. Service quotas, graph and vector data services, the Connected App and OAuth handoff from Salesforce into AWS, and cross-account deploy roles using OpenID Connect (OIDC) or workload identity federation instead of long-lived keys
- Model platforms on more than one cloud. Google Cloud Platform (GCP) Vertex AI endpoints and Gemini, alongside whatever we run on AWS. Quotas, regional availability, batch against online inference, context caching, and the integration patterns for tool and function calling
- Continuous integration and continuous delivery (CI/CD) across all three systems. Branching strategy, protected main branches with required pull requests, lint, typecheck, build, and secret scanning on every push, plus the auto-deploy path to Vercel
- More than one CI/CD toolchain, on purpose. GitHub Actions is where we are today. Some customers require everything inside the AWS boundary, which means CodePipeline, CodeBuild, CodeDeploy, and CodeCommit with the Identity and Access Management (IAM) and Virtual Private Cloud (VPC) endpoint story that comes with them. You'll be able to build the same release path either way and advise on which a given customer needs
- Repository and access governance. GitHub App permissions per customer repo, and the discipline that keeps credentials out of local environment files
- Pipeline health as a real deliverable. When a tooling migration quietly breaks a deploy step, you're the person who finds it fast and stops it happening the same way twice
- Runtime choice and the reasoning behind it. Where a managed platform like Vercel is the right answer, where it isn't, and where an Elastic Compute Cloud (EC2) instance or a container is what the workload actually needs. Function duration limits, cold starts, long-running processes, VPC attachment, and who is responsible for patching are all part of that call
- Cost sustainability, not just cost reporting. Per-customer analysis plus the cost models underneath it: Lambda priced on requests and gigabyte-seconds, Step Functions priced very differently for Standard than for Express, Secrets Manager charged per secret per month so a few hundred per-customer secrets stop being a rounding error, and inference charged per token. Tagging and per-tenant attribution, budgets and anomaly alerts, right-sizing, and the internal tooling that makes all of it legible to engineering and to the people quoting the deal
- Usage metering that billing can rely on. Collecting and aggregating usage events so consumption-based pricing reflects what actually ran. We use OpenMeter for this, and the harder part is not the tool, it's deciding what counts as a billable unit of digital worker output
- Observability for work that is wrong rather than down. Tracing and evaluation for model calls through Langfuse, so a run that completed successfully and produced nonsense is visible. Token and cost attribution per trace sits here too
- Error tracking and release health on the front end and the services. Sentry or PostHog class tooling, wired to releases so a spike is traceable to a deploy, with source maps, alerting, and enough context that the on-call engineer isn't guessing
- Onboarding time as a metric you're accountable for. We've already cut it hard through better runbooks. We want it faster still, and mostly unattended
- Agent-executable process documentation. Runbooks written and wired so a digital worker can carry out onboarding, release, and upgrade steps with a human in the loop at the checkpoints that matter
- The incident practice itself. Service level objectives (SLOs) that mean something, the on-call rotation, escalation, and blameless postmortems that produce real fixes rather than a document nobody reads again
- Audit trail and documentation. Deployment status tracking, build records, and the evidence customers ask for during security review
- Developer experience. Local setup, test speed, and the small friction points that quietly cost the team hours every week
For full qualifications & requirements, please visit our careers page