Data Engineer — Registry Data Pipelines (MCA / ROC)

Verizol

About the role

Verizol is building India's corporate-registry data layer — company search, daily new-incorporation alerts, registry watchlists, and KYB verification APIs, all sourced from MCA/ROC public records.

Everything we sell depends on one thing: a pipeline that reliably knows which companies and LLPs were incorporated yesterday, and gets that into our database before our customers wake up. You own that pipeline end to end.

This is not a dashboard-building role. It is acquisition, resilience, and data quality on a source that actively resists automation and changes without notice.

What you'll own

Daily incorporation feed

  • Build and run the ingestion that captures newly registered companies and LLPs from MCA every single day, with a defined freshness SLA (new entities live within 24 hours of appearing on the registry).
  • Implement multiple independent discovery paths so a single upstream change doesn't blind the product — trailing-window registry listings, sequential CIN/LLPIN identifier discovery, and reconciliation between them.
  • Detect and alert on silent failures: a pipeline that returns zero new rows is a P1, not a quiet day.

Historical corpus and backfill

  • Seed and maintain the full company/LLP master-data corpus (3M+ entities) from bulk open-data sources, and reconcile it against incremental daily pulls.
  • Own re-crawl strategy for stale records.

Change data capture

  • Track and diff registry changes over time — director appointments and resignations, filing activity, charge events, status changes (active / strike-off / under liquidation) — to power watchlist and monitoring products.

Scraper resilience

  • Session management, proxy rotation, rate limiting, retry and backoff, CAPTCHA handling, request fingerprinting.
  • Monitoring for upstream schema and endpoint changes, with fast turnaround when the portal shifts (it will, several times a year).
  • Keep crawl volume polite and defensible — we do not want to be the reason an endpoint gets locked down.

Normalisation and data quality

  • Transform raw registry output into our standard entity schema: CIN/LLPIN parsing, state and ROC mapping, city and PIN normalisation, NIC-code to industry mapping, capital and date parsing, address cleanup.
  • Deduplication and entity resolution across companies, directors (DIN), and addresses.
  • Automated data-quality checks: completeness, field-level validity, row-count anomaly detection, referential integrity between entity and director tables.

Serving the product

  • Land clean, query-ready data into Postgres and the search index that powers company search, alerts, and the KYB/MCA APIs.
  • Work with backend on schema contracts so API latency and correctness don't degrade as volume grows.
Requirements

Must have

  • 2+ years building production data pipelines in Python.
  • Strong SQL and hands-on PostgreSQL (indexing, query tuning, schema design for read-heavy workloads).
  • Real experience with large-scale web scraping — not tutorial-level. You've dealt with rate limits, blocks, dynamic pages, and sources that broke your parser at 3 a.m.
  • Working knowledge of an orchestration tool (Airflow, Prefect, Dagster, or similar) and comfort owning scheduled jobs in production.
  • Linux, Docker, Git, and a cloud environment (AWS preferred).
  • Ownership instinct. This pipeline has no fallback — if it's down, the product is down.

Good to have

  • Prior work with Indian registry, KYC/KYB, or financial data (MCA, GST, GSTN, credit bureau, Aadhaar/PAN verification stacks).
  • Elasticsearch / OpenSearch for search and faceted filtering.
  • Entity resolution or record-linkage experience.
  • Awareness of DPDP Act obligations and MCA terms of use as they apply to public-record data and director contact information.
  • Experience with dbt, Redis, or queue systems (Celery, SQS, RabbitMQ).
Typical stack

Python · Scrapy / Playwright / httpx · PostgreSQL · Redis · Airflow · Elasticsearch · Docker · AWS (EC2, S3, RDS) · Grafana or equivalent for pipeline observability

What success looks like

TimeframeOutcome30 daysYou understand every existing ingestion path, have documented the failure modes, and have alerting in place so we find out about breakage before a customer does.60 daysDaily incorporation feed running with a redundant discovery path and a measured freshness SLA. Data-quality checks automated.90 daysChange-data-capture live for director, filing, charge, and status events. Backfill reconciled. Pipeline is boring and you're building the next layer.

Why this role is interesting
  • You're not maintaining someone else's warehouse. You're building the acquisition layer for a product where data freshness is the moat — every competitor sells the same fields, and the one that gets them first and cleanest wins. The engineering problem is genuinely adversarial and genuinely measurable.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.