Data Scientist
Infosys
- Experience: 4–6+ years delivering end-to-end data science projects.
- Education: Master’s or PhD in Data Science, Machine Learning, Statistics, Computer Science, Applied Mathematics, or related quantitative field (required from this level onward).
- Core stack: Python, Spark, Git; ML frameworks; Databricks/MLflow (or equivalent); cloud basics.
- Translate business needs into ML/AI problem statements and measurable success metrics.
- Develop ML models and, where relevant, GenAI components (e.g., retrieval-augmented generation, prompt pipelines) with clear evaluation criteria.
- Run evaluation: offline metrics, error analysis, bias checks, and monitoring baselines; document decisions and assumptions.
- Communicate results and limitations clearly to technical and non-technical stakeholders; support adoption in workflows.
- Python (pandas, numpy) + Git for reproducible development
- Databricks (Notebooks, Workflows) for development and orchestration
- ML flow (experiments, tracking, model registry) for lifecycle management
- Azure (cloud services; where relevant Azure OpenAI and Azure AI Foundry for GenAI build/evaluation)
- Databricks Mosaic AI (including Mosaic AI Model Serving) for GenAI delivery in the lakehouse
- Databricks Vector Search for RAG retrieval patterns
- Unity Catalog for governed data and model access (where applicable)
- Lakehouse Monitoring / model monitoring for quality and drift (where applicable)