Company Description: NStarX builds the platform backbone that powers every product pod — including the ML platform, CI/CD/CT pipelines, and infrastructure that make reliable training, deployment, and monitoring of machine learning models possible at scale.
Job Description: Establishes and operates the ML platform, CI/CD/CT pipelines, and infrastructure that enable reliable training, deployment, and monitoring of ML models for the POD.
Key Responsibilities
- Build CI/CD/CT pipelines for automated model training, testing, and deployment.
- Manage ML platform infrastructure (Azure ML, Kubeflow, MLflow, or similar).
- Implement model registry, versioning, and reproducibility practices.
- Monitor model performance, data/model drift, and infrastructure health in production.
Skills and Experience
- 7+ years of experience in MLOps, DevOps, or ML platform engineering roles.
- Hands-on experience building CI/CD/CT pipelines for automated model training, testing, and deployment.
- Proven experience managing ML platform infrastructure such as Azure ML, Kubeflow, MLflow, or similar tooling.
- Experience implementing model registry, versioning, and reproducibility practices.
- Strong Python skills and familiarity with ML frameworks (e.g., PyTorch, TensorFlow, scikit-learn).
- Solid understanding of containerization and orchestration (Docker, Kubernetes) for ML workloads.
- Experience with cloud infrastructure and infrastructure-as-code (Azure, Terraform, or similar).
- Strong cross-functional communication skills, partnering closely with data science and platform teams.
Good to Have
- Experience with model monitoring and observability tools (e.g., Evidently, Prometheus, Grafana).
- Familiarity with feature stores and data versioning tools (e.g., Feast, DVC).
- Exposure to LLM/GenAI model serving and fine-tuning pipelines.
- Relevant cloud or Kubernetes certifications (e.g., Azure AI Engineer, CKA).
To apply for this job email your details to recruiting@nstarxinc.com
