SRE Engineer
Company Description:
At NStarX, we live by our mission: being the Lodestar to Success for enterprises navigating the rapidly evolving digital landscape. We build cutting-edge, production-ready platforms that integrate advanced AI, cloud-native architecture, and robust data engineering to solve complex business challenges at scale. If you thrive in a collaborative environment that values technical excellence, architectural ownership, and production-first innovation, NStarX is where your impact will shine. Job Description: We are seeking a highly skilled SRE Engineer Platform / SRE Lead to own, architect, and secure the foundational cloud infrastructure for our enterprise AI platform. In this role, you will lead the charge through our production launch, ensuring high availability (HA), disaster recovery (DR), automated DevSecOps pipelines, and robust identity integrations You will collaborate closely with development teams to deliver a resilient, scalable, and secure environment capable of supporting enterprise-grade AI workloads
KEY RESPONSIBILITIES
- Infrastructure Ownership: Design, build, and maintain high-availability cloud infrastructure on AWS, ensuring the platform is ready for production launch
- Container & Gateway Management: Manage core container platforms (EKS/ECS) and support the deployment of critical services like LLM gateways.
- DevSecOps Integration: Design and own automated CI/CD pipelines across multiple teams, enforcing quality gates and integrating security scanning tools.
- Identity & Security: Implement robust enterprise identity and SSO integrations, applying strict secrets management discipline across all environments.
- Reliability & DR: Define and execute disaster recovery (DR) strategies, backups, automated scaling, incident management protocols, and runbooks.
MINIMUM REQUIREMENTS
- Production SRE Leadership: Hands-on experience and ownership of production cloud platforms in an SRE, Platform Engineering, or DevSecOps capacity.
- Deep AWS Expertise: Experience managing core AWS services including ECS/EKS, ALB, Route 53, CloudFront/WAF, RDS, ElastiCache, S3 (including object lock), Secrets Manager, ECR, and IAM.
- Kubernetes Operations: Strong experience running production Kubernetes clusters (EKS preferred), utilizing Helm for deployment, and configuring autoscaling (HPA required; KEDA is a major plus).
- Infrastructure-as-Code (IaC): Deep proficiency with Terraform (or an equivalent production-grade IaC tool).
- CI/CD Pipeline Design: Proven capability in designing, implementing, and scaling multi-team CI/CD delivery pipelines.
- Automated DevSecOps: Direct experience embedding security tools into pipelines (SAST, dependency, and container scanning using Snyk, Trivy, SonarQube, or similar) alongside automated quality-gate enforcement.
- Identity & Access Management: Hands-on integration of enterprise SSO/Identity protocols (OIDC/SAML) against enterprise IdPs like Entra ID.
- Resiliency & Incident Response: Proven track record in designing DR/backup execution plans, alongside creating runbooks, and managing active production incidents.
GOOD TO HAVE
- AI Platform Operations: Experience deploying or operating enterprise LLM gateways (such as LiteLLM, Kong, Portkey, or Azure APIM).
- Zero-Trust Hardening: Deep understanding of Zero-Trust network architectures and production network hardening.
- Supply-Chain Security: Familiarity with container/image hardening and software supply-chain security initiatives (e.g., SBOM generation, image signing).
- Multi-Cloud Portability: Architectural awareness or direct experience supporting multi-cloud portability (particularly running Azure alongside AWS).
- Cloud FinOps: Experience implementing cost-aware infrastructure practices (tagging enforcement, automated budget alerts) to support enterprise FinOps reporting.
To apply for this job email your details to recruiting@nstarxinc.com
