|
Job Description
Role: Senior Consultant, AI/ML Ops Engineer Location: Remote Duration: 9/1/2026-2/1/2027 Rate: $79-$88/hour W2
Description of work / project: We're operationalizing machine learning across NMDP — from survival models that inform donor recommendations to LLM pipelines processing clinical documents and call transcripts. We're looking for a Senior MLOps Engineer who lives at the intersection of data science, DevOps, and platform engineering: someone who can take a model from a data scientist's notebook to a monitored, versioned, cross-account production endpoint with full CI/CD — and own the platform that lets the whole team do the same. This is a hands-on senior role. You'll build and maintain the SageMaker-based training and inference platform, the Terraform that provisions it, and the GitLab pipelines that ship it. You'll also be a force multiplier: setting the standards, patterns, and guardrails other engineers and data scientists build on. Performance Expectations: - Own the ML lifecycle end-to-end — build and operate SageMaker training pipelines and inference endpoints, model registry, versioning, promotion (validation → live), and blue/green deployment with automated rollback.
- Build MLOps infrastructure as code — author and review Terraform for SageMaker, S3, KMS, IAM, CloudWatch, and cross-account roles across multiple AWS accounts (dev/eng/prod).
- Run the CI/CD — design GitLab CI/CD pipelines covering testing, IaC security scanning (Checkov), SAST (SonarQube), dependency checks, Terraform plan/apply gates, and automated model/artifact promotion.
- Ensure train/serve parity — maintain shared feature-encoding pipelines so preprocessing is identical in training and inference; catch drift before it ships.
- Instrument for production — CloudWatch dashboards, alarms, model/data drift monitoring, and endpoint-level observability. Own the on-call story for ML services.
- Harden and govern — least-privilege cross-account IAM, KMS encryption, model artifact signing, secrets management, and cost visibility for GPU/inference and Bedrock spend.
- Operationalize LLM workloads — support Bedrock-based pipelines (batch and real-time), including throughput/cost management and guardrails.
- Partner with data scientists — turn experimental models (XGBoost, survival models, etc.) into reproducible, tested, deployable services, and coach the team on MLOps best practices.
- Bring AIOps to the platform — add anomaly detection on endpoint, drift, and cost signals, and event-driven auto-remediation (auto-rollback, scaling) so operational issues are caught and handled before they escalate.
- Reduce operational toil with AI — use LLM-assisted log/trace analysis and alert correlation to speed up incident triage and root-cause analysis, cutting MTTR and on-call noise.
Required Qualifications - 5+ years in MLOps / ML platform / ML infrastructure engineering, with production ownership of deployed models.
- Deep AWS expertise, especially SageMaker (training jobs, pipelines, model registry, endpoints), plus S3, IAM, KMS, CloudWatch, Lambda, Step Functions, and multi-account architectures.
- Strong DevOps foundation: Infrastructure as Code with Terraform, CI/CD pipeline design (GitLab CI, GitHub Actions, or similar), Docker, and Git-based workflows.
- Strong data science fundamentals — you understand model training, evaluation metrics (e.g., C-index, AUC, calibration), feature engineering, and can reason about why a model behaves as it does in production.
- Proficient in Python (production-grade, tested code) for both ML tooling and services.
- Experience with model monitoring, drift detection, and the operational realities of models degrading over time.
- Solid grasp of security best practices: least-privilege IAM, secrets management, encryption at rest/in transit.
|