Production Engineer - Applied Machine Learning
San Jose, California, United States
USD 156K-387K Entry-level Full Time
Tasks
- Build reliability mechanisms for SLO SLA observability and alerting
- Create automated inspections and pre flight checks
- Develop CI/CD canary releases and auto rollback
- Forecast capacity and manage elastic auto scaling
- Govern GPU CPU storage and network resources
- Handle fault diagnosis incident response and post mortems
- Implement auto healing disaster recovery and incident reviews
- Manage production stability for training and inference systems
- Manage quotas cost attribution and performance tuning
- Orchestrate AML pipelines
Perks/Benefits
- N/A
Skills/Tech-stack
Alerting | Auto Scaling | Auto rollback | Auto-healing | CI/CD | Canary Releases | Capacity forecasting | Cost attribution | Disaster Recovery | Distributed Training | GPU | Incident Response | Inference Serving | Kubernetes | Model Inference | NoSQL | Observability | Online Inference | Online Inference Serving | Orchestration | Parameter Server | Performance Tuning | Postmortems | Quota Management | Resource Governance | SLA | SLO | Scheduling
Education
N/A
Regions
Countries
States
Cities
Related jobs
-
Model Operations Engineer USD 107K-179KALiBi | AWS SageMaker | Amazon Lambda | Apache Airflow | AzureHealth insurance | Holiday pay | Learning and development | Life insurance | Long-term disabilitySenior-level Full TimeUSA-VA-Ashburn2h ago
-
Benchmarking | Computer Vision | Data parallelism | Deep learning | Distributed TrainingMid-level Full TimeSeattle, Washington, United States3h ago
-
Computer Vision | Data parallelism | Deep learning | Distributed Training | Model AccelerationSenior-level Full TimeSeattle, Washington, United States3h ago
-
Software Development Engineer - AI/LLM Network - Global Frontier Tech Research Program - 2027 Start (PhD) USD 202K-368KAI infrastructure | C++ | Cause analysis | Elastic scaling | Fault LocalizationEntry-level Full TimeSeattle, Washington, United States3h ago
-
Distributed Systems | Inference | Machine Learning | Model Serving | MonitoringSenior-level Full TimeSan Jose, California, United States3h ago
-
Cross-layer Optimization | Diffusion Models | Distributed Training | Distributed task scheduling | Heterogeneous computingSenior-level Full TimeSan Jose, California, United States3h ago
-
Cloud Native | Containerization | Distributed Systems | GPU Acceleration | KubernetesCareer growth | Collaborative engineering environment | Global opportunities | Open Source contributionEntry-level Full TimeSan Jose, California, United States3h ago
-
Deep learning | GPU resource management | Language Models | Large Language Models | Machine LearningEntry-level InternshipSan Jose, California, United States3h ago
-
LLM AIOps Development Engineer Graduate (Data Center Networking) - 2026 Start (BS/ MS) USD 122K-256KAPI Integration | Anomaly Detection | Cause analysis | Continuous Monitoring | Data StreamingEntry-level Full TimeSan Jose, California, United States3h ago
-
AI | Big Data | CPU | Cloud Native | Container RuntimeSenior-level Full TimeSan Jose, California, United States3h ago
-
Cloud Native | Cloud-native computing | Containerization | Distributed Systems | GPU AccelerationCareer growth | Collaborative team environment | Open source contributionsEntry-level Full TimeSeattle, Washington, United States3h ago
-
CUDA | Distributed Systems | GPU | Language Models | Large Language ModelsCareer growth | Global team collaborationEntry-level Full TimeSan Jose, California, United States3h ago
-
Backup and Restore | Blob Storage | Cause analysis | Debugging | Distributed SystemsAgile environment | Collaborative team | On-call rotationSenior-level Full TimeSan Jose, California, United States3h ago
-
AI and DB | Automation | Backup and Restore | CPU Optimization | CXLMid-level Full TimeSeattle, Washington, United States3h ago
-
Software Development Engineer - Cloud Native Databases USD 156K-316KBackup and Restore | Blob Storage | Crash diagnostics | Database Schema | Database schema managementCollaborative work environment | Intellectual curiosity | On-call rotationMid-level Full TimeSan Jose, California, United States3h ago
-
Containerd | Docker | Go | Kubernetes | LinuxSenior-level Full TimeSeattle, Washington, United States3h ago
-
Site Reliability Engineer - Data (Seattle) USD 177K-341KApache Flink | Blameless postmortems | Capacity Planning | Incident Response | KubernetesMid-level Full TimeSeattle, Washington, United States3h ago
-
Software Engineer - AI Compute Infrastructure USD 156K-387KCloud Native | Cluster management | Distributed Systems | GPU Acceleration | KubernetesSenior-level Full TimeSan Jose, California, United States3h ago
-
Site Reliability Engineer - Data USD 136K-359KApache Flink | Automation | Capacity Planning | Cloud infrastructure | Incident ManagementMid-level Full TimeSan Jose, California, United States3h ago
-
Automated Deployment | Backup and Recovery | Column-stores | Data Consistency | Database AdministrationSenior-level Full TimeSeattle, Washington, United States3h ago
-
Deep learning | GPU Computing | Language Models | Large Language Models | Machine LearningEntry-level InternshipSan Jose, California, United States3h ago
-
Software Engineer III, AI/ML Machine Learning, Core USD 147K-210KAI/ML | AI/ML Model Deployment | AI/ML model development | Artificial Intelligence | C++Senior-level Full TimeAustin, TX, USA5h ago
-
Software Engineer III, Manufacturing AI Infrastructure USD 147K-210KAgentic AI | Data Ingestion | Data Processing | Data Storage | DebuggingSenior-level Full TimeMountain View, CA, USA5h ago
-
Machine Learning Engineer USD 120K-140KAI Pipelines | AI Workbench | AI endpoints | API Development | Apache KafkaEntry-level Full TimeDenver, Colorado, United States6h ago
-
Data Engineer/ Analyst- LookML USD 90K-118KBigQuery | CD pipelines | CI/CD | CI/CD pipelines | Data GovernanceSenior-level Full TimeMadrid, MD, Spain7h ago