Site Reliability Engineer

Chicago, ILFull-timehybrid$79K – $99K / yrPosted 2026-07-09

Our SRE team keeps production up and keeps engineers sane during incidents. You'll build monitoring and alerting that actually catches problems before customers do, automate away toil, and lead incident response when things go sideways. If you like the idea of turning a 2am page into a permanent fix instead of just muting the alert, this role is for you, and we promise the pager doesn't go off nearly as much as it used to.

What you'll do

  • Build and maintain monitoring, alerting, and dashboards using Prometheus and Grafana
  • Lead incident response and write blameless postmortems after major outages
  • Automate manual operational tasks through scripting and Terraform
  • Define and track SLOs/SLIs for critical customer-facing services
  • Partner with engineering teams to review designs for reliability gaps
  • Improve deployment safety through canary releases and rollback tooling

What we're looking for

  • 4+ years in an SRE, DevOps, or production engineering role
  • Strong Linux systems knowledge and networking fundamentals
  • Experience with Kubernetes in production at meaningful scale
  • Proficiency scripting in Python, Go, or Bash for automation
  • Calm, methodical approach to high-pressure incident response
  • Comfortable being on a rotating on-call schedule

Nice to have

  • Experience with chaos engineering practices
  • Familiarity with incident management tools like PagerDuty or Opsgenie
  • Background in capacity planning for high-traffic systems

About Stacklane Robotics

About Stacklane Robotics: Stacklane builds autonomous mobile robots and fleet software for warehouse picking and inventory operations. Headquartered in Chicago, our team includes robotics engineers, field technicians, and supply chain specialists working directly inside customer warehouses to get automation live and running reliably.

Apply for this roleTakes ~4 minutes · 52 people have applied