Our SRE team keeps production up and keeps engineers sane during incidents. You'll build monitoring and alerting that actually catches problems before customers do, automate away toil, and lead incident response when things go sideways. If you like the idea of turning a 2am page into a permanent fix instead of just muting the alert, this role is for you, and we promise the pager doesn't go off nearly as much as it used to.
What you'll do
- Build and maintain monitoring, alerting, and dashboards using Prometheus and Grafana
- Lead incident response and write blameless postmortems after major outages
- Automate manual operational tasks through scripting and Terraform
- Define and track SLOs/SLIs for critical customer-facing services
- Partner with engineering teams to review designs for reliability gaps
- Improve deployment safety through canary releases and rollback tooling
What we're looking for
- 4+ years in an SRE, DevOps, or production engineering role
- Strong Linux systems knowledge and networking fundamentals
- Experience with Kubernetes in production at meaningful scale
- Proficiency scripting in Python, Go, or Bash for automation
- Calm, methodical approach to high-pressure incident response
- Comfortable being on a rotating on-call schedule
Nice to have
- Experience with chaos engineering practices
- Familiarity with incident management tools like PagerDuty or Opsgenie
- Background in capacity planning for high-traffic systems
About Fieldrow
About Fieldrow: We build sensor hardware and forecasting software that helps row-crop farmers make better irrigation and fertilization calls. Based in Denver, our team splits time between the office and the fields our Probes are installed in, and we measure success in bushels saved, not just software shipped.
Apply for this roleTakes ~4 minutes · 36 people have applied