Sherlocks.ai is your SRE teammate, handling alerts, conducting RCAs, and planning long-term stability projects. We aim to resolve incidents faster and prevent future outages. Our founding team has deep expertise in scaling startups and AI-driven ventures.
Deploying diverse environments: Deploy and manage a wide range of applications, stacks, and dummy environments that help test, scale, and stress-proof our systems.
Optimising build pipelines: Set up and improve pipelines that make builds faster, smarter, and more efficient ensuring customers get the best experience without delays or overhead.
Streamlining release workflows: Design and maintain pipelines for Sherlocks agents, making it seamless to roll out updates and new features reliably and without downtime.
Managing monitoring & alerting: Set up and refine monitoring, alerting, and observability tools.
Driving chaos experiments: Contribute to our chaos engineering framework by simulating failures in applications and services
Improving investigation pathways: analyse how Sherlocks investigates incidents and suggest improvements that enhance both performance and reliability.
K8s, IaC, AWS primitives, Go/Python
Small team, extreme ownership, high-leverage work
You'll work across a wide array of tools and systems.