
Jeddah
Eazli | Future Living, headquarter in Jeddah, Saudi Arabia is redefining how the Middle East lives, designs, and builds, an AI-native marketplace connecting homeowners with products and home services across Saudi Arabia, the UAE, and Egypt. Built on a multilanguage models and a fleet of specialized AI agents, we’re not a startup chasing trends; we’re architecting the infrastructure layer of the MENA home economy, with ambitions stretching into Europe and beyond. If you want to work on real AI that converses, reasons, and transacts at scale, and you want your work to matter to millions of people making their most personal spaces their own - Eazli is where that happens.
Role Overview:
We are building an AI-native platform where agentic systems, backend services, and real workflows operate together. This role sits at the intersection of ML Systems, Backend infrastructure and Distributed system operations. You will be responsible for ensuring that models, agents, APIs and workflows run reliably, scale predictably and remain observable end to end. This is not a traditional ML/DevOps role, but it is about operating intelligent systems in production.
What You’ll Work On:
1. Design and implement cloud architectures supporting AI/ML workloads and production-grade systems
2. Build and manage ML platforms using AWS services including EC2, ECS/EKS, Lambda, S3, RDS, VPC, and IAM
3. Leverage AWS Bedrock, SageMaker, or similar managed AI services for model training and deployment
4. Use Infrastructure-as-Code tools such as Terraform, CloudFormation, or CDK to automate cloud provisioning
5. Work with vector databases (Milvus, Pinecone, Weaviate) and graph databases (Neo4j) to support retrieval-based and knowledge-driven AI solutions
6. AI/Model infrastructure:
6.1.Deploy and manage Small and Medium Language Models (SLMs)
6.2.Manage external LLM integrations
6.3. Build pipeline for model versioning, evaluation and fine-tuning
6.4. Support RAG systems, embeddings and Vector database infra
7. Agentic System Runtime
7.1.Enable execution of multi-agent workflows
7.2.Enable orchestration layers
7.3.Ensure consistency, fault-tolerance and latency control
8. Design and operate event-driven backend infrastructure – Apache Kafka (or equivalent)
9. Handle async workflows, retries, ordering, idempotency and enable reliable communication between backend and AI
10.Own Kubernetes cluster design, scaling strategies and workload isolation
Support micro-services and model-serving workloads
Design and manage API Gateway between Frontend and Backend layers
Implement routing, auth, rate limiting and service protection at Edge layers
Build end to end CI/CD (beyond basic GitHub Actions) and support multi-service deployment, environment isolation, rollback strategies, secret/vault management, blue/green deployment, feature flag support
Designing, managing and handling data storage infra layer for RDBMS, Vector DB, Document DB, Caching, Object storage.
Instrumentation, Observability and Debugging
You will design and own end to end observability across AI + Backend + Frontend + Infra
Implement structures instrumentation across APIs, async workflows, AI agents and capture request lifecycle, agent decision paths and execution timelines
Design centralised logging – structured logs (JSON) and contextual logging (correlation IDs)
Distributed tracing across tiers and service layers – user journey, agent decisions and backend actions
Build strong monitoring with metrics for system health, API performance, LLM token burns, Queue lag and model latency
Owning and commanding triage & debugging with engineering teams for multiservice failures, AI <> Backend inconsistencies, root cause analysis and replaying failures
You will be managing security and secrets by securing APIs, model configs, infra credentials
You will be implementing RBAC, secret rotation and environment isolation
Must-have:
Expertise in AWS cloud services, EC2, S3, Lambda, ECS/EKS, managed AI platforms, model deployment and distributed systems
Strong hands-on experience with Docker, Kubernetes at Production grade
Experience with Apache Kafka or similar event-driven
Networking protocols, Security concepts, VPC, Load balancers
Strong experience in distributed systems and reliability engineering
Hands-on with CI/CD pipeline - Jenkins, GitLab CI, GitHub Actions, automated testing
Experience with Observability stack (logs, metrices, tracing) – Grafana, Prometheus, ELK, Sentry, Phoenix, Arize, Open Telemetry.
Strong ownership, communication, and code review skills.
You're based in Saudi Arabia and holding transferable Iqama.